AIHOY and happy weekend, AInauts!
Maybe you missed some of last week's AI news, tools, and hacks, or you're new here. No worries!
Here are our highlights from the past week:
To wrap up, we've got the most important Quick News, so you can catch up on the developments that matter in one place. Ready? Let's go!
AI research, explained in one morning email
Most of what shapes AI next year is sitting in a research paper today. Keeping up with them is a full-time job, and the abstracts don't help.
TLDR AI is the newsletter that does the translating. Each morning you get the handful of papers worth knowing about, curated by Anthropic and ex-Google engineers, with a plain-English summary of what they found and why it matters.
You come away smarter about AI research in the time it takes to finish your coffee.
Free, and delivered to 1.1M+ inboxes every morning.
Press a key, speak, release: the cleaned-up text lands wherever your cursor is. We know the routine from Wispr Flow, which we've used to dictate more than a million words. Now there's our own version: AInauten Voice.
Free for Mac, with speech recognition and cleanup running locally. Your audio never leaves your computer, and you can import your Wispr settings. In the article, we also show how we built the app with AI in a few hours, including tests and code review.
The Agentic Engineering Playbook top engineers are using
Most engineers are stuck babysitting AI one prompt at a time. A handful have figured out how to build agents that plan, execute, and self-correct on their own. Sign up for The Code and get The Complete Playbook for Agentic Engineering. Claude Code, subagents, MCP, all of it.
A video call used to be fairly solid proof that there was a human on the other end. Tavus is shaking that up: its Griffin avatar listens, recognizes pauses, and responds without constantly talking over you. All it takes is a photo and ten seconds of voice.
In the vendor's test, 26 of 54 people thought the person opposite them was real. Time for a new ground rule: if an unusual request for money or access arrives by video, call back through a known channel. Even if the face looks familiar.
Most apps tell you what they can do through their menus. An AI model doesn't. First, you need an idea of what you could build with it. That's where many people get stuck, long before the first line of code.
Our suggestion: pick an annoying routine from your everyday life. Write down what goes in, what needs to come out, and how you'll know it works. Then let AI build it and check the result. A small tool that saves you work every week beats the app idea you've been polishing in your head for months.
AI doesn't need to write a dissertation in its head to make an email friendlier. In our test, the highest thinking budget cost about 25 times as much as the lowest. The result was similar. The wait wasn't: around two minutes instead of seven seconds.
Invoices were different: a small model missed the VAT issue with a low thinking budget, but got it right with more thinking. That gives us a useful starting rule: strong model, lower effort. Small model, higher effort.
Market note: this sponsored sevdesk example concerns our German company and German bookkeeping rules.
Twenty-eight receipts, eight booked, three duplicates found. The other seventeen stayed open because the actual euro amount of the foreign-currency payment was missing. Exactly what we wanted: AI didn't replace missing information with a plausible guess.
Claude reads and sorts the receipts, sevdesk sets the accounting rules, and a human does the final check. Our tip: see what your bookkeeping tool already does before putting a second accountant made of tokens next to it.
Mistral Large 4 is available as an API preview, with model weights expected at the end of the month. Aleph Alpha is back with Kolibri, an open model for German and English. Both are interesting if you want more control over where your AI runs and who operates it.
Mistral offers a European alternative with its own infrastructure. Kolibri still lacks independent tests. So look at your use case: is the quality good enough for documents, support, or internal workflows? A European passport alone doesn't get a task done.
AI News Quickie: highlights from the industry
AI never sleeps!
GPT-6 reaches regular chat, while Haiku 5.5 lowers the cost of routine work. Agents are also moving deeper into documents, spreadsheets, and meetings.
🧠 OpenAI and ChatGPT
GPT-6 is now for everyone in ChatGPT: paid plans get Sol, while Free and Go get Luna.
There's also the Intelligent UI, which builds answers as interactive comparisons, calculators, or forms. Checking whether the numbers are right is still your job.
Paid ChatGPT plans now let you upload audio files, transcribe them, summarize them, and ask questions about them. Availability depends on your region and app version.
GPT-6.1 Sol Ultrafast brings a faster mode to Codex and Work, initially for the $500 Pro subscription and eligible Enterprise and Edu accounts.
The OpenAI API is moving from five usage tiers to three: Build, Launch, and Grow. If you run your own apps, check where your limits now stand.
ChatGPT is testing ads during image creation, initially in the US for Free and Go. It isn't an issue for us in the German market yet, but the direction is clear: free users will soon see more advertising.
🌐 Google and Gemini
The models in personal Gemini accounts are changing: Free defaults to Auto and mostly gets Flash-Lite, while AI Pro and Ultra have all models.
Nano Banana 2.1 is generally available in the Gemini API. Image edits are meant to keep characters, instructions, and lettering more consistent, at significantly lower cost per image.
At Gemini at Work, Google is showing a business agent that follows tasks across Gmail, Drive, Docs, and Calendar and can also use Claude models. When a project hits its spending limit, the agent pauses.
Foresight is an experimental Mac app that transcribes meetings locally and turns your keywords into notes. Audio and transcripts stay on your computer.
🤖 Anthropic and Claude
Haiku 5.5 is here: a fast, small Claude model for summaries and routine work. Up to 100,000 input tokens, a million tokens cost $0.10 in and $0.50 out; above that, five times as much. Watch out: the model thinks at medium effort by default, so our effort rule applies here too.
Sonnet 5.5 gets half-price cache reads. That makes automations with lots of reused context cheaper, but not your Claude subscription.
Max and Team subscriptions get monthly API credit: $100 for Max 5x, $200 for Max 20x, and up to $500 shared across Teams.
Docs, Slides, and Design are leaving beta and are now included in all plans, even Free. There's also better exporting and collaborative editing.
New dashboards turn business data from Salesforce, Snowflake, or BigQuery into views you can refresh.
Motion builds animations from code and exports them as MP4, initially in beta for Team and Enterprise. Think explainer video rather than Hollywood.
Still working on the separate Claude Design site? Plan your move by December 14. You can take your design systems with you, but not old chats, comments, or public links. Have we mentioned files over tools?
Claude now works directly in Google Docs, Sheets, and Slides, in public beta for paid plans. By default, Claude asks before every change.
💼 Companies, startups & creative tools
In Copilot Studio, you can now build apps and workflows from a description, initially in public preview. Take onboarding: the form, data, and agent are created together.
Reflection is announcing Beam, an open model for coding agents. Weights are expected at the end of October. Until then, there's just a waitlist.
Midjourney is testing a Thinking mode: on the alpha website, restart existing image jobs to get complex instructions, lettering, and composition closer to what you wanted. Compare using the same subject.
TikTok is adding tools for advertisers: an AI shopping assistant is meant to guide people from search to purchase, while new Smart+ tools put your own, AI-generated, and creator ad assets into one approval workflow. When and where this starts is still unclear.
Uber and Pony.ai plan to test robotaxis in London in the coming weeks. They've been operating in Zagreb since August, with 2,000 robotaxis planned for Europe. Driverless transport is getting closer to home for our European readers.
🔐 Security, privacy & research
OpenAI is introducing invisible watermarks for AI text because of the EU AI Act's transparency rules. In the coming weeks, ChatGPT and Codex text in the EU will carry the signal across all plans.
The SynthID Detector is now available worldwide in English and checks images, video, and audio for Google's watermarks. OpenAI has a similar tool, Verify, for its own images and audio.
Norway wants to temporarily ban camera-equipped AI glasses in many public places: parks, beaches, museums, shopping centers, and events, potentially also schools, doctors' offices, and gyms.
Claude is getting new usage rules. Important for businesses: when Claude helps make decisions about health, finances, or rights, a human must be able to review and override them, and affected people must know AI was involved. Extreme, sustained torment of the model is also newly prohibited.
The Anthropic Cyber Mission offers critical open-source projects free security scans with suggested fixes. That still isn't a reason to apply an unchecked patch.
OpenAI has stopped disguised influence operations from Russia and Iran, using invented journalists and think tanks to slip articles into real media outlets.
A sixteen-year-old had to be rescued from the mountains in Canada after a Claude route description led him to the wrong place.
OpenAI is publishing new math results from an internal model on open problems, with many proofs formally checkable in Lean. The math community's verdict is still pending.
A UV map of the sky built with Claude combines existing data and predictions. Around a third is estimated, and the researcher had to spot errors himself.
Google Earth AI helps estimate disease risks from satellite and health data. During the Ebola response in Congo, a prototype found 48 remote settlements with around 45,500 people at risk.
Made it! No reason to be sad: we'll be back on Monday with a fresh round of news, hacks, and insights.
Your AInauten
Fabian & Reto




