In partnership with

AI-HOY, AInauts,

Welcome to a new edition of your favorite newsletter.

Whew, this week has brought another flood of updates, and we can barely keep up with testing everything.

New models and AI agents are moving in everywhere. Even into your car.

So today, we have the news, the practical implications, and the question of what all this means for your setup.

Here are today's topics:

  • ⚑ First they shout β€œslow down,” then they ship: Opus 5.5 vs. GPT-6

  • πŸ› οΈ Four rules for getting the most from the new models

  • πŸš— Agents everywhere: Why your context needs to belong to you

Let's go!

Some teams never seem to stop moving. They're on Attio, the agentic CRM.

Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.

With Attio, you’ll get:

  • Leads automatically prioritised and routed to the right rep

  • Expansion and risk signals caught the moment they land

  • Follow-ups written in your voice, already there when you arrive

Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?

⚑ First They Shout β€œSlow Down,” Then They Ship: Opus 5.5 and GPT-6 Are Here

Last week, we were writing about the big AI safety debate.

Everyone was shouting β€œSLOW DOWN.” Anthropic CEO Dario Amodei wrote a long essay about it.

Its title: β€œWe Must Pace the Frontier.” His argument was that AI labs should reduce the pace so safety research can catch up.

Sam Altman responded the same day: β€œI agree with Dario.” Elon Musk was even shorter: β€œDario is right.”

AI leaders agreeing that frontier development should slow down

Ten days later. Tuesday.

Anthropic ships Claude Opus 5.5. OpenAI ships GPT-6 Sol and GPT-6 Luna on the very same day. 😁

To be fair: None of these models raises the absolute frontier. Opus 5.5 is roughly as capable as Anthropic's large Fable 5.1 model, but significantly cheaper. Sol and Luna bring technology from the large GPT-6 Astra model into more affordable models.

So development at the frontier may genuinely be slowing down. But across the broader market, everyone is stepping on the gas.

For users, that is great news. Cheaper, faster, and more capable is exactly what we need in daily work.

Opus 5.5: We Are Friends Again

We have written before that we never really warmed to Opus 5.

For many tasks, we stopped using Claude entirely, or used only Fable instead. That was not exactly cheap.

Opus 5 argued when you wanted something changed. Its explanations needed explanations. Simple tasks became projects, and it made too many mistakes.

And wow, the writing style was bad.

We were not alone. Many power users moved to ChatGPT Work and Codex during the past few months.

And now you read it everywhere, or at least all over X: They are coming back.

What Anthropic promises:

  • Fable-level performance on most tasks

  • 40% cheaper in daily use than Opus 5 and 30% faster

  • Clearer writing: the key point first, with less jargon

  • Higher limits for Pro, Max, and Team users

After only a few hours of work, we can confirm much of that. It is faster, clearer, smarter, and uses surprisingly little of the usage limit.

Claude Opus 5.5 usage and model interface

Of course, the first annoying traits are already visible:

  • It does not know when to stop. Without a boundary, it keeps building and building, even when it would be useful to ask whether the direction is still right.

  • Still too much prose. The writing is strong, but the core points are often wrapped in paragraphs before and after.

  • Green does not always mean correct. We ask AI to verify everything it builds. Even when it says all tests passed, we still find errors.

We have practical fixes for those problems below.

GPT-6 Sol and Luna: Half the Price

OpenAI launched GPT-6 Astra, its large model, at the start of the month. It is incredibly strong at computer use, extremely smart, and also extremely hungry for tokens and money.

Sol and Luna are the faster, cheaper siblings.

  • Sol is for complex work, coding, and agents. In one benchmark for business workflows, OpenAI says it beats the old Opus 5 at 9% of the cost.

  • Luna is for volume: automations, straightforward tasks, and everything that runs frequently.

  • Both cost half as much as their predecessors and respond more briefly and clearly.

Language note: The following price-comparison graphic is in German; the key figures are translated above.

German-language price comparison for Claude Opus 5.5 and GPT-6 models

German-language price comparison; key figures are translated in the text.

Both models are available in ChatGPT Work and Codex for all paying users.

One small catch: OpenAI compares Sol with the old Opus 5, while Anthropic compares Opus 5.5 with the old GPT-5.6 Sol.

Neither company provides a real Opus 5.5 versus GPT-6 Sol comparison. Of course. 😁 That is why we are here. But it is genuinely difficult right now because both models are very strong, and for ordinary tasks the difference is often hard to notice.

What We Would Use for What

Based on our first tests and what we are seeing from others:

  • Opus 5.5 when you want an AI that actively thinks with you: plans, opinionated analysis, prototypes, designs, and long agent runs.

  • GPT-6 Sol when there is a deadline, you need a strict editor, or the task needs to run cheaply. We still prefer it slightly for writing.

  • Fable and Astra only for the genuinely difficult problems.

Our Take: You Cannot Really Choose the β€œWrong” Model

We are very happy that Anthropic has caught up again. OpenAI still offers the best user experience and clearly had the better models during the past few weeks.

Now the balance is shifting again, and Claude is absolutely back. The best part: Whatever you are using right now, both options are excellent.

Learn AI in 5 minutes a day

You don't have to scroll every AI thread, track every new tool, or watch every demo.Β 

The Rundown AI breaks it all down for you β€” the latest AI news, tools, and tutorials in one free 5-minute email every morning.Β 

Trusted by 2M+ professionals at Apple, Google, and NASA.

πŸ› οΈ Four Rules for Getting the Most From the New Models

A new model often means that old prompts, instructions, and skills need a quick review.

In August, we removed old β€œcheck and think” instructions. Last week, we recommended no more than three rules per prompt.

Anthropic and OpenAI both published guides for their new models this month. The good news for us: They say almost the same thing.

The models work independently for longer, think on their own, and check their own work.

You no longer need to explain how they should work.

You need to explain what finished looks like and when they should stop.

1. Say What β€œFinished” Looks Like

Put the full task in one message. Add the finish line. Then let the model run.

Turn these three Excel lists into one customer list. Finished means: one file, no duplicates, every company includes a contact and the most recent interaction. Ask me only when two records contradict each other.

You can finally remove β€œthink carefully” from your prompts. Opus 5.5 thinks before every answer by default. According to Anthropic, the response is faster without that line and no worse.

2. Say When It Should Stop

The new models, especially Opus 5.5, keep going and going. That is useful for large projects and expensive when you only wanted a quick answer.

Anthropic recommends adding a stopping rule to CLAUDE.md, or to AGENTS.md for ChatGPT Work and Codex. Here is our English version:

If a step does not need my input, continue. Ask only when you cannot proceed without me or before doing something irreversible: deleting data, sending something, or changing anything outside this project.

When time is short, tell the model what comes first: β€œDo the timeline first. Everything else only if time remains.”

Important: Keep approval gates active for critical actions. The rule does not replace the seat belt.

3. Say Exactly What You Do Not Want

β€œDo not make it generic” barely helps, according to Anthropic. The model simply swaps one default for another.

A list works better. For design work, Anthropic suggests something like:

No cream background, no italic accent words in headings, no β€œ01 / 02 / 03” labels, and no rounded pill buttons.

If you have built a landing page with Claude during the past few months, you know every one of those habits. πŸ˜‰

OpenAI does the same for writing and even published a list of slop words: β€œdelve into,” β€œit's worth noting,” and contrast clichΓ©s such as β€œnot X, but Y.”

Our tip: Maintain your own list. Whenever something annoys you, add it.

4. Send Your Skills Through an Inspection

This is the step most people forget. Your skills and instructions were written for older models. For newer models, they are often too narrow, too long, or internally contradictory.

Claude Code now includes a command for this: /skill-doctor. It shows which skills you actually use, which you never use, and how much context each one consumes.

We ran it on our own setup yesterday.

36 skills loaded. Eleven had not been used for ages, according to the report. One was installed twice.

Claude Code skill doctor report

If you only see β€œUnknown command,” update Claude Code first.

OpenAI recommends a smart way to improve the contents of your skills: Let the model clean them up.

Read my CLAUDE.md or AGENTS.md and every skill in this folder. Mark each instruction that is unnecessary, contradictory, or too narrow for a current model, such as β€œthink step by step,” β€œcheck your work,” or vague triggers in the description. Quote every passage and propose a shorter version. Do not change anything yet.

The same approach works in ChatGPT Work and Codex.

Our Take: Define β€œFinished” Clearly

We used to explain to AI how it should work, step by step.

Now we describe exactly what finished looks like and let the model run.

That sounds like a small change, but it creates a very different way of working. Less like a boss directing every hand movement and more like a boss defining the outcome, occasionally correcting the direction, and saying when the workday is over.

To be fair, not every rule applies to every model. Smaller models such as GPT-6 Luna may still need more guidance, according to OpenAI.

That is often how it works: The more capable the model, the less direction it needs.

πŸš— Agents Everywhere: Why Your Context Needs to Belong to You

To close, here is a topic that connects directly to the model changes above.

The race for AI, and for your personal AI agent, is still wide open. It will probably stay that way for a while.

Since Tuesday, Grok, SpaceXAI's assistant, can manage your email, organize your calendar, and review your files inside a Tesla. By voice, while you drive.

Grok Bot, SpaceXAI's agent, can do even more: order your usual coffee, reserve a table, or make an appointment.

For now, it is available only in the US and requires the $300-per-month SuperGrok Heavy plan. But the direction is clear.

We tested Grok Bot in depth in August. Now it is sitting in the car with you.

The Race for Your Personal Agent

Grok is not alone. Within only a few weeks, we saw:

  • Meta Muse, Meta's personal AI agent

  • Instinct, which messages you proactively and can even make calls for you

  • The new Siri, although it is not yet available on iPhone in the EU

  • OpenAI, which is also preparing a release in this direction

Agents are moving beyond the laptop, into phones, WhatsApp, and cars.

We no longer use one agent. We use several: Instinct and Muse for personal tasks, Claude Code and ChatGPT Work or Codex for complex work, and Grok Bot for experiments.

Honestly, we think it is unlikely that one company will ever cover everything.

A landscape of personal AI agents across devices and services

The Agent Is Replaceable. Your Context Is Not.

Axios summarized it well this week: The fight for your personal agent is also a fight over who manages your personal data.

The more an agent knows about you, the better it becomes, and the harder it becomes to leave.

That is why we have been preaching Files over Tools for years.

Your context should live in simple text files: who you are, what you are working on, who you work with, and what matters to you. Keep it in folders, in a place you control.

Then every new agent can connect in minutes and immediately understand your world.

Our own proof came this week. When Opus 5.5 launched on Tuesday, our Second Brain was running on it that same evening.

No export and no migration. Codex reads the same folder. If a better agent arrives tomorrow, it can read it too.

Several agents can work from the same context only when that context is not locked inside one tool.

Multiple AI agents using the same user-owned context

Our Take: Your Context Belongs to You

Models are changing every few weeks. Agents are changing too.

If you build your context inside one tool, you start from scratch every time you switch. If you own it yourself, you simply plug in the next agent.

The agent is rented. The context belongs to you.

That is it for today. Thank you, as always!

See you next time.

Reto & Fabian
AInauten Team

X Β· LinkedIn Β· Facebook Β· Instagram Β· YouTube Β· TikTok

AInauten

Login or Subscribe to participate