Hey AInauts,
Welcome to a new edition of your favorite newsletter. Today's might be a little more advanced than usual.
We're talking money. Specifically: what do AI models actually cost, and how do we protect our usage limits as those limits keep getting tighter?
We gave different models the same tasks 100 times and discovered how wasteful some of our own habits have been.
We're also kicking off a short series on bookkeeping with AI. And Europe is back with two new models.
Here's what's inside:
ποΈ AI costs, money, and limits (and why they matter)
π§Ύ Bookkeeping tools & AI? Part 1 of our series
πͺπΊ Europe is back: Mistral Large 4 & Kolibri
Let's get into it!
AI research, explained in one morning email
Most of what shapes AI next year is sitting in a research paper today. Keeping up with them is a full-time job, and the abstracts don't help.
TLDR AI is the newsletter that does the translating. Each morning you get the handful of papers worth knowing about, curated by Anthropic and ex-Google engineers, with a plain-English summary of what they found and why it matters.
You come away smarter about AI research in the time it takes to finish your coffee.
Free, and delivered to 1.1M+ inboxes every morning.
ποΈ AI costs, money, and limits (and why they matter)
We write a lot about different AI models. If you don't spend all day in the AI bubble, it can get confusing.
Different providers offer different models, and each model can be told to work at different levels of intensity.
And every combination costs something different. Confusion guaranteed.
But we're seeing costs and limits become increasingly important.
The way you and we use AI today is usually heavily subsidized. We're not paying the real cost.
We think those subsidies will eventually end. Understanding the different layers matters, so that's what we're digging into today.
On Monday, we talked about choosing the right model for the right step. Today, it's about the second dial many people overlook: effort.
We run almost everything on Extra High. That's our default for Claude Opus 5.5. For quick jobs in between, we use Sonnet 5.5. Also on Extra High. π
In ChatGPT, it's GPT-6.1 Sol, also on High or Extra High.
It feels thorough. It is. But it's expensive.
Cheap isn't always cheap
Two weeks ago, we wrote that GPT-6 Sol is the model to use when you want a lower-cost option than, say, Claude Opus 5.5.
Per token, that's true. But tokens are often an abstract comparison. "Cheap tokens mean inexpensive results" can be true. Quite often, the opposite happens.
Researchers at Stanford and Berkeley ran the numbers: in one out of three comparisons, the model with the cheaper listed price ended up costing more.
SemiAnalysis measured how much work you get from your subscription this week.
Take the midrange models: for $100 a month, Claude gives you Opus 5.5 API tokens worth $5,725. With ChatGPT, it's $1,055 worth of GPT-6.1 Sol.
You might do some back-of-the-envelope math and say, "So I get roughly five times as much from Claude as from ChatGPT." But that isn't right either, because Anthropic and OpenAI charge different API prices.
Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. GPT-6.1 Sol costs just $2 and $10, respectively. The same token usage already has twice the API dollar value with Claude, without getting you twice the work.
Now OpenAI is also halving the limits on its $200 subscription. We mentioned that in last week's P.S.

On top of that, each model spends a different amount of time thinking about the same task.
Bottom line: you can't simply turn the price per token into a price per task. Each model needs a different number of tokens to complete it.
We tested it. 100 times.
Those were the token costs of the different models. To make things even more complicated, there are effort settings too.
We don't want to confuse you. We want practical ways to get more out of your subscription. So we gave different models the same tasks at different effort levels:
Make a customer email friendlier.
Check an expense report with 30 receipts.
With traps built in: a duplicate receipt number, a taxi on a Sunday, a hotel billed in Swiss francs. And a printing invoice whose total was right but whose VAT wasn't.
Here's what happened (this is interesting):
The same email costs 25 times as much on Max effort as on Low. 14,698 tokens instead of 435, almost two minutes instead of seven seconds. The result? Practically identical.
Opus 5.5 checks the expense report without errors even on Low. Max costs 7.6 times as much. Nothing gets better.
Gemini 3.8 Flash is 62% cheaper per token than GPT-6.1 Sol. Per task, it's 73% more expensive. It simply thinks five times as long. π
The small GPT-6 Luna fails on Low and Medium. It misses the printing-invoice trap every time. On Extra High: three out of three correct. For 0.27 cents. That's 28 times cheaper than Opus on Low.
How we tested: through the API via OpenRouter, using the same task every time and changing only the model and effort. Small samples, two to three runs per level. In the apps, your subscription may give you fewer options. Luna and its effort levels are available in ChatGPT Work and Codex, for example. And with a subscription, what counts isn't the dollar amount but how quickly your usage meter rises.
Two dials you need to know
Enough numbers. How do we get more out of our subscriptions?
1. The model
Pick one for each task and stick with it. Each model has its own token cache. Switch models halfway through a chat, and the new model reads the entire conversation again. At full price!
Switching between two steps, as we described on Monday: fine. Switching in the middle: better not. For example, make a plan with the strongest model, then switch to a midrange model for implementation.
2. Effort
In Claude, you'll find it in the model menu next to the send button. In ChatGPT, it's called Thinking and ranges from Instant to Pro (or Ultra, which you first need to enable in settings).
Important: effort doesn't control how smart the model is. It controls how thoroughly it checks its work.
You're giving the model more time to think. That helps catch hidden edge cases. But only when there are any.
Low: emails, translation, rewriting, brainstorming. Anything you can immediately check yourself.
Medium: the default for everything else.
High or Extra High: when mistakes are expensive and hard to spot. Contracts, numbers, expense reports.
Max: only when the AI needs to work on its own for a long time.
For tricky tasks, this can make a difference. For a friendlier email, it often just takes longer. Our rule of thumb from the test: strong model, lower effort. Small model, higher effort.
Our take: find the sweet spot for your tasks
It's a complex topic, but one we all need to understand. Especially as the labs start cutting back their subsidies.
For our long agent runs in Claude Code, Extra High can stay. Nobody checks in halfway through. That's exactly what it's for.
But for quick jobs, we're burning through our limits. Sonnet on Extra High costs as much for an email as Opus on Low. And 2.5 times as much as Sonnet on Low. Same result. So we're switching to Low.
For you, that means: experiment more with the effort settings. It'll pay off.
There are plenty of other levers too: old chats, screenshots, projects, connectors. We'll put those details into a separate Deep Dive.
200+ Proven Ways to Make Money With AI in 2026
The next wave of millionaires will be people who figured out how to make AI work for them.
The window to get ahead is still open. But not for long.
Here are 200+ proven ways to make money with AI in 2026.
Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.
Advertisement | With our partner sevdesk
π§Ύ Bookkeeping tools & AI?
Market and language note: this example, the sevdesk plans and discounts, and the linked setup guide are for the German market. The linked pages are in German. VAT filing and DATEV references below concern our German company, not US bookkeeping rules.
Last week, we wrote about the tools we're canceling. Gamma was the latest.
But there was something on our "keep" list too: bookkeeping. That's where the agent connects to the tool. It doesn't replace it.
Bookkeeping with AI is a recurring topic among the solopreneurs and freelancers in our community.
So today we're starting a short, three-part series showing exactly how it works at one of our small companies.
We're doing this with today's partner, sevdesk, which we've been using for that company since the start of the year.
Don't worry: part one is a little more advanced. The next installments will be simpler.
Last week's bookkeeping run
Late September. Twenty-eight receipts in the incoming folder on our Google Drive. Software subscriptions, telecom bills, parking, public transit tickets. The usual.
We've had a skill for the initial bookkeeping run for a while. We type /belege-buchen. Done.
We connected our sevdesk account to Claude through the sevdesk API. We explained that in an earlier article (in German).
Claude reads each receipt, creates it in sevdesk, selects the account and tax rule, and matches it against the bank transactions.
The result:
Eight booked. With account, tax, and payment.
Three duplicates spotted. Invoices that had already been booked but were in the folder again. Set aside.
Seventeen deliberately left alone.
Why seventeen were left alone
All foreign-currency purchases. API credits, domains, tools from the US.
Our handbook for Claude contains a rule: never convert currencies yourself. Nothing gets booked until the actual euro amount from the credit card is available.
A model that "just quickly" converts at today's exchange rate produces entries that look plausible. And are still wrong. Nobody notices. Until the accountant asks.
The discrepancies and problems AI flags for us are also incredibly useful. We can quickly see which receipts don't add up.
The tool sets the guardrails
Now the important part: why we're not canceling sevdesk, even though Claude could theoretically do the bookkeeping itself.
sevdesk gives Claude a framework it can't step outside:
Claude can't invent accounting entries. The API only accepts receipts assigned to valid accounts. No free-form journal entries. Sometimes annoying. But that's how it should be.
What's booked stays booked. Changing it requires a deliberate step back to draft. Nothing gets casually overwritten.
The advance VAT return and final submission stay with you and your accountant. Those happen in the interface, not through Claude. The accountant can also correct anything that needs fixing in DATEV.
AI is great at reading, categorizing, and spotting anomalies. A brilliant coach. But in bookkeeping, you don't want it writing its own rules and doing whatever it thinks is right.
The tool supplies the rules. Claude just operates it.
Do you even need AI for this?
Honestly? Most people don't.
We like experimenting with AI and seeing what's possible.
sevdesk already has AI built in.
It automatically reads receipts. And when your bank account is connected, incoming payments are booked automatically.
For freelancers and new founders, that's often all you need.
Claude is the optional extra for AI power users who want edge cases handled, analyses, and reports. We showed the setup step by step here (in German).
In our case, a human still reviews everything. Claude only does the prep work. And the accountant checks it again at the end.
Getting started (German-market plans)
Free: β¬0 ongoing, up to three invoices per month, including e-invoices. A way to try it without risk.
With API access for Claude: the Buchhaltung Pro plan. Use AINAUTEN60 to save 60% on a twelve- or twenty-four-month term. Prefer monthly billing? Use AINAUTEN3M for 60% off the first three months.
P.S. Next month in part two: a prompt that makes Claude read only, touch nothing, and explain where your business stands in five minutes.
πͺπΊ Europe is back: Mistral Large 4 & Kolibri
To wrap up, a short but important update from Europe.
Back in August, we were wondering whether Europe was giving up on building its own frontier model.
Mistral had suddenly put a Chinese model on its platform. Many people assumed it was throwing in the towel.
This week came the answer. Twice!
Mistral Large 4. Nickname: Le Chonk. π One trillion parameters, trained and operated entirely in Europe.
Available through the API now. Weights for self-hosting are expected at the end of October.
Mistral calls it the best open model from the US or Europe. Artificial Analysis's numbers support that.
To be fair: with 38 points in the Artificial Analysis ranking, it's behind seven open models from China.
And twenty points behind Opus 5.5. So roughly at the level of GPT-6 Luna. Yes, the small Luna from earlier.

Still. We think having a good open model from Europe is IMPORTANT!
The second piece of news is one we can't fully assess yet:
Kolibri-1 from Aleph Alpha.
Yes, Aleph Alpha. The once-hyped German lab many had written off is back with a model.
It's the first new model from Heidelberg since August 2024. It's small (3.5 billion active parameters), supports German and English, and is freely usable under Apache 2.0.
It's built for government agencies and industry, not the top of the leaderboard. So far, the only benchmarks come from the vendor.
We wanted to cover it because we so rarely get AI news from Germany. But it isn't exactly blowing us away either.
A small irony: according to heise, some of the training data comes from Chinese models. π
Our take: reports of Europe's demise were premature
Europe isn't out. But it isn't out front either.
For most companies here, something other than leaderboard position matters anyway. Where does the model run? Can I run it myself? And is it good enough for my task?
Our test above shows that a Luna-level model can be enough if you turn up the effort. (Disclaimer: we haven't tested Mistral or Kolibri ourselves yet.)
Made it! Another longer edition, with some more complex topics than usual. Maybe we should have written it on Low. π
See you next week!
Reto & Fabian from AInauten





