AI-HOY, AInauts!
You wanted to hand work over to AI. Ten minutes later, you are hunting for files, pasting context, and explaining everything again. Congratulations: your AI has successfully delegated the job back to you and made you the bottleneck again π.
Today, we will show you how to turn a chatbot into a real AI employee with its own workspace, clear assignments, and the right model for each job.
Here is what we have for you today:
π¨βπ» How to turn your chatbot into a real AI employee today
π The Gauntlet Loop: Give your AI a clear assignment
π Luna is 80% cheaper: Which jobs should you hand off?
π€£ AI Fun: The funniest AI crashes in movie history
Letβs get started.
Your prompts are leaving out 80% of what you're thinking.
When you type a prompt, you summarize. When you speak one, you explain. Wispr Flow captures your full reasoning β constraints, edge cases, examples, tone β and turns it into clean, structured text you paste into ChatGPT, Claude, or any AI tool. The difference shows up immediately. More context in, fewer follow-ups out.
89% of messages sent with zero edits. Used by teams at OpenAI, Vercel, and Clay. Try Wispr Flow free β works on Mac, Windows, and iPhone.
π¨βπ» You Are the Bottleneck: Turn Your Chatbot Into a Real AI Employee Today
You open a chat, explain the task, upload files, answer follow-up questions, and correct the first draft. Then the next task starts from scratch.
That is not delegation. It is supervision in a chat window.

For blog posts and articles, we often needed an average of 13 rounds in Chat and Cowork before the result was usable. The AI was capable, but it had no persistent workplace and no reliable operating system.
Give Your AI a Computer
The key change is simple: give your AI a fixed workspace on its own computer.
Instead of starting with a blank chat, the agent works with durable files, clear work orders, reusable skills, and visible evidence at the end. A customer proposal, for example, should no longer begin with you pasting the client background, pricing, and previous offers into a prompt. The workspace already contains the approved context and examples. Your assignment only needs to describe the outcome.

The Tool Is Not Your System
The desktop apps for ChatGPT Work and Codex or Claude Cowork and Code are the door to the office. But you still have to furnish that office.

A useful AI workspace needs six things:
A short context file with the role, goals, and workflow
Approved examples and relevant knowledge
Access to the tools required for the job
Clear boundaries for actions with external impact
Reusable skills for recurring tasks
Visible evidence that the task is actually complete
The app provides the interface. Your workspace provides the operating system.
More than one billion people use OpenAI models. Yet only around ten million use Work and Codex. If you build a real workspace now, you are already among the roughly one percent who have moved beyond ordinary chat.

Our Take: Go From Chatter to Client
The goal is not to write longer prompts. It is to become the client: you define the direction, the benchmark, and the final approval. The AI does the work in between.
Start with one recurring task. Build the smallest workspace that gives the agent the context, examples, tools, and boundaries it needs. Then require a visible receipt at the end. Once that works reliably, expand from there.
Stop making AI decisions in the dark.
Leadership is asking: where is AI delivering value for us and where is it creating risk? Right now, most teams have no idea.
With Harmonic Securityβs Usage Explorer, you get a complete picture of how your organization actually uses AI, automatically categorized into custom use cases with complete tool-level granularity.
π The Gauntlet Loop: Give Your AI a Clear Assignment
Even the best workspace can still operate like it is 2023: you request snippets, review every intermediate step, and keep rescuing the process.
An agent needs more than a goal. It also needs a benchmark that tells it what βgood enoughβ means.
That is the idea behind Matt Shumerβs Gauntlet Loop. You give the AI one demanding assignment, let it work through subproblems, test the result against a concrete quality bar, and keep improving until it passes.
Shumer used this approach to trigger multi-hour subagent runs. One experiment produced a browser shooter with around 55,000 lines of code. Others created games and complete applications. The impressive part is not the line count. It is the process: the agent receives a clear target and is responsible for closing the gap.
The Principle in Five Steps
Define the outcome. Describe the finished result, not a list of vague activities.
Set a benchmark. Give the agent an example, test, score, checklist, or other concrete quality bar.
Let it build. The agent researches, creates, and uses subagents or tools where useful.
Test the actual result. The agent compares its output against the benchmark and gathers evidence.
Repeat until it passes. Weak points become the next work order instead of a reason to stop.
The benchmark must be concrete. βMake it excellentβ is not a benchmark. βMatch these three approved examples, pass these tests, and show the receiptsβ is.
The Prompt to Steal
Use this short meta-prompt to turn a rough idea into a bounded Gauntlet assignment:
First read Matt Shumerβs explanation of the Gauntlet Loop:
https://somethingbig.ai/gauntlet-loop
My goal:
[Describe the desired finished result.]
Check whether this goal is suitable for a Gauntlet Loop.
Ask only the questions that materially change the assignment, benchmark, or safety boundaries.
Then turn my answers into the shortest possible assignment that includes:
- the finished outcome
- the available files, tools, and sources
- a concrete quality benchmark
- required tests and evidence
- a loop that continues until the benchmark is met
Actions such as sending, publishing, paying, deleting, or changing permissions remain blocked until I explicitly approve them.
Show me the final assignment for confirmation before execution.You can also study Shumerβs longer Claude of Duty prompt and follow @mattshumer_ for more experiments.
π Luna Is 80% Cheaper: Which Jobs Should You Hand Off?
OpenAI has lowered the API price of Luna by 80 percent and Terra by 20 percent. For Work and Codex users, that mainly means the same usage limits can now cover more work.
But cheaper does not mean interchangeable.
Luna Max Is Still Luna
More reasoning gives a model more time and tokens to think. It does not remove the modelβs fundamental limits. Luna with maximum reasoning does not become Sol, just as giving Einstein more paper would not turn someone else into Einstein.
The useful question is not βWhich model is best?β It is βWhat is the cheapest model that can complete this specific job reliably?β
How to Enable Maximum Reasoning
In the ChatGPT desktop app, open Settings β Configuration β Available reasoning levels and enable the additional levels.
Language note: The screenshot below shows the German ChatGPT desktop interface. The setting is under Settings β Configuration β Available reasoning levels.

Luna and Terra are not regular web-chat models. You use them inside Work and Codex, where they can operate on files, tools, and longer assignments.
A practical division of labor looks like this:
Sol High: planning, unclear assignments, important decisions, and final review
Luna High or Max: repetitive, bounded tasks with a clear benchmark
Terra High or Max: more complex execution when Luna is not reliable enough
Start with High. Use Max only when the extra reasoning produces fewer errors or fewer review rounds.
This prompt helps decide which model should own a task:
Review this task and assign it to Sol, Terra, or Luna.
Task:
[Insert the task.]
Evaluate:
1. How clear is the desired outcome?
2. Is there a concrete benchmark or test?
3. How costly would a wrong result be?
4. Does the task require planning, judgment, or mainly execution?
5. Can the result be checked quickly and objectively?
Choose the least expensive model likely to complete the task reliably.
Recommend High or Max reasoning.
State what Sol must still plan or review.
Give one short acceptance test.Our Take: Sol Has to Earn Its Jobs
Test a recurring task twice. If the smaller model meets the benchmark, it keeps the job. Use Sol when the benchmark is unclear, Terra has already failed, or the decision is important enough to justify the extra capability.
In short: Sol decides and reviews. Luna and Terra do the bulk of the work with higher reasoning when it actually helps.
For more detail, see OpenAIβs notes on the price-performance frontier, the model guide, this practical breakdown, and the open-source Sol Advisor.
π€£ AI Fun: The Funniest AI Crashes in Movie History
What happens when AI systems try to recreate famous movie scenes and lose track of physics, faces, or the plot? The results are gloriously chaotic.
Here are two clips that made us laugh this week:
Thatβs it for today. See you Thursday!
Reto & Fabian from AInauten






