"You sleep, AI does research" — Why did this one sentence stop the entire AI world?
March 2026. Karpathy dropped a repo: autoresearch.
Three files. No fancy UI, no complex infrastructure. It went viral on X instantly.
Three reasons why.
Layer 1: The Karpathy effect. He's an early OpenAI member, former Tesla AI lead. When he says "this is interesting," the entire community stops to look.
Layer 2: The narrative is irresistible. "You sleep, AI does research" — anyone can understand it, anyone would want it.
Layer 3, the most important: It previews a new way of working.
Previous AI tools were "help you write code" — Copilot, ChatGPT, Claude. You give instructions, it produces, you review. Back-and-forth interaction.
autoresearch is different. You design the rules, then leave, and let AI run by itself. The interaction becomes: design once, execute infinitely.
But wait. What's the actual principle? Why can something this simple pull this off?
Three files, one loop
The entire autoresearch system has three core files:
prepare.py — Data preparation, tokenizer, evaluation. Locked down. Agent can't touch it.
train.py — Model architecture and training logic. The one file the agent is allowed to modify.
program.md — The research SOP written for the agent. How to run, how to evaluate, how to rollback.
The workflow is a single loop:
No RAG, no vector database, no multi-agent orchestration. One loop, one number, one git command.
That's it.
It captures the most critical trait of AI
"That's it" works because Karpathy nailed the fundamental strengths and weaknesses of agents:
Agents don't mind repetition. They don't get tired. They don't feel like giving up after the 30th failure.
But agents struggle with: problems that are too big, no clear direction, no way to know if they're doing well, no way to undo mistakes.
Every design decision in autoresearch eliminates an agent weakness:
| Agent weakness | How autoresearch solves it |
|---|---|
| Problem too big, don't know where to start | Can only modify one file |
| Don't know if it's working | One number: val_bpb |
| Afraid of breaking the whole system | git reset auto-rollback |
| Need a human to judge results | Compare numbers, program decides |
| Stop mid-work to ask for permission | NEVER STOP — no asking, no pausing |
Each solution is extremely simple. Combined, they create a perfect operating environment for agents.
This isn't making AI smarter. It's making the problem dumber. Dumb enough that AI can handle it every time.
The most counterintuitive design: driven by .md, not scripts
This is the part I find most worth discussing.
The core control logic isn't in a Python script. It's in program.md — a Markdown file.
Wait. Markdown is just documentation, right? How does a document drive an agent?
It does. And in certain scenarios, it's more powerful than a script.
Here's why: Scripts control machines. Markdown controls AI judgment.
# What a script can do:
if profit_factor > best:
keep()
else:
discard()
Scripts handle conditional logic. But they can't express:
- "If the improvement is tiny but the code got much uglier, it's not worth keeping"
- "If it runs longer than 10 minutes, kill it — don't waste time"
- "When you run out of ideas, re-read the config and try combining near-miss experiments"
- "Don't chase noise — a profit_factor difference of < 0.05 with under 100 trades is probably random"
All of these live in program.md. In natural language.
Because LLMs understand natural language. You describe judgment criteria, strategic direction, what to do and what not to do — and the agent follows.
Scripts are hard logic: if/else, loops, thresholds. Markdown is soft logic: judgment, priorities, trade-offs, strategic direction.
They don't replace each other. They each manage a different layer:
| Layer | Tool | What it controls |
|---|---|---|
| Execution | Script | Run training, read numbers, git operations |
| Decision | Markdown | Should we keep this? What to try next? What counts as "good enough"? |
When to use which? My rule of thumb is simple: Deterministic logic, no judgment needed → script. Trade-offs, AI needs to decide next steps → markdown. The strongest setup uses both. autoresearch is the textbook example.
Karpathy's real insight: In the agent era, the decision layer matters more than the execution layer. And the decision layer is better written in natural language than in code.
Real-world application: From stuck in Q&A loops to 42 autonomous iterations
While learning quantitative trading, I hit a wall.
Not because I couldn't write code. Because of how I was interacting with AI.
Every time I wanted to adjust strategy parameters, switch training data periods, or try a different backtest configuration, I had to go back and forth with Claude: "Now change the train period to this." "Switch this parameter." "Results are bad, roll it back."
The problem was, many of these operations were things I couldn't articulate clearly myself. Something like "try a different data window and see what happens" involves knowing which settings to change, how to slice time periods, how to compare results — I had a fuzzy direction in my head, but I couldn't describe every step precisely.
Claude wasn't incapable, but it needed me to spell out each step. When I couldn't, it would stop and ask. I'd explain for a while, it would produce something that wasn't quite right. Then another round.
This wasn't an AI capability problem. It was an interaction model problem.
The old model was instructional: I had to tell it what to do step by step. Everything bottlenecked on me — thinking of the next experiment, describing the operation, judging the result, deciding whether to rollback. The entire "think → communicate → execute → evaluate" cycle was on my shoulders.
After studying autoresearch, I tried something: applying the exact same design pattern.
I wrote an autobacktest-program.md that defined:
- The one file that could be modified (strategy parameters)
- The evaluation metric (profit_factor)
- Rollback rules (no improvement → git reset)
- One crucial instruction: NEVER STOP
Started it, went to do other work.
When I came back, it had completed 42 experiments on its own. Changed parameters, ran backtests, judged results, rolled back failures — all autonomously.
I no longer needed to articulate each step. Because program.md doesn't define "what to do" — it defines "search space + evaluation criteria + rollback mechanism." The agent explores within that framework, no step-by-step instruction needed.
During the process, it made several judgments I didn't expect:
- When slippage was set to zero, theoretical returns looked great, but the agent marked it "unrealistic" and discarded it
- A strategy with only 3 trades showed high profit_factor, but it chose a version with more trades instead — treating low sample sizes as untrustworthy
- A trailing exit strategy lost heavily; after discarding it, the agent never went back to retry
These judgments weren't if/else statements in a script. They were natural language rules in a Markdown file. The agent understood them and followed them.
Important caveat: This was a learning exercise in backtesting, not real trading. Backtest results and live performance are separated by many variables — slippage, liquidity, market regime changes — all of which can dramatically reduce backtested numbers in practice. The point of this case study isn't the profit numbers at all. It's this: a learning workflow that was previously stuck in Q&A loops became a self-running closed loop using the autoresearch design pattern.
From "I instruct AI step by step" to "I design rules, AI runs on its own" — that shift is the real takeaway.
Full automation isn't about tools — it's about design
OpenClaw (the "lobster" agent) recently went viral, promising AI that can operate your entire computer — send emails, book restaurants, manage calendars. Unlike developer tools like Claude Code, it targets general users and controls the whole machine, not just code.
But here's what I've noticed: many people haven't even properly used the tools they already have — Claude Code, Codex, Gemini — and they're already rushing to install the lobster.
That's like hailing a taxi before you've decided where you're going. No matter how good the car is, you'll just drive in circles.
The bottleneck isn't the tool. It's whether you can define: what AI should do, how to judge good vs. bad, and how to rollback when things go wrong. That ability doesn't magically appear when you switch to a fancier tool. The entire design of autoresearch — one file, one metric, one rollback rule — is practice for exactly this skill.
Learn to design closed loops first. You can swap tools anytime. Without that design ability, a more powerful tool just gives you a more powerful way to be lost.
Where's the ceiling?
autoresearch is powerful, but its ceiling is clear:
It can do local optimization, not breakthrough innovation.
It can find the best parameter combination within the search space you define. But it won't question whether the search space itself is correct.
An analogy: autoresearch is an extremely diligent research assistant that will search every corner of the map you drew. But it won't say "wait — we should be looking at a different map."
The ceiling = the upper bound of the search space you define.
So the truly valuable work isn't running the loop. It's:
- Defining the right problem
- Choosing the right metric
- Designing the right search space
- Knowing when to switch maps
These are still human jobs. For now.
Conclusion: Not a stronger model, but a better loop
If you ask me for the single most important takeaway from autoresearch:
An agent's power isn't in how human-like it is, but in whether you can shape work into a closed loop where it continuously receives feedback.
The principle is simple: one file, one metric, one rollback command, one Markdown SOP.
But this simple structure captures AI's most essential trait: it doesn't mind repetition, it doesn't get tired, it doesn't need motivation — it only needs direction.
Give it a small enough space, fast enough feedback, and clear enough rules, and it will finish what you couldn't while you were awake.
This doesn't only apply to AI research or quant backtesting. Marketing, content, testing, operations — any work that can be framed as "change parameters → check metrics → keep/discard" can use this pattern.
The future won't be defined by stronger models alone. It will be defined by better feedback loop design.
Whoever can shape a task into a stable closed loop for agents will be first to capture the efficiency dividend.
