After Reading Karpathy's llm-wiki (1): AI Memory Doesn't Expire on Its Own — It Quietly Leads You to Wrong Decisions
AI Development Practice·10 min

After Reading Karpathy's llm-wiki (1): AI Memory Doesn't Expire on Its Own — It Quietly Leads You to Wrong Decisions

Karpathy's llm-wiki sparked a wave of knowledge base implementations. Lex Fridman uses ephemeral wikis — build, use, delete. I use the same concept to manage enterprise projects for a year and a half, but my version added one thing: verification. Because stale knowledge doesn't disappear — it makes you confidently wrong.

Y
Young Tsai

First, what Karpathy actually said

Andrej Karpathy — OpenAI co-founder, former Tesla AI lead, the person who coined "vibe coding" — published an idea file called llm-wiki in early April

He noticed his AI usage shifting: less time generating code, more time organizing knowledge

So he proposed an architecture:

  • raw/: dump raw material in (papers, articles, web clippings), don't modify
  • wiki/: AI reads raw, compiles into structured knowledge pages. One new source can touch 10-15 wiki pages
  • schema/: config files that tell AI how to operate (like CLAUDE.md)

Three operations: Ingest (compile new material into wiki), Query (search + feed back new knowledge), Lint (quality check — catch contradictions and stale info)

Core concept in one line: compile knowledge once, keep it current, don't re-derive it every query

Karpathy only published the concept, no implementation — deliberately letting the community build it. Within two days, a dozen implementations appeared on GitHub, 84 tools compiled into an awesome list

He later followed up with a bigger point: in the age of agents, you don't share apps — you share idea files. You write down the idea, the other person's agent customizes and builds it for their specific needs. No need to share code, because everyone's situation is different and the agent will grow the right version for them

That framework inspired a lot of people, myself included. My context is freelance consulting — managing over a dozen projects simultaneously — so I extended his concept in a direction shaped by that reality. Rules grew out of every mistake, SOPs out of every time a client asked something I wasn't ready for. A year and a half of that, and it became a system. Same seed, different soil, different shape

Same concept, different people took it in completely different directions. The one that surprised me most was Lex Fridman — he took Karpathy's architecture and did the exact opposite of what I do


How Lex uses Karpathy's architecture: build it, delete it

According to VentureBeat's coverage, Lex uses a similar architecture in an interesting way

Before recording a podcast, he has AI crawl the guest's papers, tweets, controversies, technical background — compiled into structured markdown

Then he goes running, headphones on, talking to this knowledge base in voice mode. Comes back with a feel for the topic

That knowledge base? Deleted

He calls it an ephemeral wiki — born for one task, gone when the task is done

Smart — no maintenance cost, clean slate every time


My approach: keep it

I use Claude Code to manage projects for over 5 companies. every day for a year and a half

My knowledge isn't built for a single session. It grew from bugs, client meetings, and postmortems after things went wrong

A real scenario:

Three months ago, I deployed a healthcare project and said "it's live" — without actually curling the URL to check. Got a 404. After that I added a rule: must show proof before claiming completion

Three months later, deploying an education project. AI tried to say "deployment complete" and move on. But that rule was still there — it got blocked, had to run verification first

Session recovery can't solve this — /catchup picks up where you left off, but it won't remember a lesson from a different project three months ago

Cross-project, cross-time knowledge is what needs a memory system


The real difference isn't save or don't — it's verify or don't

Karpathy's original text mentions three operations: Ingest, Query, Lint (quality check)

Most people do the first two and stop. Build a wiki, query it, done

I added one more: verification

I built a quality checker that scans all knowledge files:

🔴 Critical

  • TODO items with passed deadlines, still unchecked
  • Index summaries contradicting actual file content

🟡 Warning

  • Files that exist but aren't in the index
  • Files unchanged for 30+ days but marked "in progress"

First run: 4 critical, 12 warnings, 10 formatting suggestions

One critical: a task marked "in progress" that was long finished — outcome never written back

Without this check, that stale info stays there. Next session, AI gives advice based on a wrong assumption — and it sounds confident doing it


Hooks: don't rely on AI remembering — enforce with code

Beyond quality checks, I have a layer called hooks

Hooks aren't suggestions. They're gates — code-level enforcement

Example: AI says "done" but provides no evidence (screenshot, actual response, test output) — hook blocks it

Data from 92 sessions: this hook fired 516 times. Over half of all sessions had AI trying to claim completion without proof

Another: if it's been 7+ days since the last knowledge cleanup, a reminder fires at session start. Not AI remembering — a script reading the log and counting days

Lex doesn't need any of this. His wiki won't survive a day. If it breaks, who cares

Mine needs to survive months. Without verification, you're not accumulating knowledge — you're accumulating confidence in outdated information


But internal checks aren't enough

Quality checks and hooks are my own metrics. If that's all there was, I'd be no different from anyone sharing "my AI workflow" online

So I hold myself to one more standard: whether the system works isn't for me to decide — it's for the people paying me

This isn't a standard I'm imposing on anyone. Learning AI for fun is a perfectly good reason

But I eat with this system

Currently: 15+ projects working with me simultaneously, spanning healthcare, education, elder care, career services. All handled after hours — evenings, weekends, while I sleep the system keeps running. Once the system is built right, most tasks are near-automated. No Notion, no Trello — Claude Code + GitHub handles everything

Clients don't know what tools I use. They know: deliverables arrive, quality is consistent, responses are fast

That's the real quality check — the market runs one for you every month. No renewal means you failed


Back to Karpathy

Reading that idea file, my reaction wasn't "I should build this"

It was: "this is what I've been doing"

He started from concepts. I started from problems. His idea file is clean. My system is patches everywhere. But every operation he described, I was already running

Something I realized: a system that truly fits your needs is grown, not copied. But while you're growing it, you can distill what others have figured out and graft it onto yours. Karpathy gave me the language. The system was already mine

His line means:

Knowledge compiled once, kept current, not re-derived on every query

The key word isn't "compiled"

The key word is "kept current"

And keeping current requires verification. Otherwise you compiled once, then slowly rotted


Next: What my system actually looks like — six-layer architecture, four tools, the screw-up behind every rule, and what I'd do if starting over

claude-codekarpathyllm-wikiknowledge-managementlex-fridmanmemory-system