I Set Up Codex by Talking to Claude — Without Outsourcing My Decisions
AI Development Practice·10 min

I Set Up Codex by Talking to Claude — Without Outsourcing My Decisions

Conversational setup beats reading docs, but it can outsource your brain. My afternoon setting up OpenAI Codex CLI with Claude, questioning every step.

Y
Young Tsai

Would you let an interior designer decide every single chair in your home, without your input? Or would you stay involved in every choice?

Most people would pick the second option Because you know this is your home, not the designer’s

So why do so many people let AI make all their technical decisions?


One Afternoon, One Story

Here’s what happened I wanted to try installing Codex CLI in the afternoon It’s an OpenAI tool, and recently a lot of people in the community have been playing with hybrid setups (Claude + Codex together)

I opened Claude Code and typed my first line “Help me install the codex plugin”

Two years ago, I probably would’ve thought AI should just do it for me and save me the trouble

But Claude didn’t move right away It asked back “Do you mean codex-plugin-cc (Anthropic plugin version), Codex CLI (standalone binary), or codex-mcp (MCP server wrapper)? These are three different things”

I froze for three seconds

Right, I hadn’t even figured out which one I actually wanted

That’s the first value of conversational setup It’s not here to help you be lazy, it forces you to think clearly


Conversational Setup Is a New Move, But Don’t Outsource Your Brain

Professor Lee-Feng Chien has a line I really like “Don’t outsource your brain to AI”

He adds one more line “Always think for 10 minutes first, then ask AI” (from that HBR Taiwan piece)

What I did that afternoon was basically applying this principle to AI tooling setup

Conversational setup is neither GUI (clicking buttons) nor CLI (reading docs) It sits in between, faster than docs and more flexible than GUI

But the biggest trap is this It can easily become “you say it, AI does it, and you don’t know what happened”

So I gave myself one rule AI is not allowed to make decisions for me, and I challenge every suggestion


5 Things I Did Right

Claude’s first suggestion was “You already maintain 208 skills on the Claude side, do you want to symlink ~/.claude/skills/ so Codex can use them too, and avoid maintaining two copies?”

Sounds reasonable, right? Convenient

I asked back “Could this break Claude? What are the cost and risk?”

Claude paused and answered

  • It won’t break Claude, symlink is one-way read
  • But ~/.claude/skills/ includes ref-skills with client CRM info / family info
  • Codex sandbox runs on OpenAI side, so that means OpenAI can see that data

I pulled my fingers off the keyboard and made coffee

This is exactly what Naval says about leverage and accountability Hybrid setup gives leverage, but accountability always stays with me

If I had been lazy for 30 seconds, OpenAI sandbox would’ve gotten keys to my safe

2) I asked for multiple angles (“How do top experts do this?”)

When deciding hook translation scope, I asked “Do we translate all 74 hooks?”

Claude gave me an “all of them” plan

I pushed again “How do top people in the industry do this? What has Boris Cherny said?”

Only then did it add this Boris said on the Latent Space podcast “Claude Code is not a product as much as it's a Unix utility” Meaning tools should be composed, not adopted wholesale

So we changed strategy Translate only 5 high-risk hooks (cloud-run secret / persona gate / client anonymization / LLM hardening / SQLAlchemy FK)

Leave the other 69 on the Claude side, because Codex shouldn’t commit / push / modify schema anyway

3) I rejected one-sided conclusions (“What does the broader community do?”)

When Claude gave me the first hybrid division-of-labor proposal, it felt a little too marketing-ish

So I asked “How does Karpathy view this?”

Then it added this Karpathy has said one of the biggest AI pitfalls is sycophancy A single model can self-convince inside its own reasoning loop

In other words, the core value of hybrid is not “using two tools is stronger” It’s “switch to another model that isn’t trapped in the same loop and can say ‘this is wrong’”

That framing completely changed my division matrix Codex isn’t “upgraded Claude” It’s another engineer who can spot Claude’s blind spots

In Karpathy’s Y Combinator 2025 “Software 3.0” keynote, he also mentioned the Centaur framework (human + AI) The point is “human + AI,” not “AI replaces human”

4) I proactively added blind spots myself

Halfway through the conversation, I caught a problem “Wait, after translating hooks over, will Codex actually run them? Or will files just sit there unregistered?”

Claude initially assumed Codex would pick them up I insisted on verification

So we listed 30 test cases and ran all of them For each case we captured exit code + stderr, then checked expected behavior

The numbers looked like this

  • cloud-run secret guard: 6/6 correct, 4 malicious cases blocked (exit 2), 2 normal cases passed
  • client anonymization scan: 6/6 correct
  • persona gate: 4/6 correct + 2 known false positives (I accepted this, fixing next week)
  • LLM endpoint hardening: 6/6 correct
  • SQLAlchemy FK: 6/6 correct

Overall 28/30 Not perfect, but known limitations were documented instead of buried

If I hadn’t pushed for verification, I’d think I had 5 guardrails In reality, persona-gate misfires on memory/feedback-blog-*.md and scripts/blog_publisher.py That’s a full 33% miss rate, and I would’ve had no idea

This is exactly Professor Chien’s “60→80 with AI / 80→100 with your own brain” Claude gave me the 60 (hook translation finished) The next 40 required me to insist on running tests to reach 80

5) I demanded proof that it can block secrets

Final step, I asked for an attack simulation

Claude wrote 6 bypass cases (yaml-dump / json-dump / case variations / pipe grep bypass) and fed them into hooks

5/6 were blocked with exit 2 The 6th was a legitimate gcloud command and should pass So hook logic was correct

If I had skipped this step, hooks might still have had holes and I’d only discover them after a real leak


3 Things I Almost Got Wrong

To be honest, I almost messed up three times that afternoon

  1. I almost symlinked the entire skill pack to Codex My first instinct was convenience, thankfully I asked “will this blow up?”
  2. I almost hard-coded routing rules before real usage Claude gave me a “Codex does X / Claude does Y” list, I insisted on running one week before locking it
  3. I almost accepted the marketing conclusion that “Codex is stronger” There are lots of “Why We Switched From Claude Code to Codex” testimonials, but after talking to three friends, most had only used Codex for one week and compared before building any real stack, which isn’t fair

Common pattern in all three Claude’s initial suggestions were reasonable, but incomplete

If I had accepted them directly, each one would leave a delayed landmine that explodes six months later


What I Learned During the Process

While writing this, I realized something

That whole afternoon, I wasn’t actually “installing Codex” I was designing a hybrid setup with Claude through active thinking

Codex was just one component in the setup The real deliverable wasn’t “installed” It was a hybrid role matrix in my head plus a 30-case validation report

That’s the value of conversational work Not AI doing things for you AI helping you think the problem through

Professor Chien has a concept called “π-shaped talent” Deep in one area + broad in multiple areas + cross-domain integration + AI leverage You can’t miss any of these

Conversational setup is one concrete way to combine cross-domain integration and AI leverage


Ending

Conversational is not lazy mode

AI gives you the speed to get to 80

The last 20 still comes from your own brain

Next time you open Claude and type “help me install X,” try pushing back on every suggestion Ask “why?”, “how do experts do it?”, “what are the risks?”

You’ll find the process changes from 30 minutes to 3 hours

But what you learn will be 10x the 30-minute version

What about you? In your most recent “doing things through AI conversation” session, how many times did you challenge it?


Extra: One Thing I Only Realized After Writing This (Real-world Epilogue)

After writing the five things above, I ran another real E2E test I registered the 5 hooks in ~/.codex/config.toml, then used codex exec "..." to send a command that should be blocked

Result The hook didn’t fire

I froze for 30 seconds Because the previous 30 fixture tests had passed, so hook script logic was definitely correct

I checked the official OpenAI Hooks docs A small line at the end says:

「PreToolUse doesn't intercept all shell calls yet, only the simple ones. The newer unified_exec mechanism allows richer streaming stdin/stdout handling of shell, but interception is incomplete」

Translation Codex runs shell with the new unified_exec, and PreToolUse interception is incomplete

In other words I translated 5 hooks to Codex, but when Codex runs in exec mode, those hooks are basically bypassed

This discovery directly breaks my earlier hybrid guardrail plan


The Real Implication

I originally thought the architecture looked like this

Codex runs shellPreToolUse hook checkblock or pass

In reality, it’s this

Codex runs shell (unified_exec path)⚠️ executes directly, hook doesn't fireCodex runs shell (simple Bash path)✅ goes through hook

The problem is that non-interactive codex exec mostly uses unified_exec path That means my guardrails were effectively wasted

So I did two things

  1. Removed the 208-skill symlink from Codex Guardrails are unreliable, so client data cannot be exposed for Codex to read
  2. Repositioned Codex’s role in my stack From “guardrailed worker” downgraded to “unguarded review-only tool”

In practice, Codex can now only be used for

  • /codex:review to inspect PR diff (read-only, no file writes)
  • /codex:adversarial-review to pressure-test design (same)

Disabled: codex exec for writing code, codex /goal for long-running tasks No guardrails means real risk


Biggest Lesson I Learned

Passing Mode A tests ≠ production actually works

Mode A is “feed fixtures directly into hook script” That validates script logic Production E2E is “when Codex really runs commands, does hook fire?” That validates the full chain

I got 28/30 pass in Mode A, then assumed “OK, guardrails are ready”

If I hadn’t insisted on running Mode B, I would keep using hybrid with fake confidence and only realize it after a real leak

That echoes Professor Chien’s line “60→80 with AI / 80→100 with your own brain” Mode A is 60→80 Mode B is 80→100

And it echoes the main point of this post Conversational is not lazy mode

If I had accepted Claude’s “setup complete” conclusion right after Mode A pass, the whole afternoon would have been wasted


Research I Did for This Post (Usually I’d write by gut feel, this time I wanted it grounded)

Before writing this, I spawned 4 research agents in parallel to check:

  • Expert perspectives (Boris Cherny / Karpathy / official recommended patterns)
  • Industry FAQ + real pitfalls (reddit / dev.to / GitHub issues)
  • Risk / cost / learning curve (Cymulate security research + GitHub issues)
  • Multi-source references (YouTube / podcasts / official docs / GitHub demo repos)

In total I reviewed roughly 30+ articles / 5 GitHub demo repos (codex-plugin-cc 18.1k stars / compound-engineering-plugin 16.4k / claude-codex-settings 674 / Z-M-Huang/claude-codex / everything-claude-code) / 3 podcasts / 5 YouTube videos

Arguments supporting hybrid

  • Karpathy: cross-model review catches sycophancy (same model can self-convince in its own reasoning loop)
  • Boris Cherny: “Claude Code is not a product as much as it's a Unix utility” — Codex is another composable utility
  • Naval: more leverage layers are usually stronger than fewer (if you can maintain them)
  • 18.1k stars on codex-plugin-cc, so community momentum is real
  • Anthropic officially promoting plugins is basically endorsement of the cross-model concept

Arguments against / skeptical of hybrid

  • Many people on Reddit: “I tried for a week then gave up, too complex and not worth the learning cost”
  • Cymulate security research: Codex sandbox is still immature, hybrid increases attack surface
  • In that Every podcast “Why We Switched” testimonial, interviewees had only used Codex for one week and compared without maintaining a stack, which is unfair
  • Anthropic promoting plugins comes with bias, they won’t officially say “Codex is weak, don’t use it”
  • Two subscriptions at $400/month is a high threshold for solo developers

When hybrid is a good fit

  • Heavy coding usage (4+ hours/day of actual AI-assisted development)
  • Already paying Anthropic Max and actually consuming the full monthly quota (no marginal cost)
  • You want to learn cross-model review workflows
  • You already have your own hook / pipeline / skill stack (like me)
  • You run multiple projects and want parallelism

When you should NOT do hybrid (my own counter-case)

  • Light AI coding usage (under 5 hours/week)
  • You want one subscription and lower cost
  • You’ve never built your own hooks / pipeline / skills
  • You want “simple and just works” (first-week hybrid productivity drop of ~30% is real)
  • Only 1-2 small repos, overhead is not worth it
  • You’re in a high-pressure deadline period (learning while shipping is unrealistic)

Blind spots I may still be missing (honest version)

  • I may be overestimating the value of cross-model review, I only ran a few small experiments and no controlled study
  • I may be underestimating plugin update breakage, the 30 tests were run once, behavior can shift after updates
  • I wrote this after already deciding to do hybrid, so strictly speaking this is rationalization, not deliberation
  • In 6 months, Codex might ship Plan Mode + commit hooks, then my entire role matrix needs redesign
  • Maybe Cursor / Cline / Aider is actually the more valuable third path, and I haven’t evaluated that seriously

Feel free to comment with your blind spots I want to know mine too


Further Reading

Viewpoints supporting “conversational + don’t outsource your brain”

Counterpoints / skeptical takes (read these before deciding)

Community synthesis and integration

claude-codecodexai-toolingconversational-setuphybrid-aiharness-engineering