aicoding agentscodexclaude codesoftware engineeringengineering workflow

Play the Card, Don’t Collect the Deck: Why Codex Fits My Engineering Workflow Better Than Claude Code

By Anthony Kung
Picture of the author
Published on
Domain
AI literacy
Focus
Coding-agent scope, judgment, and engineering ownership
Perspective
First-hand workflow comparison

Play the Card, Don’t Collect the Deck: Why Codex Fits My Engineering Workflow Better Than Claude Code

Claude Code is one of the most popular coding agents available today. A lot of developers swear by it. When I tell people that I have generally had a much better experience with OpenAI Codex, the response is often predictable:

“Then you configured Claude wrong.”

I do not think that is a particularly useful explanation.

I have configured Claude Code. I have written project instructions. I have spent time correcting its behavior. I have tried to teach it how I want repositories handled.

And I still repeatedly encountered the same problems: unnecessary web searches for questions answerable from local files, excessive investigation before simple changes, enormous context growth, prompts to manually prune context only after it had already become unwieldy, new tests for seemingly every small modification, and narrow bug fixes turning into changes across a surprising number of files.

Meanwhile, Codex has generally behaved much closer to what I expected a coding agent to do in the first place.

That does not mean Codex is universally smarter than Claude. It does not mean Claude Code is a bad product, nor does it mean everyone who prefers Claude is wrong. But it does mean my experience cannot reasonably be reduced to a configuration mistake.

The more interesting explanation is that Claude Code and Codex have different behavioral tendencies, and those tendencies matter enormously depending on how you use coding agents.

The simplest analogy I have found is a card game.

Codex often feels like an experienced player who looks at the cards already in hand, studies the state of the game, considers which information matters, and then plays a card.

Claude often feels like an extremely energetic player who keeps asking whether another card might help.

What about this card? What about that one? Should we draw another? Maybe another combination exists. What if there is another deck we should check first?

Sometimes that behavior is exactly what you want. If the problem is poorly understood, the solution space is genuinely open, or you specifically want somebody to challenge every assumption you have made, gathering more information can expose something important that everyone else missed.

But engineering is not a competition to see who can collect the largest hand.

At some point, good engineering means recognizing that the evidence already available is sufficient, that the correct move is understood, and that it is time to play the card.

That distinction explains much of why Codex fits my engineering workflow better than Claude Code.

I do not want the coding agent to become the engineer

A lot of the disagreement around coding agents starts with a deeper disagreement about what we expect the agent to be.

There is a style of agentic development that begins with something like:

I want a SaaS application that manages organizations. Build it.

The agent is then allowed to choose a database, authentication system, ORM, API architecture, cache, job system, validation library, state-management strategy, testing approach, deployment model, and whatever other dependencies it thinks are appropriate.

A sufficiently capable coding agent can absolutely produce something impressive from that level of instruction.

That is not how I want to build serious commercial software.

If I am responsible for a product, I usually want to know what database it uses and why. I want the framework to be an intentional choice. I want authentication boundaries to be deliberate. I want to know which dependencies are part of the platform. I want the data model, deployment assumptions, interfaces, and subsystem responsibilities to match a technical vision rather than whichever implementation happened to be most convenient for the agent.

My interaction with a coding agent is therefore much closer to:

This is the architecture. These are the constraints. These technologies are intentional. This subsystem owns this responsibility. This is the behavior we need. Understand the existing system and implement the requirement within those boundaries.

The agent can still perform an enormous amount of autonomous work. It may write most of the implementation. It can search hundreds of files, trace an unfamiliar call path, diagnose failures, run tests, iterate on a difficult bug, and work for hours.

I am not arguing against autonomy.

I am arguing for directed autonomy.

The human owns the engineering direction. The agent dramatically accelerates execution.

OpenAI has described its own agent-first software-development experiment with almost exactly that division of responsibility: “Humans steer. Agents execute.” Its engineers describe their role less as manually writing code and more as defining intent, creating acceptance criteria, shaping the environment and repository, and constructing feedback loops that allow Codex to perform implementation reliably.

What is particularly interesting is that Anthropic's own research points in essentially the same direction. In a privacy-preserving analysis of roughly 400,000 Claude Code sessions, Anthropic found that users made about 70% of planning decisions while Claude made roughly 80% of execution decisions. Anthropic also found that greater task-specific domain expertise was associated with higher success and a better ability to recover when the agent misunderstood something.

So this is not really an OpenAI philosophy versus an Anthropic philosophy.

Both companies have evidence pointing toward an important role for human direction.

The difference for me is how much effort I have to spend keeping the agent inside that direction.

Rendering diagram...

That is the relationship I want with an agent. I do not want to manually implement everything, but I also do not want the technical identity of the system to emerge accidentally from whatever the agent decided to do.

If I cannot recognize when the agent is wrong, I cannot really supervise it

This matters even more because a substantial part of my work sits closer to low-level software and hardware than to mainstream application development.

An agent can generate SystemVerilog that compiles and still be wrong.

A simulation can pass and still fail to represent the intended microarchitecture.

A proposed change can look completely reasonable while altering handshake semantics, pipeline timing, reset behavior, cycle ordering, inferred structures, synthesis characteristics, or some experimental assumption that the immediate code does not make obvious.

The agent may then form a plausible hypothesis, change something, discover the failure remains, form another plausible hypothesis, modify another subsystem, create a new problem, and continue investigating.

From inside the agent loop, every individual step may look reasonable.

A human who understands the design can sometimes stop the entire loop with one sentence:

That cannot be the cause because this interface is synchronous.

Or:

Stop changing that block. The failure occurs before that boundary.

Or:

Revert this. You changed the microarchitecture, which means we are no longer comparing the same experiment.

That is not micromanagement. That is engineering judgment.

The fact that Anthropic sees a measurable relationship between domain expertise and successful Claude Code sessions is particularly relevant here. Its research concludes that coding agents are not simply substituting for domain expertise; people who understand the problem are better able to direct the agent, recognize misunderstandings, and recover when things go wrong.

This is also why I am uncomfortable with the idea that coding agents should eventually make it unnecessary for anyone on a commercial project to understand the code they are shipping.

I am perfectly comfortable using code I did not personally type. Engineering teams have always maintained systems containing code written by other people.

The requirement is not authorship.

The requirement is ownership.

Someone needs to understand the architecture well enough to review it, debug it, secure it, modify it, and decide whether what the agent produced actually satisfies the requirement.

That becomes difficult when the agent is constantly broadening the work beyond the engineering problem I intended to solve.

Claude too often makes the investigation itself part of the problem

One of the most obvious examples for me has been web search.

If I ask why a configuration behaves a particular way in my repository, the repository is usually the authoritative source. I want the agent to inspect the configuration, find the code that consumes it, follow the relevant types and call path, and determine what the system actually does.

If the answer depends on a current external API, a recently changed library, or documentation that is not available locally, then searching the web makes complete sense.

What repeatedly frustrated me with Claude was seeing external research happen before that threshold had been reached.

A local question would produce a search. The search would introduce another possibility. Another page would be fetched. That page would suggest another path of investigation. Before long, a question whose answer may have existed in three local files had become a much larger research exercise.

To be precise, this is not because Claude has web tools and Codex does not.

Claude Code explicitly exposes web-fetching capabilities and gives users permission controls that can allow, deny, or require confirmation for tool use.

Codex has web capabilities too. OpenAI's current Codex environment includes cached web search under its default sandbox model and allows broader network access under configurable permission controls.

So my complaint is not:

Claude searches the web while Codex stays local.

That would be inaccurate.

My complaint is about tool-selection judgment.

In my repeated use, Codex has been more likely to decide that the repository already contains enough information to answer a repository-specific question. Claude has more often behaved as though another potentially useful source is worth investigating simply because that possibility exists.

That extra research costs time and quota, but it also has a secondary cost: everything the agent gathers becomes more material that may compete for attention later.

Claude keeps drawing cards.

Codex more often decides that the existing hand is sufficient.

Sometimes I am asking a question, not commissioning an investigation

The same behavior appears even when I am not asking the agent to change anything.

Sometimes I just want to know what a configuration option does, where a value is inherited from, why an existing function behaves a particular way, or which component owns a piece of state.

Those are not always difficult questions.

Yet Claude has sometimes treated them as though the safest response is to explore every potentially relevant direction before answering.

This is one area where the evidence goes beyond my own frustration.

Anthropic's current prompting documentation contains an explicit section titled “Overthinking and excessive thoroughness.” It says that recent Opus models can perform more upfront exploration, gather extensive context, and pursue multiple research threads without being prompted, especially at higher effort settings. Anthropic recommends narrowing tool instructions and lowering effort when that behavior becomes counterproductive.

Anthropic also documents the token consequences directly. Extended thinking is enabled as part of normal Claude Code operation, and depending on model and settings, thinking can consume tens of thousands of output tokens on a request. Its cost guidance explicitly recommends lowering effort for simpler work where that reasoning depth is unnecessary.

Those controls are useful.

But they also undermine the simplistic response that excessive investigation must be something I invented by configuring Claude badly.

Anthropic itself acknowledges the behavioral tendency.

The fact that a behavior can be mitigated through configuration does not mean experiencing that behavior is a configuration failure.

If my normal workflow repeatedly requires me to tune down an agent's desire to investigate things another agent already handles at an appropriate level of effort, that difference belongs in the product comparison.

A small bug should not casually turn into a twenty-file project

This becomes much more consequential when the agent starts modifying code.

Suppose I have a narrow bug. The architecture already exists. The relevant subsystem is known. After understanding the surrounding implementation, the actual problem turns out to be a small conditional error.

What I generally want is simple: understand enough context to be confident in the diagnosis, make the smallest correct change, run the relevant validation, and stop.

Claude has repeatedly pushed beyond that boundary in my projects.

It finds the problem, but notices another nearby issue. That issue suggests an abstraction. The abstraction suggests a helper. A neighboring path looks similar, so perhaps it deserves inspection too. Another edge case appears worth testing. Before long, a fix I expected to touch a couple of files has expanded across a surprising portion of the repository.

Again, I am not arguing that a large change is inherently bad.

Sometimes correctness genuinely requires changing twenty files.

What matters is why the scope expanded.

I want scope to grow because the engineering requirement demands it, not because the agent discovered twenty other things it is capable of improving.

This point also has unusually direct support from Anthropic's own guidance. In the same current prompting documentation that discusses overthinking, Anthropic describes an “overeagerness” tendency in recent Claude models: creating extra files, adding unnecessary abstractions, or building flexibility that was never requested. Its suggested mitigation explicitly tells Claude that a bug fix does not require surrounding cleanup and that one-off operations do not automatically deserve new abstractions.

That is remarkably close to the behavior I have spent time trying to suppress.

And this is where Codex has simply felt more mature in my workflow.

Not because it never makes unnecessary changes. It absolutely can.

But it more often behaves as though the requested scope is itself part of the requirement.

When Codex gets something wrong, I usually feel like I am correcting the solution.

With Claude, I have too often felt like I am correcting the strategy it chose for approaching the problem at all.

That distinction matters enormously when an agent is doing hours of work inside a large repository.

More work is not automatically better work

This is one of the reasons I dislike evaluating coding agents by how much initiative they display.

Imagine that a bug can correctly be solved by changing two lines.

One agent changes those two lines, runs the relevant validation, and reports what it found.

Another agent changes twenty files, introduces a new helper, restructures a neighboring abstraction, adds several tests, and updates related code because it found opportunities to make the area “better.”

The second agent clearly did more.

That does not prove that it performed better engineering.

Every additional file increases the surface I have to review. Every additional subsystem introduces another opportunity for regression. Every new abstraction becomes something future developers need to understand. Every unrelated cleanup makes it harder to identify which part of the diff was actually required to solve the original problem.

Restraint is not the absence of intelligence.

In mature engineering, restraint is often evidence of judgment.

The important question is not:

Could I improve this too?

It is:

Does correctly completing the requested task require me to change this?

Codex has generally been better aligned with that distinction for me.

The test suite should describe the system, not the history of every edit

Claude's behavior around testing has produced a similar kind of friction.

In my repositories, Claude has repeatedly created additional tests around very small implementation changes, even where I believed the existing feature-level validation already covered the behavior.

I do not believe regression tests are bad. Quite the opposite.

If a bug exposes a genuine coverage hole, that hole should be fixed. If new externally observable behavior is introduced, it should be validated. If an important architectural invariant was previously unprotected, it deserves explicit coverage.

What I do not want is a test philosophy that effectively becomes:

Code changed, therefore another permanent test should exist.

A healthy test suite should describe durable behavior, contracts, and system invariants. It should not gradually become a chronological record of every internal line that somebody once modified.

If an existing feature-level test would fail when a one-line bug returns, then the correct validation may simply be to run that feature-level test.

This is an area where I have to be particularly careful not to overstate the comparison.

Claude Code does not have a documented rule saying “create a new test for every edit.” Codex does not have a documented rule saying “never create one.” Both systems are explicitly designed to verify their work, and both companies emphasize testing as part of reliable agentic development. Claude's own documentation describes verification as part of the agent loop, while OpenAI's engineering write-up describes automated tests and feedback loops as fundamental infrastructure for Codex.

So this is a description of my observed experience:

Claude has repeatedly been more aggressive about generating additional test artifacts than I want, while Codex has more often worked with the feature-level validation structure that already exists.

For my projects, I strongly prefer the latter.

Context management should preserve the state of the work, not the transcript of the work

The card problem becomes particularly obvious in long-running sessions.

A coding-agent session naturally accumulates information. The agent reads files, runs tests, receives command output, explores hypotheses, makes changes, discovers failures, and learns things about the repository.

Some of that information remains important.

Much of it does not.

What I want retained from a long engineering task is the useful state of the work: architectural decisions, requirements, current implementation state, important constraints, and concise conclusions from hypotheses we have already tested.

I do not need hundreds of lines of old test logs occupying attention indefinitely. I do not need an entire failed investigation to remain prominent after its conclusion has been captured. I do not need irrelevant research branches simply because the agent once considered them.

Claude sessions have repeatedly felt to me as though too much of that material remains active for too long, until context management itself becomes part of the workload.

Again, accuracy matters here: Claude Code does automatically manage context.

Anthropic documents automatic compaction and provides /clear, /compact, and /context controls. Its own best-practices documentation warns that long sessions can fill with irrelevant conversation, file contents, and command output, reducing performance and distracting Claude. Anthropic recommends clearing context between unrelated tasks and customizing what should survive compaction.

Codex also automatically compacts long-running conversations. OpenAI describes the Codex agent loop replacing accumulated input with a smaller representation once a threshold is crossed so the agent can continue with an understanding of what has happened without carrying the entire original context.

The comparison, therefore, is not:

Codex manages context and Claude does not.

Both do.

The difference is how those systems have behaved in my work.

Codex has generally produced a working state that feels closer to what I want preserved. Claude has more often made me feel as though the hand keeps filling with cards until the size of the hand becomes its own engineering problem.

There is also credible evidence going the other direction, which is worth acknowledging precisely because I do not want this article to become a rant disguised as a comparison. In an extended hands-on comparison, Composio's author actually preferred Claude Code for some extremely long, tool-heavy sessions and described an architectural detail surviving a very aggressive context compaction that they felt Codex would have lost.

That is a genuine Claude strength.

It simply does not erase the cost I experience from carrying so much context in the first place.

Running out of quota can mean two very different things

This distinction changed how I think about paying for more capacity.

I run out of Codex quota.

A lot.

It is annoying.

But the important thing is why I run out.

I usually run out because I keep giving Codex useful engineering work and want to continue giving it more.

That is primarily a capacity problem.

My experience exhausting Claude usage felt substantially different. A meaningful portion of the activity appeared to be going toward searches I did not need, investigations that expanded beyond the question, more reasoning than the task warranted, additional tests, broader changes, and the context those activities accumulated.

That is partly an efficiency problem.

Buying a larger Claude Max allowance would certainly give the agent more room to operate.

But if I dislike how a significant portion of that room is being used, buying more room does not fix the underlying mismatch.

It gives the behavior more fuel.

There is some independent evidence that this distinction is real rather than merely psychological. In Composio's 2026 comparison, Claude's Fable model completed slightly more benchmark scenarios than GPT-5.6 Sol in Codex, but Sol used fewer runtime tokens. On a separate software-engineering evaluation where both reached the same pass rate, the Codex/Sol run used substantially fewer output tokens and fewer steps, with a much lower estimated task cost.

That does not prove Codex is always more efficient.

It does show that “more computation and exploration” can be a genuine tradeoff rather than a free improvement.

“You configured Claude wrong” explains only part of the problem

None of this means configuration is irrelevant.

Claude Code is extremely configurable. It supports persistent CLAUDE.md instructions, user and project scopes, nested directory instructions, path-specific rules, permissions, skills, subagents, hooks, and automatic memory.

Codex has its own configuration hierarchy and agent harness. OpenAI describes Codex as a configurable system whose behavior is shaped by repository instructions, sandbox policies, permissions, skills, and the surrounding engineering environment.

So I am not arguing:

Codex is configurable while Claude is not.

Both products give engineers significant control.

What matters is how much behavioral correction I have to encode before the agent works the way I want.

Anthropic's own documentation provides an important limitation to the “just write a better CLAUDE.md” argument. Claude treats CLAUDE.md and auto-memory content as context, not enforced configuration. Anthropic says more specific and concise instructions are followed more consistently, recommends keeping each CLAUDE.md below roughly 200 lines, and explicitly notes that long files reduce adherence. If something must be enforced regardless of what Claude decides, Anthropic recommends using settings, permissions, or hooks instead.

That is a sensible architecture.

But it also illustrates my problem.

If my project instructions continually grow into:

Do not search externally unless necessary.

Do not create a test merely because code changed.

Do not refactor neighboring systems.

Do not create additional files unless required.

Do not introduce dependencies casually.

Do not continue exploring after enough evidence exists.

Do not turn a narrow bug into general cleanup.

then at some point I am no longer mainly documenting my project.

I am building a behavioral containment layer around the agent.

Could I keep going?

Of course.

Claude gives me more controls. I could tune effort, constrain tools, add hooks, break workflows into skills, disable or adjust memory, and create more specialized agents.

But the existence of increasingly elaborate ways to make a tool fit my workflow does not obligate me to keep using that tool.

If another agent already behaves closer to the way I want, choosing it is not evidence that I failed to configure Claude.

It is evidence that the other tool fits me better.

My workflow is probably exactly why other people reach the opposite conclusion

I think this explains a large part of why so many developers can sincerely love Claude Code while I have repeatedly found it frustrating.

Imagine that your preferred way of using an agent is:

Here is the result I want. Investigate whatever you need, decide how to implement it, improve related things if they need improvement, add whatever tests make sense, and take responsibility for delivering the outcome.

Claude's willingness to expand the search space becomes an enormous strength.

The solution space is intentionally open.

You want it to draw more cards.

My normal workflow is different.

I usually already have an architectural direction. I know the database I intend to use. I know the framework. I know which dependencies are intentional. I know the subsystem boundaries. I know what kind of validation the repository uses. I often know what must not change just as clearly as what needs to change.

I want the agent to understand those decisions and execute effectively inside them.

The solution space is intentionally constrained.

That makes Claude's broad initiative less valuable to me and sometimes actively counterproductive.

It also explains why Codex's more bounded behavior feels so much better.

I am not asking it to discover the technical identity of the product.

I am asking it to help me build the product whose technical identity I already understand.

Dependencies are a good example of why this distinction matters commercially

An agent can solve many immediate problems by installing another package.

That may be the correct decision.

But every dependency becomes part of the system I eventually have to own.

It can introduce transitive packages, vulnerabilities, licensing questions, compatibility constraints, upgrade burden, supply-chain risk, abandoned maintenance, and another API that somebody on the team will eventually have to understand.

For a prototype, my threshold for accepting that tradeoff may be very low.

For commercial software intended to exist for years, it is much higher.

So I do not want the agent to treat:

There is a package that solves this.

as equivalent to:

This package belongs in our architecture.

Those are different questions.

The first is implementation knowledge.

The second is engineering judgment.

I want to tell the agent what ecosystem I am building, what dependencies are intentional, and when adding another permanent component requires justification.

That is not preventing the agent from being intelligent.

It is part of owning the system.

Do not collect every potentially useful card simply because it exists.

Build the deck intentionally.

I cannot accept a deliverable I do not understand

This ultimately comes back to accountability.

If I commission engineering work, I need to know what I commissioned.

There must be requirements. There must be constraints. There must be acceptance criteria. Somebody must know enough about the problem to decide whether the delivered result actually satisfies them.

AI-generated code is no different.

“The agent says it works” is not acceptance criteria.

“The tests pass” is valuable evidence, but tests only validate the properties somebody knew to encode.

A commercial codebase cannot responsibly reach a point where production-critical components are treated as:

Nobody really knows what this does. The AI wrote it, and nothing seems broken.

That is not an AI strategy.

That is unmanaged technical debt.

I am completely comfortable with an agent writing code I did not personally type. The requirement is not authorship.

The requirement is ownership.

Someone responsible for the product needs to be capable of understanding the architecture, reviewing the implementation, recognizing incorrect assumptions, debugging failures, evaluating security and dependency implications, and ultimately accepting responsibility for shipping it.

That is why an agent that respects a deliberate engineering vision is much more valuable to me than one that continually expands the vision on my behalf.

The comparison outside my own experience is more nuanced—but it points to the same tradeoff

My own experience is enough to decide which tool I prefer, but it is not enough to make product-wide claims. That is why I think independent comparisons are useful.

They do not show a universal winner.

That is actually the interesting part.

Composio's 2026 comparison, written after extensive use of both systems, found Claude Code stronger in some long-context and tool-heavy sessions. Claude also completed all 47 scenarios in one agent benchmark where GPT-5.6 Sol in Codex completed 45. At the same time, Sol used fewer runtime tokens, and the author described Codex as steadier and more economical for their everyday delegation. They also reported that Codex tended to follow instructions more reliably, even while preferring Claude's broader skills ecosystem overall.

Zapier's comparison similarly does not produce a simple “Claude wins” or “Codex wins.” It describes meaningful differences in how the products encourage developers to work and notes that each is attractive for different styles of agentic coding.

Those comparisons do not prove that my preference is objectively correct.

They do something more useful.

They show that it is perfectly plausible for two competent engineers to use both tools and prefer different ones because the harness, context behavior, autonomy, and working style are different.

Which brings me back to the response that started this whole article.

No, I do not think this is simply my fault

Could I configure Claude better?

Almost certainly.

I could restrict web access more aggressively. I could lower effort. I could add stronger testing instructions. I could build hooks around file scope. I could restructure context handling. I could disable or reshape memory. I could move more workflows into skills. I could create bounded subagents for noisy exploration.

Anthropic gives me the tools to do all of those things.

But “there are more ways you could configure the tool” is not the same statement as “therefore the tool's poor fit is your fault.”

At some point a coding agent is simply a tool, and I am allowed to evaluate it by how much useful engineering work it produces relative to how much effort I spend managing its behavior.

Claude Code repeatedly makes me manage the agent's behavior in addition to managing the engineering problem.

Codex more often lets me focus on the engineering problem itself.

That is the difference.

When Codex gets something wrong, I generally feel that I am correcting a solution.

When Claude goes wrong for me, I too often feel that I am correcting the entire strategy it chose for approaching the problem: why it searched, why it expanded scope, why it created those tests, why it touched those files, why it is still investigating, or why the session now contains an enormous amount of material that was never important to the task.

That is why my answer is not to keep adding rules until Claude eventually resembles Codex.

My answer is to use the agent whose behavior already fits the engineering process I want.

Play the card

The best coding agent for my workflow is not necessarily the one capable of doing the most.

It is the one with enough judgment to understand how much actually needs to be done.

Sometimes the correct solution really does require broad research, a large refactor, several new tests, a new dependency, and twenty changed files.

When that is true, I want the agent to do all of it.

Sometimes the correct solution is one line.

When that is true, I want the agent to change one line.

The hard part is knowing the difference.

That is the judgment I want AI to amplify.

Claude is extraordinarily capable of finding more cards.

For developers who intentionally leave the solution space open, that may be exactly what makes it their preferred tool.

My engineering process usually begins from a different place. The architecture exists. The constraints matter. The dependencies are deliberate. The human still owns the technical vision.

I want the agent to understand the game, study the hand, recognize when it has enough information, and make the correct move.

For the way I build software and hardware today, Codex does that much better for me than Claude Code.

And that is not because I failed to configure Claude until it behaved like Codex.

It is because the agent should fit the engineering process—not require the engineering process to be continually redesigned around the agent.

Stay Tuned

Want to stay up to date with the latest posts?
The best articles, links and news delivered once a week to your inbox.