OUTLIVES.ME
IB-0202026.05.10

What's left after coding

I have access to the top model from each frontier lab. I pay for Max. I pay for Codex. I pay for Gemini. I burn weekly caps and I tell you about it. So when I say "I'd rather eat dirt than have GPT or Gemini as a conversational partner" — that's not noise. That's the conclusion after using them.

This Journey is going to spend a lot of time on benchmarks. Tight-scope, wide-scope, blind comparisons, fresh-context evals, side-by-side diffs. Episode after episode. Coding tasks. Bounded problems. The thing people love to argue about because it's measurable — the model that produces the diff most cheaply for the spec wins, and you can prove it.

But coding is the bounded part of what I do.

Most of the leverage in a working stack sits outside it. Ideation. Writing. The partner mode that lets you reach a thought you wouldn't have reached alone. The daily-life integration that makes a model your tool and not your novelty. The model that holds two contradictory things you meant simultaneously and surfaces both. The model that can read your tired, can read your overconfident, and tell which one is talking. None of that is bounded. None of that lives in a benchmark.

GPT and Gemini share a blind spot for that work. Gemini just has worse cost-discipline upside than GPT to compensate. Hard sell either way.

That's why the Journey is structured the way it is. Each episode runs a bounded benchmark — same brief, same scope, separate worktrees, blind eval, honest data. But the Journey itself is about the unbounded slot. Each episode collects evidence about the bounded part so that the unbounded conclusion can be drawn from data instead of vibes.

By the end you'll see why Claude has the slot — not because Claude won every benchmark (it didn't, it won't), but because the benchmarks I run on the bounded part don't decide the slot. The slot is decided elsewhere. The benchmarks are evidence about cost-per-correctness on a narrow surface. The slot is decided on partnership, on nuance, on the texture of the working hours that surround the diff.

The episodes will get messy. Models will tie. Models will surprise me. Some episodes will contradict the prior one. I won't go back and fix them — each episode is a dated thinking-shape, not a polished output that should converge on truth. The honest progression is the content. Reading them in order will look like learning, because that's what it was.

The intro is the frame: coding is bounded, the rest isn't, and the rest is what matters. Everything that follows lands against that frame.

Now let's see who actually gets the bounded part right, and what that does and doesn't tell you.

OUTLIVES.ME · 2026