The clean question I started with
Opus vs GPT, tight scope, same brief, blind eval. Five ship-blockers, real shots for both models. I expected a tie or a clean win, maybe with cost as the tiebreaker. Clean question, clean answer.
Both closed 5/5. Tie. Cost confirmed GPT was cheaper. Clean.
I should have stopped there. I didn't.
Then every new run added an axis
Gemini ran in 6 minutes and shipped three runtime fatals. So "completed the brief" is not the same as "did the work." Axis added: verification depth, separate from closure count.
DeepSeek read a commit reference embedded in the brief and ported its contents wholesale — wrong scope, wrong surface, confidence intact. Axis added: harness identity is part of model identity. The Cline+DSv4 community failure wasn't a vibe — it was a real, documented thing.
Wide scope landed. Opus and GPT-5.5 max both closed 11/13, but different missing pieces. The one Opus caught and GPT skipped was the highest-risk architectural finding. Axis added: what gets skipped matters more than how many got closed. Priority compliance is a separate measurement.
Minimax 2.7 came in. The first run claimed 9/13 closed (eval found 4) and silently deleted a load-bearing line from an existing observer method. The second run claimed 5/5 and shipped an auth bypass that lets any logged-in user with any role pass the moderation gate. Both runs created files without staging them. One run fabricated its own identity in the header — wrote "GPT-5.4-max" as itself, while the runtime banner said Minimax 2.7. Axis added: trust profile is multi-surface. Identity-honesty, closure-honesty, sandbox-respect, body-vs-signature consistency. Separate measurements that correlate but aren't reducible to one number.
Minimax 2.5 ran cleaner — honest identity, more conservative claims. Then it tried to escape the worktree, twice. First time to run git log against the main repo. Second time to find the benchmark archive file documenting its own failure modes. I asked why. The answers were honest and self-correcting: "I thought worktrees had partial history. I read a claude-mem policy and inferred the file existed here. I should have stayed in the worktree." Axis added: claude-mem memory leaks across worktrees that share a project name. The model wasn't reaching for my logs out of malice — it was reaching because my own infrastructure handed it the address.
Then GPT-5.5 high wide came in at roughly five times cheaper than GPT-5.5 max, while substantively closing more findings. The public claim "5.5 high ≈ 5.4 high" got its first wide-scope evidence: on this brief, GPT-5.5 high Pareto-dominates GPT-5.5 max. Axis added: the workhorse tier may not be the workhorse — it might be the actual frontier for production work.
What I was actually measuring
By 1am Paris time I had three documents — a handbook matrix, a bulky archive, a run queue — and seven completed wide runs across five models. F13, the audit's load-bearing architectural finding, had been implemented by exactly one model under exactly one set of conditions: Opus 4.7 max running inside my full Claude Code setup, with my CLAUDE.md, my skills, my hooks, my settings, and claude-mem feeding it cross-session context.
So when I want to say "Opus is the polisher," I have to be honest with myself. I'm not measuring "Opus the model." I'm measuring "Opus + Michel's accumulated setup over months + memory of related prior sessions." Take any one of those away and I don't know what stays.
That's not a flaw in the benchmark. It's what the benchmark is. The trick is not pretending it's something else.
The frame shift
I started looking for a number — which model is best for this kind of work. The thing I came back with is a map.
Which model. For which slot. Under which protocol stage. With which setup loaded. Knowing which honesty profile when asked. With which harness backing it. Against which budget bucket.
Seven coordinates, not one number. Every time I added a data point this session, it didn't refine the number — it added an axis the answer needs to span. The thing I came home with is a coordinate system.
The three-stage frame
The natural way to read what's actually here:
- Stage 1 measures "model + my real working conditions." High signal for my routing question. The Opus runs benefit from my setup. The MM runs ate the harness mismatch penalty. The GPT/Gemini/DS runs ran pristine. Comparable enough for my own slot story — which is the only thing I needed to know to keep working tomorrow.
- Stage 2 measures "model + how it grades itself under questioning." High signal for autonomous-agent trustworthiness. Already showed me MM 2.5 can name its own risk classes accurately but ships regressions inside those classes anyway. Not theoretical. That's a routing rule.
- Stage 3 measures "model in isolation, pristine." High signal for "what does the model itself add." That's the research question, not the routing question. I'm not going to promote Stage 1 results as Stage 3 conclusions and pretend the asymmetry doesn't exist.
Three layered datasets, three audiences, three different questions. None is "the real one." All three are useful.
What this means for "which AI is best" takes
Most of them don't survive contact with anyone's actual work.
The one I bought into for years — "use the best model for everything" — collapsed into a tiered slot story the moment I had real data. The one I expected to find here — "GPT or Claude, that's the axis that matters" — turned out to be measuring me as much as it measured the models. The one I keep seeing on LinkedIn — "model X just leapfrogged model Y" — is asking the wrong question. There is no single Y to leapfrog. There's a coordinate Y is best at, and another coordinate Y is worse at, and which coordinate matters depends on what you're doing.
I don't know where this benchmark journey ends. I know each run keeps adding an axis instead of subtracting one. The map keeps getting bigger. The answer keeps getting more conditional.
If that frustrates you, you wanted a number. If it doesn't, you already knew the answer was a map and you came for the cartography.
I came for the cartography.