Which AI is actually better?
We ran the tests.
Everybody says their AI is the best. Nobody says at what. So we sat every model down in front of the same exam — 40 questions, 147 separate things to get right — and had two rival AIs mark every single answer independently. Then we wrote down what happened, including the bits that make our own conclusions look shaky.
They nearly all score the same. They do not nearly all cost the same.
Here is every model's mark against what that mark cost. The left column is almost flat — modern AI is, annoyingly, mostly quite good. The right column is not flat at all.
Five skills separate them. Five don't.
We expected the models to split on coding. They don't — 37 of 38 tie on it. They split on things nobody puts in a benchmark, and only on a handful of questions. Here's where the difference is real and where you're overthinking it.
✅ These sort them out
Most of the field falls behind. Worth choosing on.
😴 These barely tell them apart
Most of the field finishes in a heap. Pick on price instead.
Five things we did not expect
Every one of these came out of the data rather than out of us, which is why a couple of them are inconvenient.
Remember when Windows 10 came out?
We were told it was "better." We were told it was "more secure." What we were never told was what was better, or which part was more secure, or how anyone would ever know. It took years for regular people to work out what had actually changed.
And we still miss XP, damn it.
AI is doing the exact same thing right now — "more capable," "more aligned," "state of the art" — and the models themselves are the worst possible source, because every one of them will cheerfully tell you it's excellent at something it just failed. So we stopped asking and started marking.