● Updated · not one number came from a vendor

Which AI is actually better?

We ran the tests.

Everybody says their AI is the best. Nobody says at what. So we sat every model down in front of the same exam — 40 questions, 147 separate things to get right — and had two rival AIs mark every single answer independently. Then we wrote down what happened, including the bits that make our own conclusions look shaky.

Two cartoon robots sitting at school desks taking the same paper exam. One writes confidently; the other chews its pencil and sweats.
Same paper. Same grader. Turns out they're good at different questions.
The one chart that matters

They nearly all score the same. They do not nearly all cost the same.

Here is every model's mark against what that mark cost. The left column is almost flat — modern AI is, annoyingly, mostly quite good. The right column is not flat at all.

How well it did What it cost to find out

So what does vary?

Five skills separate them. Five don't.

We expected the models to split on coding. They don't — 37 of 38 tie on it. They split on things nobody puts in a benchmark, and only on a handful of questions. Here's where the difference is real and where you're overthinking it.

✅ These sort them out

Most of the field falls behind. Worth choosing on.

😴 These barely tell them apart

Most of the field finishes in a heap. Pick on price instead.

Things we are not going to pretend about
🤖 "Was this site made by AI?" Of course it was. We'd be idiots not to. That's rather the whole point.
✍️ "Did AI write the test questions?" Yes. You think we're clever enough to catch out a multi-billion-pathway model that's had more learning experience in the last five minutes than the entire history of the universe?
🧠 "Were the questions dumbed down?" Some of them, so us mere mortals could follow the results. The nasty ones are still in there. They count.
🔌 "Were any datachips harmed?" No datachips were fried in the gathering of this information. …to our knowledge.
💸 "Who paid for this?" We did, out of our own pocket, and we have the receipts to prove it — literally, they're on the costs page.
🙋 "Are you experts?" God no. We're regular people who got sick of paying for the wrong subscription and guessing. That's exactly why the answers here are in English.
🕵️ "Who are you, then?" Nobody you've heard of, and we're keeping it that way. Don't trust us — check the method. It's all on the method page precisely so you don't have to take our word for anything.
🏢 "Is anyone paying you to say this?" Nobody has approved, reviewed, or been warned about a single word here. We rather suspect a couple of them would like a quiet chat.
The good bits

Five things we did not expect

Every one of these came out of the data rather than out of us, which is why a couple of them are inconvenient.

Remember when Windows 10 came out?

We were told it was "better." We were told it was "more secure." What we were never told was what was better, or which part was more secure, or how anyone would ever know. It took years for regular people to work out what had actually changed.

And we still miss XP, damn it.

AI is doing the exact same thing right now — "more capable," "more aligned," "state of the art" — and the models themselves are the worst possible source, because every one of them will cheerfully tell you it's excellent at something it just failed. So we stopped asking and started marking.