Don't trust us — check it

How we tested

This is the page where we hand you the stick to hit us with. Every method, every caveat, and the numbers that make our own conclusions look shakier.

The obvious objection

"An AI graded this. Of course it picked its favorite."

Fair. So we had two AIs mark every answer — one from each company — separately, against the same written answer key, neither knowing what the other said. Then we checked whether either one went easy on its own side.

How this is measured
🤝
Awkward but true

Half our own questions barely tell them apart

A benchmark question that everyone gets right measures nothing. Here's how well each of our questions actually separates the field — including the ones that don't, which we're leaving in and labeling rather than quietly dropping.

The unglamorous one

How many times we actually ran each thing

Much of this board rests on a single run. A chart that hides that is overclaiming, so here it is. More runs means we know how consistent a model is, which is a different question from how good it is.

The rules we set ourselves

The boring but important bit

If you don't say how you measured, you haven't published a result — you've published an opinion with numbers attached.

🛠️ Tools stay switched on

Both sides can search the web and use their tools, because that's how people actually use them. Testing with tools off measures a setup nobody runs. We mark the answer and record what it cost.

🔒 One question, one fresh chat

First answer counts. No retries, no re-wording, no "try again but better." That's how you'd actually use it.

🙈 Marked blind

Answers are saved under random filenames. The model's name lives only in the results file and gets joined back after marking is done.

🧾 We store tokens, not dollars

Prices move constantly. Baking them into the data would make every old result un-recalculable. Costs are applied at display time from a dated price table — which you can edit.

🖼️ Identical images

Picture questions use the same pre-rendered file for every model. Re-screenshotting per model would mean different crops and meaningless comparisons.

📖 Answer keys stay shut

Nobody reads the key until collection is finished. Knowing the answer changes how you mark an ambiguous response, and you won't notice it happening.

Where to hit us

What you should be suspicious of

The receipts

Every question, every setup

Nothing above is a vibe. Pick any two setups and compare them question by question. Turn on Nerd mode for the graders' own reasoning.

ID Question Skill A B

Remember when Windows 10 came out?

We were told it was "better." We were told it was "more secure." What we were never told was what was better, or which part was more secure, or how anyone would ever know. It took years for regular people to work out what had actually changed.

And we still miss XP, damn it.

AI is doing the exact same thing right now — "more capable," "more aligned," "state of the art" — and the models themselves are the worst possible source, because every one of them will cheerfully tell you it's excellent at something it just failed. So we stopped asking and started marking.