Skip to content

Published results

Prompt experiments, shown with their working: how many test cases, how answers were graded, the best prompt and model, and how sure the result is.

How to read a result

Score
Out of 10: the median of several graders, each blind to which prompt or model wrote the answer.
Range
Where the score would likely land on other test cases like these. Overlapping ranges mean too close to call.
Cost
What the winning prompt and model cost per 1,000 calls, at the model maker's price.
Test cases
How many inputs every candidate answered. More cases, narrower ranges.

Standing experiments

Fixed test cases and grading guides we re-run every month as models change: regression tests for the whole market.

Shared results

No published results yet.

When someone publishes an experiment, it appears here with its best prompt and model, its score, range and cost. Yours could be the first.

Try your first case free