Published results
Prompt experiments, shown with their working: how many test cases, how answers were graded, the best prompt and model, and how sure the result is.
How to read a result
- Score
- Out of 10: the median of several graders, each blind to which prompt or model wrote the answer.
- Range
- Where the score would likely land on other test cases like these. Overlapping ranges mean too close to call.
- Cost
- What the winning prompt and model cost per 1,000 calls, at the model maker's price.
- Test cases
- How many inputs every candidate answered. More cases, narrower ranges.
Standing experiments
Fixed test cases and grading guides we re-run every month as models change: regression tests for the whole market.
- Standing experiment · re-run monthly
Best model for a support reply
Refund and order questions for a small online shop, graded on policy, tone and next step.
First run pending - Standing experiment · re-run monthly
Best cheap model for a support reply
The same case, limited to models under a set price per 1,000 calls.
First run pending - Standing experiment · re-run monthly
Best model for a sales email
First-touch cold email to a small business, graded on relevance, clarity and a clear ask.
First run pending
Shared results
No published results yet.
When someone publishes an experiment, it appears here with its best prompt and model, its score, range and cost. Yours could be the first.
Try your first case free