Same method on every plan. You choose how big you test.
Every plan defines success first, grades blind and puts a range on every score. Plans differ only in how large your experiments are and how often you re-test. Model calls go on your own key, with nothing added.
Trial Pass$19once14 days of access, no renewalSettle one decision with a full-size experiment. No subscription.Buy a Trial Pass | SoloMost builders$29a monthKeep improving your prompts, with regression tests when models change.Start Solo | |||
|---|---|---|---|---|
| Size of each experiment | ||||
| Test cases per trialMore cases give narrower score ranges | 25 | 200 | 500 | 5,000 |
| Candidates per trialPrompt × model combinations compared | 3 | 6 | 6 | 12 |
| Graders per trialMore graders give a steadier median | 2 | 3 | 3 | 5 |
| DeliberationRounds of rival prompts critiquing each other | 1 round | 3 rounds | 3 rounds | 5 rounds |
| How often you iterate | ||||
| CasesA prompt and its tests, saved together | 1 | 1 | 10 | Unlimited |
| TrialsEach test of a change uses one | 3 a month | 10 in total | 30 a month | 150 a month |
| RetrialsAutomatic regression tests on a schedule | Not included | Not included | Monthly | Weekly, plus new-model alerts |
| Evidence and reports | ||||
| Reports | Public verdict page with PromptJury branding | Public verdict page, plus PDF | Private pages, PDF | White-label PDF and pages |
| RetentionHow long answers and grader reasons are kept | 14 days | 30 days | 6 months | 2 years |
| Data importBring your own test cases | CSV | CSV | CSV, OpenRouter logs | CSV, OpenRouter logs |
| Seats | 1 | 1 | 1 | 5 |
The method is the same on every plan
A bigger plan never buys a different standard of proof. Free and Pro run the same controls; Pro just runs more, larger experiments.
- Success defined firstYou approve the grading guide before anything is scored.
- Same test for everyoneEvery prompt and model answers the same test cases.
- Graded blindGraders never see which prompt or model wrote an answer, and no model grades itself.
- Median of several gradersOne generous or harsh grader can't swing a score.
- A range on every scoreAnd a plain “too close to call” when the top two overlap.
- A suggested fix per modelWhat the prompt failed to tell each model, ready to re-test.
- Cost before every runAn estimate you confirm before anything spends money.
- Every number has a receiptOne click from any score to the answer and the grader's reason.
Questions about plans
How many test cases do I need?
Enough for the ranges to separate. A clear winner often shows with 25. When the top candidates are close, add cases: the range on a score shrinks roughly with the square root of the number of cases, so four times the cases about halves it. The verdict tells you when it's too close to call.
What counts as a trial?
One run of every chosen prompt and model against your test cases, with grading. Testing a change, re-testing after a suggested fix and each scheduled retrial each use one.
What happens if I hit a limit?
The trial is checked before it starts, so nothing runs halfway. You'll see which limit you reached and can upgrade or wait for the next period.
Who pays for the model calls?
You do, directly, through your own OpenRouter account at the model makers' prices. We add nothing, and every run shows its estimated cost first.
Does the Trial Pass count toward Solo?
Yes. If you move to Solo within 14 days of buying a Trial Pass, its $19 comes off your first Solo invoice.
Can I cancel?
Any time, from Manage billing in Settings. You keep your plan until the end of the period, and nothing is deleted when you move down a plan.
Not sure yet? Run your first case free. No card.