Skip to content
Pricing

Same method on every plan. You choose how big you test.

Every plan defines success first, grades blind and puts a range on every score. Plans differ only in how large your experiments are and how often you re-test. Model calls go on your own key, with nothing added.

Free$0Run your first experiment on one prompt.Try your first case free
Trial Pass$19once14 days of access, no renewalSettle one decision with a full-size experiment. No subscription.Buy a Trial Pass
SoloMost builders$29a monthKeep improving your prompts, with regression tests when models change.Start Solo
Pro$99a monthProve your choices to clients, under your own brand.Start Pro
Size of each experiment
Test cases per trialMore cases give narrower score ranges252005005,000
Candidates per trialPrompt × model combinations compared36612
Graders per trialMore graders give a steadier median2335
DeliberationRounds of rival prompts critiquing each other1 round3 rounds3 rounds5 rounds
How often you iterate
CasesA prompt and its tests, saved together1110Unlimited
TrialsEach test of a change uses one3 a month10 in total30 a month150 a month
RetrialsAutomatic regression tests on a scheduleNot includedNot includedMonthlyWeekly, plus new-model alerts
Evidence and reports
ReportsPublic verdict page with PromptJury brandingPublic verdict page, plus PDFPrivate pages, PDFWhite-label PDF and pages
RetentionHow long answers and grader reasons are kept14 days30 days6 months2 years
Data importBring your own test casesCSVCSVCSV, OpenRouter logsCSV, OpenRouter logs
Seats1115

The method is the same on every plan

A bigger plan never buys a different standard of proof. Free and Pro run the same controls; Pro just runs more, larger experiments.

  • Success defined firstYou approve the grading guide before anything is scored.
  • Same test for everyoneEvery prompt and model answers the same test cases.
  • Graded blindGraders never see which prompt or model wrote an answer, and no model grades itself.
  • Median of several gradersOne generous or harsh grader can't swing a score.
  • A range on every scoreAnd a plain “too close to call” when the top two overlap.
  • A suggested fix per modelWhat the prompt failed to tell each model, ready to re-test.
  • Cost before every runAn estimate you confirm before anything spends money.
  • Every number has a receiptOne click from any score to the answer and the grader's reason.

Questions about plans

How many test cases do I need?

Enough for the ranges to separate. A clear winner often shows with 25. When the top candidates are close, add cases: the range on a score shrinks roughly with the square root of the number of cases, so four times the cases about halves it. The verdict tells you when it's too close to call.

What counts as a trial?

One run of every chosen prompt and model against your test cases, with grading. Testing a change, re-testing after a suggested fix and each scheduled retrial each use one.

What happens if I hit a limit?

The trial is checked before it starts, so nothing runs halfway. You'll see which limit you reached and can upgrade or wait for the next period.

Who pays for the model calls?

You do, directly, through your own OpenRouter account at the model makers' prices. We add nothing, and every run shows its estimated cost first.

Does the Trial Pass count toward Solo?

Yes. If you move to Solo within 14 days of buying a Trial Pass, its $19 comes off your first Solo invoice.

Can I cancel?

Any time, from Manage billing in Settings. You keep your plan until the end of the period, and nothing is deleted when you move down a plan.

Not sure yet? Run your first case free. No card.