Solo builders
“Tell me which prompt and model to ship, and show me why.”
You can wire an API call. You don't have time to build a test harness. Get a tested prompt in one sitting.
PromptJury treats your prompt like an experiment. Define what a great answer looks like, test every prompt and model on realistic cases, and keep the one that scores highest.
No card. You bring an OpenRouter key, and you see the estimated model cost before anything runs.“The vase arrived cracked. It's been 40 days because I was travelling. Can I get my money back?”
I'm sorry it arrived broken. Damaged items are always refunded, whatever the date, so you'll get the full amount back. Reply with a photo of the vase and I'll process it today.
prompt v3 · openai/gpt-4.1Unfortunately our refund window is 30 days, so we can offer store credit instead.
prompt v1 · same modelWhy v1 lost: “Misses that damaged items are always refunded.” 3 of 3 graders, graded blind.
Most prompts are tuned by feel.
You change a line, read three outputs, and decide it's better. That's a hunch, not a result. It holds until a customer asks something you didn't try. PromptJury writes the test first, runs the tries you didn't have time for, and keeps the receipts.
Six steps, in one sitting. Each starts with a draft you can accept or change, so you never face a blank form, and none of it needs code.
Two sentences on what your AI should do and who reads the answer.
Several models each draft one, then critique each other's. You start with rivals to test, not a single guess.
Approve what a great answer looks like, dimension by dimension, with a scale for each. The test exists before the prompt is judged.
Realistic inputs are written for you, including the hard and awkward ones. Or upload your own.
Every prompt and every model answers every test case. A separate panel grades each answer blind.
See the best prompt, the best answer and its score. Each model gets a suggested fix; re-test it and watch the score move.
“One model grading its own work is a student marking their own exam.” So the grading is a controlled experiment.
The verdict says what each prompt failed to tell each model. Apply a fix, run the same test cases again, and keep the change only if the score goes up. Every version has to earn its place.
Saved cases become regression tests. Schedule a retrial and PromptJury re-runs them when models change, so you hear about a drop before your customers do.
An illustrative experiment: support replies for a home goods shop, five models, 40 test cases, three graders. Every model's score sits next to its cost, so you can see the best answer and the cheapest one that is nearly as good. Pick a model to see its numbers.
“Holds the 30-day policy and offers store credit instead.”
You connect your OpenRouter account, so calls go through it at the model makers' own prices.
We add nothing to model costs. Your plan pays for PromptJury; your key pays for the calls.
Nothing that spends money starts until you've seen the estimated cost and confirmed it.
“Tell me which prompt and model to ship, and show me why.”
You can wire an API call. You don't have time to build a test harness. Get a tested prompt in one sitting.
“I have nothing to back up my model choice.”
Hand clients a report with your logo on it: the method, the test cases and the scores behind your choice.
“I only hear about it when it's bad.”
Schedule a retrial and get a plain pass or fail by email, before a customer notices.
You fix the definition of a good answer before anything is scored, every prompt and model faces the same test cases, grading is blind, and each score comes with a range. Change one thing, re-run the same test, and you can see whether it actually helped.
It finds the best prompt among the ones tested, against the standard you approved, and shows how sure that result is. Then it suggests what to try next. Nothing can promise perfect; this gets you measurably better, with proof.
You approve the grading guide, and the verdict shows sample answers so you can check the scores against your own reading.
It can, which is why the jury mixes makers and never sees which model wrote an answer. No model grades itself. You can see every juror's score and reason, and we flag dimensions where they disagreed.
You can. It takes a day per prompt, and you will not do it again when the next model ships.
Every trial shows an estimate first, and nothing that spends money runs until you confirm it. You can also cap spending on the key you connect.
That is the case PromptJury is built for. Subjective quality is graded on dimensions you approve.
Yes. It takes a couple of minutes, and one key reaches models from every major maker. We walk you through connecting it on your first case.
Yes. Upload a CSV and map its columns, or edit the cases we generate. Most people do a bit of both.
No. Your prompts, test cases and answers are used only to run your trials, and you can delete a case at any time.
We say so in plain words instead of crowning a winner. The verdict shows the range for each model, and the cheaper one is usually the sensible pick.
Your first case is free, with no card. If the result does not tell you something you did not know, you have lost half an hour.