Write an LLM-as-Judge Scoring Rubric
Create a clear rubric for judging model outputs on a task — dimensions, a scoring scale, and anchored examples.
0 likes
0 dislikes
Sign in to rate this prompt
Prompt
Write a scoring rubric for judging outputs of the task below, usable by a human reviewer or an LLM judge.
The task whose outputs I'm judging:
{{task}}
What "good" means to me (optional — the qualities I care about most): {{criteria}}
Produce:
1. **Dimensions** — the 3–6 things to judge (e.g. correctness, completeness, faithfulness to the source, format, tone). Define each in one line so it's unambiguous.
2. **Scale** — a small scoring scale per dimension (e.g. 1–5, or fail / partial / pass) with a short description of what each level means.
3. **Anchors** — for the most important dimension(s), a brief example of a high-scoring vs. a low-scoring output so the judge calibrates.
4. **Judging instructions** — a short prompt a judge could follow, including telling it to justify each score with evidence from the output and to score only what's present (not to reward intentions).
Keep the dimensions independent so they don't double-count. If my notion of "good" is vague, propose a reasonable default and flag it.
Why this prompt works
-
How to Trace and Debug an AI Agent
Print statements don't survive a system whose control flow is decided at runtime by a model. What every span must carry, the OpenTelemetr...
Read the guide → -
Why AI Models Hallucinate
Hallucination isn't a malfunction: benchmarks reward confident guessing and penalise "I don't know", so training selects for it. Purpose-...
Read the guide → -
Claude vs Gemini vs ChatGPT
Google's Flash rate doubles on 1 January 2027 and its flagship is still preview; OpenAI spans 50x from Luna to Astra with no retirement d...
Read the guide →