Write an LLM-as-Judge Scoring Rubric
Create a clear rubric for judging model outputs on a task — dimensions, a scoring scale, and anchored examples.
0 likes
0 dislikes
Sign in to rate this prompt
Prompt
Write a scoring rubric for judging outputs of the task below, usable by a human reviewer or an LLM judge.
The task whose outputs I'm judging:
{{task}}
What "good" means to me (optional — the qualities I care about most): {{criteria}}
Produce:
1. **Dimensions** — the 3–6 things to judge (e.g. correctness, completeness, faithfulness to the source, format, tone). Define each in one line so it's unambiguous.
2. **Scale** — a small scoring scale per dimension (e.g. 1–5, or fail / partial / pass) with a short description of what each level means.
3. **Anchors** — for the most important dimension(s), a brief example of a high-scoring vs. a low-scoring output so the judge calibrates.
4. **Judging instructions** — a short prompt a judge could follow, including telling it to justify each score with evidence from the output and to score only what's present (not to reward intentions).
Keep the dimensions independent so they don't double-count. If my notion of "good" is vague, propose a reasonable default and flag it.