Skip to content
All comparisons
Standard vs. ChatGPT-as-judge

Lyric Scoring Standard vs. ChatGPT-as-judge

ChatGPT can score lyrics if you ask it to. So can Claude, Gemini, and any LLM. The honest comparison isn't 'does it work' — it does. The honest comparison is 'is the score grounded in anything reproducible.'

Calibration

ChatGPT-as-judge

ChatGPT-as-judge has no calibration corpus. Two prompts to the same model on the same lyrics produce different scores. The model's internal sense of 'this is an 85' has no anchor — and tends to land in the 70-90 range for almost any input.

Lyric Scoring Standard

The Lyric Scoring Standard ships with a hand-scored reference corpus (currently 17 entries — F-band through S-band). The 1949 Hank Williams song 'I'm So Lonesome I Could Cry' is the canonical 95-point anchor. Every score lands relative to those anchors via the four anti-inflation rules (Gravity / Burden of Proof / Antagonist Ceiling / Historical Context).

Anti-inflation discipline

ChatGPT-as-judge

ChatGPT defaults to inflation. 'Score these lyrics' returns a 78 even for plain bad lyrics, because the model's training data rewards encouragement.

Lyric Scoring Standard

The Gravity Rule sets the default expectation at 50 (population mean). Burden of Proof requires evidence to score above 70. Antagonist Ceiling caps inflation when one specific aspect is failing. Historical Context Anchor pins band labels to era-appropriate exemplars. A weak draft scores in the 30-45 band; a strong draft earns 80+.

Reproducibility

ChatGPT-as-judge

Same lyrics + same prompt + 24h later: different score, possibly by 10+ points. Model drift, system prompt drift, no version pinning.

Lyric Scoring Standard

Every /api/v1/score response carries the ed25519-signed seal. Pin the rubric version + the model id + the temperature; replay produces the same score within the published reproducibility-audit thresholds (±1pt median, RFC-0007).

Cost + speed

ChatGPT-as-judge

Effectively free for individual use. Fast — same speed as any LLM call.

Lyric Scoring Standard

Free tier: 5 songs/month covers full pipeline (forge + score + refine). Score-only API at /api/v1/score is metered; comparable cost to the underlying model call. Speed: scoring takes 5-10 seconds, gauntlet refinement adds 30+ seconds, full forge takes 60-120 seconds.

Verdict

ChatGPT-as-judge is the right tool for a one-off vibe check. The Lyric Scoring Standard is the right tool when you need a number that means the same thing tomorrow as it does today, that another implementer can reproduce, and that doesn't drift toward 'looks good' just because the lyric is well-formed. The trade-off is calibration discipline + reproducibility for free + ubiquity. Pick the right tool for what you're doing.