Skip to content
2026 inaugural edition
Preview stub · v0.1-draft·stubbed 2026-04-26

State of AI Lyric Output 2027Preview · full report ships 2027-Q4

This page is a stub published at the same dated URL the full report will ship at. It enumerates the dimensions the 2027 report will measure against the 2026 baseline + links the live data sources feeding each one. The discipline of naming the measurements before publication is the contract that they'll be honored at publication.

Why publish a stub

Reports that materialize once a year tempt their authors to choose what to measure based on what already looks favorable. Pinning the dimensions in advance (with public URLs to the live data sources) removes the temptation. The 2026 edition did the same with its observation list; the 2027 edition extends the contract by adding the new measurement surfaces that landed this year (Hum Test, triangulation, cohort divergence).

If you are reading this page in late 2027 expecting the full report and instead see this stub, the report is overdue and someone has dropped the cadence. Email support@songforgeai.com and ask why.

What the 2027 report will measure

  • 01

    Year-over-year movement on the same Lyric Scoring Standard rubric.

    2026 baselines published in the inaugural edition. 2027 will report the same five observations re-measured against rubric v1.x — including how MINOR/MAJOR bumps under RFC-0001 affected the comparison and what that says about category-wide movement vs rubric-level movement.

    2026 baseline
  • 02

    24-hour-delayed Memorability (Hum Score) drift from fresh M11.

    The Hum Test cron began writing hum_score in 2026. By Q4 2027 the corpus will have ~12 months of paired (fresh M11, hum M11) data. The 2027 report measures the corpus-wide median delta and what it implies about whether the rubric is over-scoring memorability that doesn't survive the day. Methodology pinned in RFC-0003.

  • 03

    Cross-family triangulation divergence (Sonnet ↔ GPT-4o).

    Every public /api/v1/score request triangulates against GPT-4o (B1279). The 2027 report aggregates the corpus-wide divergence distribution — % of scores agreeing within 3pts, 8pts, >8pts — and identifies any score band where the two model families systematically disagree. The cross-family number is the bias-defense any third-party citation will look for.

  • 04

    Per-cohort listener engagement vs composite score divergence.

    Five-persona listener panel (B1281+) produces engagement scores per song. The 2027 report measures how often the composite over-scores or under-scores the mean cohort engagement and which cohort is the most pessimistic listener corpus-wide. The signal is the second-half answer to "did the rubric reward what the audience actually wants."

  • 05

    Open standard adoption: third-party citations of the Lyric Scoring Standard.

    The standard shipped CC BY 4.0 in 2026 (B940). The 2027 report counts third-party citations — npm package installs, GitHub repos referencing the rubric, blog posts that cite a specific metric definition. The number is the leading indicator of whether the standards moat is converting from "published" to "used."

    standard + license
  • 06

    Velocity moat: build cadence vs quality ratchets.

    2026 baseline: ~36 builds/day average, 0 explicit-any allowlist, 4 forward-only ratchet gates. The 2027 edition reports the cadence change + the ratchet count change + the % of builds that triggered a ratchet movement. The compound question is "did we ship faster AND tighter, or did one trade off against the other."

    live engineering report

Year-over-year discipline

The 2026 inaugural edition published numbers and methodology. The 2027 edition adds the year-over-year column. Each observation in the 2027 report will read as: 2026 baseline → 2027 measurement → delta + commentary. Where the rubric version differs (per RFC-0001 versioning policy), both the old-rubric and new-rubric numbers will be reported so the comparison stays honest.