Skip to content
Back to Roadmap
Engineering Punch List

The discipline is the moat.

Prospective-discipline list driving SongForgeAI engineering. 50 items across 4 tiers, each ratcheting a CI gate or unlocking a real capability. Started Build 1173 (renamed from "Linear List" at B1216). Updated with every notable build — this page reads docs/PUNCH-LIST.md at request time so the floor is per-deploy fresh.

Total

245

Shipped

214

Partial

2

Progress

88%

The Punch List

Renamed from "The Linear List" at Build 1216. The activation phrase Continue Linear List still works for backwards-compat with old session prompts; the canonical phrase going forward is Continue Punch List. Internal commit subjects from B1173-B1215 retain the "Linear List #N" tag; they're matched by the engineering-report regex alongside "Punch List #N" and "Excellence #N".

Mission: 100th-percentile engineering discipline as a public-facing artifact. The discipline is the moat; this list ratchets it deliberately.

Status: Started Build 1173 (post-extraction-batch). Roll new sessions with this file as the reference.


Activation

When the user types `Continue Punch List` (or the legacy Continue Linear List), the agent should:

1. Read this file. 2. Find the next unchecked item (highest leverage given current state). 3. Scope it as one or more builds. 4. Ship it (typecheck → tests → commit → push). 5. Mark the checkbox here, reference the build number(s).

Order is not strictly sequential. Tier A first, but skip ahead when:

  • A Tier B/C/D item is obviously cheaper and unblocks others.
  • Surface an explanation when reordering.

Each item carries a one-line scope. If a scope expands beyond ~3 builds, split it into sub-items here first.


Song Surgery — 12-18 build flagship feature (DEFERRED to post-Release-Dossier)

Triggered by: operator-proposed concept + B3148 100-expert / 100-round WAR Room (2026-05-23). Verdict: "Asymmetric upside; bounded downside. Build it" — but NOT until the Release Dossier closes through Phase 8.

The reframe: "Put this song through Song Surgery" replaces "regenerate this song." Same machinery (the Release Dossier engines from Phases 2-6 already do 80% of the work — runAllReportQaGates + computeRevisionRoi + computeLineLevelSurgery + Refine flow); different UX wrapper. The marketing value is high (category-creating positioning: "Don't regenerate your song. Save it.") and the engineering cost is low (the engines exist; what's needed is the workflow + one new Haiku primitive + tier gating).

Tagline candidates:

  • "Song Surgery: line-by-line lyric repair until your song is release-ready."
  • "Don't regenerate your song. Save it." (the WAR Room favorite)

The 12-18 build sequence (3 phases):

Phase A — Foundation (4-5 builds)

  • A1 (B3154) — `judgeRevisionDelta` Haiku primitive shipped. src/lib/claude/judge-revision-delta.ts + 18 tests in judge-revision-delta.test.ts. Returns { verdict: 'improves' | 'changes' | 'risks', dimensions: { specificity, singability, memorability, vulnerability } on -2/+2 scale, rationale, judgeError? }. Fail-open contract: when the Haiku call fails (network / parse / timeout), returns verdict='changes' + zero dimensions + judgeError set; caller treats judgeError as "unjudged." Short-circuits the Haiku call on identical-input and deletion cases. Includes aggregateSessionDeltas() helper for end-of-session per-axis totals + net composite. Uses the same getAnthropicClient() + getModel('focus-group') (Haiku) + temperature: 0 pattern as the existing compliance-check + witness-details primitives. The four dimensions map intentionally to the load-bearing Release Dossier axes (specificity → Lyric Quality, singability → Singability, memorability → Hook Clarity, vulnerability → SA#23 AID rule).
  • A2 (B3155) — `SurgerySession` data model + Supabase table shipped. Two-table schema in supabase/migrations/create_song_surgery_sessions.sql: song_surgery_sessions (status state-machine + frozen diagnostic snapshot + verification record + final_lyrics + branched_from_session_id for alternate-path branching) + song_surgery_edits (per-edit log with verdict + dimensionDeltas + operator status). Full RLS — operators see only their own sessions. updated_at trigger so timeline UI shows accurate last-touched timestamps. TypeScript types + state-machine helpers in src/lib/surgery/session-types.ts: SurgeryStatus union (6 values matching SQL CHECK), SurgeryEditStatus union (4 values), SurgeryDiagnosticSnapshot, SurgeryVerificationRecord, SurgerySession, SurgeryEdit. canTransition() + isTerminal() enforce the state-machine contract (diagnosing→planning→repairing→verifying→complete, any non-terminal→abandoned). buildEditRecord() converts a judgeRevisionDelta judgment into a DB-row-shaped record; applyEditToLyrics() does 1-based line-number replacement preserving all other lines. 17 tests pin the state machine + the edit-record builder + the lyric-apply helper.
  • A3 (B3156) — Surgery diagnose builder shipped. src/lib/surgery/diagnose.ts provides buildDiagnosticSnapshot(song) reading three engines: runAllReportQaGates (B3143) + runTasteSensitivityScan (B3138) + computeLineLevelSurgery (B3147). Returns { hasFindings, snapshot, cleanMessage }. The load-bearing restraint contract enforced here: when ALL three signals come up clean, hasFindings=false and cleanMessage carries the operator-facing copy ("This song doesn't need surgery. Composite X/100, Release Readiness Y/100. Ship it.") instead of an empty Surgery Plan. Three flavors of clean message: high-scoring → "Ship it"; mid-scoring with no flags → "Surgery only triages flagged issues"; unscored → "Score first, then return." prioritizeSurgeryItems() helper enforces the severity-ladder + line-number sort contract callers can rely on. 10 tests pin the restraint contract + the snapshot shape + the prioritization order. The HTTP route layer comes in a forthcoming build — this ships the pure-compute orchestrator that the route will call.
  • A4 (B3157) — Lyric-diff UI primitive shipped. src/components/surgery/LyricDiff.tsx. Two render variants: 'card' (vertical, for the per-issue repair surface) + 'inline' (compact one-row, for the timeline). Per-verdict color treatment (improves=green / changes=yellow / risks=red / unjudged=neutral). Shows: Before (strike-through italic), After (bold), verdict chip (top-right), per-axis dimension delta chips (when non-zero), one-line rationale (italic, dimmed). Reusable across the future workflow page (B1) + repair card (B3) + session timeline (B4) + Final Vitals screen (B5).
  • A5 (B3158) — Verified-rescore trigger + Final Vitals math + trust verdict shipped. src/lib/surgery/verify.ts. Three functions: shouldVerifyNow() fires the verified rescore when acceptedSinceLastVerify >= 3 OR explicit operator verify OR session closing. buildVerificationRecord() composes the diagnostic snapshot baseline + the session's accepted-edit judgments + the verified post-rescore scores into the SurgeryVerificationRecord (per-axis aggregate dimensions + net composite delta + net readiness delta + verifiedAt). computeTrustVerdict() produces operator-facing copy comparing projected (per-axis sum) vs. verified (composite delta): clean improvement, projection variance flag ("Projected +8; verified +2 — review which axes the edits actually improved"), negative-direction movement honesty ("wrong direction — review the edits"), no-change reporting. VARIANCE_THRESHOLD = 4 (conservative: 2-point gap is noise, 4-point is signal). 14 tests pin trigger rules + Final Vitals math + every trust-verdict path. Phase A CLOSED — A1+A2+A3+A4+A5 all shipped. 41 tests green across all 4 Phase A modules.

Phase B — Workflow UX (CLOSED; shipped 31 builds B3159-B3203)

Build 3272 re-scope (Deep Audit Tier 3 #21). Original header said "5-6 builds." Actual: B1-B6 shipped the core workflow UX (B3159-B3163) — the 5-6 estimate was honest for that scope. Then B21-B31 (the AUDIT-1 through AUDIT-12 follow-on items, B3172-B3203) piggybacked under the same Phase B header rather than opening Phase B2. The 2026-05-24 Deep Audit Section 5 agent flagged this as artifact-drift: the header stopped tracking ground truth, which eroded the planning artifact's load-bearing function. All 31 items are done. This re-scope is a honest-relabel (not a re-open). Future Song-Surgery work opens Phase B2 with a fresh header that matches its scope from day one. The B21-B31 wave's lesson: when a follow-on item set exceeds the original scope by ~3x, open a new phase rather than appending under the old one. The B2435 "one inferred fix per build, max" rule has a sibling here: "one phase per scope estimate, max."
  • B1 (B3159) — `/surgery/[songId]` page shipped. Server-component shell at src/app/surgery/[songId]/page.tsx. Auth via createServerSupabase() + owner-check on the songs row; unauth → sign-in CTA; unconfigured-supabase → graceful fallback page; non-owned song → notFound(). Composes PremiumReportData from the songs row JSONB columns (eval_data / first_listen / prosody_report / focus_group / forge_settings) and runs buildDiagnosticSnapshot(song) (B3156). Two branches enforce the load-bearing restraint contract: when !hasFindings, renders ONLY the green CheckCircle banner with diagnosis.cleanMessage + "Song Surgery only triages flagged issues — restraint is the trust signal" footer (no plan, no list, no polish suggestions). When findings exist, renders the Surgery Plan: counts row (critical/major/minor), QA gate failures section, ordered surgery items with severity color borders (red/orange/yellow), and a "Begin Surgery" footer noting the interactive flow ships in B3. data-testid="surgery-clean-message" pins the restraint branch for e2e tests. Force-dynamic + noindex metadata.
  • B2 (folded into B3159) — Diagnosis screen shipped as the read-only branch of the `/surgery/[songId]` page. The page renders the prioritized counts row (critical / major / minor), the QA-gate failures section, and the ordered Surgery Plan as part of the same shell. No separate component — buildDiagnosticSnapshot output rendered inline. The "Begin Surgery" placeholder was retired in B3160 when the RepairFlow took over the items list.
  • B3 (B3160) — Per-issue repair card + RepairFlow shipped. src/components/surgery/RepairFlow.tsx is the client-component interaction primitive. Single-issue card: severity left-border (red/orange/yellow), Problem header, Current line (italic), Why it matters, Suggested revision (with inline Edit button → editable textarea, max 500 chars, font-serif lyric typography), and four actions: Accept (green; calls /api/surgery/judge-edit, applies the possibly-edited revision, shows loading spinner during judgment), Reject suggestion (the fix is wrong), Skip line (leave the line alone), Continue to next issue (after judgment renders). On accept, the LyricDiff component (B3157) renders inline showing Before/After + verdict chip (improves / changes / risks / unjudged) + per-axis dimension delta chips + the one-line rationale. Session summary view fires when all items exhausted: 4-tile counts grid (Accepted / Skipped / Rejected / Improves) + verdict-timeline of all accepted edits as inline LyricDiffs + footer noting verified rescore + Final Vitals ship in B5. The companion API route shipped alongside: src/app/api/surgery/judge-edit/route.ts — POST, auth-gated, CSRF-validated, 120/hour rate limited via the new surgeryJudge bucket, 500-char per-line cap. Wires through to judgeRevisionDelta (B3154) with { oldLine, newLine, genre? }. Page now ends with the RepairFlow instead of the placeholder block; the diagnose-clean branch is unchanged (restraint contract preserved).
  • B4 (B3161) — Session timeline component shipped. src/components/surgery/SessionTimeline.tsx is the reusable version-history surface. Pure presentation (no state, no fetch — parent owns the data + revert callback). Takes versions: TimelineVersion[] (each carrying { label, createdAt, caption?, edits, verifiedComposite?, verified? }) + optional currentLabel + optional onRevert(label) callback + compact mode for the dashboard strip. Renders a vertical stack with a left spine line + bullet dots (forge-gold ring on the current version, dimmed circle on the rest). Each version card shows: label + caption + Current/Verified chips + edit count + per-verdict-color count chips (improves/changes/risks/unjudged) + verified composite when present + per-edit LyricDiff (inline variant, hidden in compact mode) + Revert button when onRevert is passed AND the version is not current. Companion bumpVersionLabel('1.x.y') → '1.(x+1).0' semver helper (Song Surgery sessions bump minor, roll patch to 0). The RepairFlow session-complete view now uses SessionTimeline (via the buildTimelineVersions(accepted) helper that yields v1.0.0 baseline + v1.1.0 post-session) instead of the ad-hoc accepted-edits list. 8 tests pin the bump-version semantics (undefined → 1.0.0, unparseable → 1.0.0, patch-rollover, major-preservation).
  • B31 (B3203) — Operator-reported Genre Radio gap closeout: auto-audit on forge-finalize. Operator created a Latin song, uploaded audio, shared it — and it didn't appear on the Latin Genre Radio. Root cause: /api/genres/[slug]/radio requires an entry in song_audit_runs at audit_version='1.2.0' with the matching detected_arc_slug. That table was populated ONLY by the manual scripts/backfill-song-audits.ts script. Fresh forges never appeared until the operator re-ran the backfill. Fix in `src/app/api/songs/forge/finalize-song.ts`: inline auto-audit IIFE inside the DB-update success branch. Calls auditSong({id, lyrics, genre, sunoStyleString}) (pure-compute; no LLM), inserts the resulting SongAuditRunDraft into song_audit_runs via service-role client (RLS on the table is admin-only). Fire-and-forget; load-bearing forge path unaffected if the audit or insert fails. SF_FORGE_AUTO_AUDIT_DISABLED=1 env toggle is the emergency rollback. Structured-log events: forge.auto_audit.{disabled_via_env,skipped_no_audit,no_service_role,insert_failed,persisted,threw}. Every new forge now appears on its arc's Genre Radio within 2 minutes (the s-maxage=120 cache TTL). Pre-existing songs that never got audited still need the operator's one-time backfill: npx tsx scripts/backfill-song-audits.ts --admin-only. tsc clean.
  • B30 (B3202) — AUDIT-8 follow-up: first-forge email trigger in finalize-song. Fires when the user's complete-song count BEFORE the just-finalized one is 0 (= this is their first finished forge). Fire-and-forget; gated by SF_EMAIL_RETENTION_ENABLED inside the template. Includes the song's composite score in the email body when present. Structured-log events on count failure / no-email / dispatch / catch.
  • B29 (B3201) — AUDIT-8 follow-up: welcome email trigger on onboarding completion. New POST /api/onboarding/complete endpoint fires sendWelcomeEmail fire-and-forget. Client posts to it from src/app/onboarding/page.tsx after updateProfile() resolves. Idempotent in practice (onboarding form redirects away once username is set). Gated by SF_EMAIL_RETENTION_ENABLED — until flipped, every trigger no-ops with retention.welcome.gate_closed logs the operator can grep to verify wiring.
  • B28 (B3200) — Surgery WAR ROOM P1 #9 closeout: rate-limit split into 3 lanes. Pre-3200 all 6 surgery endpoints shared one surgeryJudge bucket (120/hr). A real session = ~1 init + 5-10 per-edit writes + 5-10 judge calls + 0-3 suggest-rewrites + 1 finalize = 12-25 calls EACH. 4-6 sessions/hour locked operators out mid-session AND silently shred the per-edit persistence path (fail-open → produces the "Session record: Not saved" surprise that B3198 now surfaces honestly via the failed-writes banner). Fix in `src/lib/api-auth.ts`: split into 3 RATE_LIMITS entries: surgerySession (30/hr; init + finalize + PATCH/GET session — once-per-session-lifecycle), surgeryEdit (600/hr; per-edit append — load-bearing persistence path, generous for burst-handling), surgeryJudge (200/hr; judge-edit + suggest-rewrites — Haiku calls, bumped from 120). Each endpoint routed to its matching lane: session/init → surgerySession; finalize → surgerySession; session/[id] (GET + PATCH) → surgerySession; session/[id]/edit → surgeryEdit; judge-edit → surgeryJudge (unchanged); suggest-rewrites → surgeryJudge (unchanged). Each routing change carries a Build 3200 comment naming the rationale per endpoint. tsc clean.
  • B27 (B3199) — Surgery WAR ROOM P0 #3 closeout: canTransition() enforcement on finalize UPDATE. Pre-3199 the finalize route's UPDATE-mode at route.ts:393-402 was update({status:'complete'}) UNCONDITIONALLY. If the session row was already in a terminal state ('complete' from an idempotency replay, or 'abandoned' because the operator clicked "Start fresh" in another tab), the route silently re-wrote the verification_record + final_lyrics OVER the terminal state — a state-machine bypass that the canTransition() helper in session-types.ts was supposed to prevent (but no caller enforced). Two-tab race + idempotency-replay both produced silent corruption. Fix in `src/app/api/surgery/finalize/route.ts`: imported canTransition + SurgeryStatus at module level. UPDATE-mode now SELECTs the current status first (with .eq('user_id', auth.userId) for RLS-equivalent owner gate), runs canTransition(currentStatus, 'complete'), and only proceeds when legal. On illegal-transition, logs surgery.finalize.illegal_transition (warn, with fromStatus/toStatus) + nulls sessionId so the UI's "Saved · X" chip surfaces "Not saved" honestly — the operator sees something didn't land. Success-log now carries fromStatus so the audit trail shows the transition shape. SELECT-then-UPDATE has a TOCTOU race window (operator could abandon in another tab between SELECT and UPDATE) but it's at the audit/failure-mode tier, not the "single tab happy path" tier; the in-code comment names a future hardening pass that moves the check into a Postgres CHECK constraint on status transitions. 17/17 session-types tests still green (canTransition test surface unchanged); tsc clean.
  • B26 (B3198) — Surgery WAR ROOM P1 #10 closeout: appendLiveEdit failure surfacing. The audit's #2 top-leverage fix. Pre-3198, appendLiveEdit fires across 5 call sites (onAccept / onSkip / onReject / onUndoJudgmentAndSkip + resume hydration) as fire-and-forget — failures only logged a console.warn. Operators walked through the whole session unaware their edits weren't persisting; "Session record: Not saved" surprised them 5+ minutes later at session-close, by which point the audit trail was already corrupted (UPDATE-mode finalize writes verification_record + final_lyrics but does NOT bulk-insert the missing edits). Fix in `src/components/surgery/RepairFlow.tsx`: new failedEditIndices state queue captures every failed write (network error / HTTP non-2xx / session-init failed); new retryingFailedEdits boolean tracks the retry-in-flight state. appendLiveEdit now queues the payload on every failure path (including the session-init-failed case where we don't have an editIndex yet — queued with placeholder -1, retry assigns a real index on re-fire). New retryFailedEdits() function walks the queue, re-fires each write, keeps failures in the queue + drops successes. Sticky banner UI renders ABOVE the resume banner + progress strip whenever the queue is non-empty: yellow border, AlertCircle icon, copy distinguishes 1 vs N edits, explains the contract honestly ("in-memory session is intact, contributes to Final Vitals, but won't be resumable from dashboard until X succeeds"), single primary "Retry failed writes" button (RotateCcw icon, loading state with Loader2 spin). role="status" aria-live="polite" so screen readers announce new failures. 93/93 surgery tests green; tsc clean. Closes the "Session record: Not saved" surprise — operators see failures at the moment they happen, not after a 10-edit walk.
  • B25 (B3196) — Song Surgery WAR ROOM audit + P0 #6 fix shipped (rescorePending stuck-state). Operator screenshot showed a completed surgery session with the Release Clearance chip stuck on "VERIFYING" + the Verified pane chip stuck on "Rescoring" (spinning Loader2). Root-cause: the server-side /api/surgery/finalize route at route.ts:339 sets rescorePending = verified === null — meaning "verified composite is null" (TERMINAL failure / no-edits-to-verify / env-disabled state). Pre-3196 FinalVitals.tsx lines 110-117 + 215-220 treated rescorePending=true as in-progress — rendered spinning loaders + "Verifying" + "Rescoring" copy on a finished session. The route had already returned; the rescore was DONE; the UI lied about that. Operator's audit ask: WAR ROOM deep audit of all Song Surgery code. Dispatched a 100-expert-style Agent against ~5,400 LOC; produced 25 findings across 4 tiers documented in docs/SURGERY-WAR-ROOM-AUDIT-B3196.md. Highlights: P0 #1 (clearance chip falls to "No net change" when verified ≠ null but startingComposite = null); P0 #2 (resume hydration silently rewrites 'unjudged' verdicts to 'changes' + fabricates zero dimensions); P0 #3 (finalize UPDATE-mode skips canTransition() — two-tab race produces silent state-machine corruption); P0 #4 (liveEditIndex closure-capture race produces duplicate edit_index rows); P0 #5 (resume currentIdx advance is wrong when surgeryItemId is null — puts operator back at item 0); P1 #10 (the appendLiveEdit fire-and-forget pattern is the load-bearing weak link producing the "Session record: Not saved" surprise — silent data loss). Also: ZERO test coverage on RepairFlow.tsx (1802 LOC), FinalVitals.tsx (334 LOC), LyricDiff.tsx, AND all 6 API routes. Architecture observations: state machine split across 3 layers with no single owner; 3 persistence modes routed by implicit client state; fail-open signals inconsistent across 5 fetch sites. Top 5 fixes by leverage identified for the next 6-8 builds. Ship in B3196: collapsed the FinalVitals.tsx clearance ladder + Verified-pane chip into the !verifiedAvailable branch with "Projection only" copy + Sparkles icon + a tooltip explaining the terminal cause; removed the Loader2 import (now unused). Long inline comment names the rename for a future build + cites SURGERY-WAR-ROOM-AUDIT-B3196.md. NEW src/components/surgery/FinalVitals.behavior.test.tsx — 10 tests pin the B3196 fix: "Projection only" chip renders when rescorePending=true (with no spinner / no "Rescoring" copy / no "Verifying" copy); same chip when rescorePending=false; no .animate-spin elements at all; clearance chip is not "Verifying"; all 4 verified-available chip states render correctly; dimension scoreboard always shows the 4 axes + net delta. First test coverage for FinalVitals + first surgery-component behavior test ever. 10/10 green; tsc clean.
  • B5 (B3162) — Final Vitals + Release Clearance screen shipped. src/components/surgery/FinalVitals.tsx is the end-of-session presentation component. Reads a SurgeryVerificationRecord (B3155 type) + VerificationTrustVerdict (B3158 type) + the baseline composite + readiness scores, then renders three scoreboards: (1) Projected — per-axis dimension sum (specificity / singability / memorability / vulnerability) with up/down/neutral trend icons + net delta header; (2) Verified — composite + Release Readiness paired panes showing start → end → delta, with "Pending / Rescoring" chip when verifiedComposite is null; (3) Trust verdict callout — full operator-facing message from computeTrustVerdict() with yellow-bordered variance treatment when projected >> verified; (4) Release Clearance chip computing the structural state (Cleared / Mixed result / Wrong direction / No net change / Verifying / Projection only). Companion API endpoint at src/app/api/surgery/finalize/route.ts — POST, auth-gated, CSRF-validated, rate-limited via the surgeryJudge bucket. Composes the verification record from accepted-judgment dimension deltas using buildVerificationRecord() + computeTrustVerdict(). Projection-only path: verifiedComposite + verifiedReadiness ship as null; the actual verified rescore (re-evaluating the rewritten lyrics via the eval engine — ~$0.10/call) wires through in a follow-up build. The Final Vitals screen handles the null case cleanly via the rescorePending flag. RepairFlow now: (a) auto-fires the finalize call when the session reaches completion via useEffect; (b) renders FinalVitals + SessionTimeline + summary tiles in the completed view; (c) shows loading + error states cleanly. Phase B B1+B2+B3+B4+B5 all shipped — only B6 (revert + branching) remains.
  • B24 (B3193) — Critical bug fix: Final Vitals auto-abort race. Operator-reported: spinner stuck at 196s+ with no Retry surfacing. B3190 thought it added a 75s client-side timeout. It DID add one, but never actually got to fire because of a useEffect dep-array race that's been latent since B3170. Trace: the finalize useEffect had finalizing in its dep array. The effect body calls setFinalizing(true) near the top. React processes the state update → re-render → useEffect deps compared → finalizing changed false → true → CLEANUP function of the just-started run fires → cleanup calls controller.abort() (B3190 addition) + clearTimeout(abortTimer) (B3190) + cancelled = true (B3170). The fetch then aborts within microseconds of starting. .catch sees AbortError but cancelled=true returns early. .finally checks !cancelled → false → setFinalizing(false) never runs. State is stuck: finalizing=true forever, spinner spins forever, the 75s timer was killed in the same cycle, no path back. Operator saw 196s+ elapsed with no Retry button (because finalizeError was never set — the .catch returned early). Fix in `src/components/surgery/RepairFlow.tsx`: remove finalizing from the useEffect dep array (with eslint-disable-next-line react-hooks/exhaustive-deps + a long-comment explaining why). The guard inside still reads finalizing via closure — which is fine because the closure captures the value at effect-scheduling time, and the only way the effect re-fires legitimately is when a different dep changes (e.g. operator clicks Retry → finalizeError changes from non-null → null), at which point finalizing SHOULD be false anyway (.finally set it). Now setFinalizing(true) doesn't trigger a re-fire, the cleanup doesn't run mid-fetch, the abort timer survives to fire at 75s, and the catch path properly surfaces the timeout error + Retry button. 83/83 surgery tests green; tsc clean. The pre-B3190 code had the same latent race but was harmless because the cleanup only set cancelled=true without calling controller.abort(); the fetch ran to completion (eventually) and .finally's !cancelled guard suppressed the state update. Bug surfaced when B3190 added the abort call — operators with slow finalize calls would have hit it. Sacred Accident candidate: "An effect's deps array should never include a state variable the effect mutates." Future builds should hold this line.
  • B23 (B3192) — Operator-requested: post-judgment back-out paths in Song Surgery. Operator hit the case where they picked rewrite #1 (an R&B chorus line), got the judgment ("RISKS" with -1/-1/-2 dimension deltas + a rationale explaining the loss of emotional choreography), and had no way to undo. The only post-judgment action was "Continue to next issue," which commits the un-improved edit. They asked for "go BACK and change to another option or SKIP." Fix in `src/components/surgery/RepairFlow.tsx`: two new client-side functions + a refreshed post-judgment button row. onUndoJudgmentToEdit() pops the last in-memory accepted entry, clears lastJudgment (returns the card to its action-row state), and flips isEditing=true so the textarea reopens populated with what they tried — operator can pick a different ranked suggestion (still rendered above) or edit further. onUndoJudgmentAndSkip() pops the accept + records a skip + fires appendLiveEdit with status='skipped' + advances. The post-judgment button row is now a flex justify-between with two back-out buttons on the left (Try different revision — RotateCcw icon + dark-800 styling; Skip this line instead — SkipForward icon + transparent styling) + the existing primary "Continue to next issue" forge-500 button on the right. Visual hierarchy keeps the happy-path primary dominant; back-out is available but de-emphasized. DB persistence trade-off documented in-code: the B3185 per-edit endpoint already wrote the accept row when the judge call returned, and we leave that row in place — the session timeline + activity feed accumulate an honest "tried + reverted" history. The verification record composed at finalize() reads the in-memory accepted array, so the metrics reflect only the operator's final decision, not their first try. A future build could add a DELETE endpoint that cleans the superseded edit row; the load-bearing UX (operator can undo) ships here without it. 83/83 surgery tests green; tsc clean.
  • B22 (B3191) — Operator-reported follow-up: forge-result autoscroll didn't go "ALL the way up." B3188 shipped autoscroll via resultRootRef.scrollIntoView({ block: 'start' }) — aligns the result root's TOP with the viewport top. Operator reported they were still landing partway down on the analysis sections instead of at the title + score + lyrics. Likely causes: (a) lazy-mounting post-result cards (SamplePlayer, cover art, wound summary, heat density) shift the page mid-animation and the single 50ms scroll loses the race; (b) block: 'start' on a nested element can leave the page short of true top when there's a wrapper above it. Fix in `src/app/forge/v2-design/ForgeV2Result.tsx`: switched from scrollIntoView to window.scrollTo({ top: 0, left: 0, behavior: ... }) — unambiguous "go to the very top of the page" — plus a safety re-fire at 350ms to catch layout shifts that happen after the first scroll lands (~300ms after the smooth-scroll completes). Smooth by default; 'auto' (instant) under prefers-reduced-motion. Try/catch fallback to window.scrollTo(0, 0) for older browsers without the options-bag API. Removed the now-unused resultRootRef ref + its attachment on the outer <div>. SSR-safe via typeof window === 'undefined' guard. 11/11 ForgeV2Result behavior tests green; tsc clean.
  • B21 (B3190) — Operator-reported: "Composing Final Vitals…" hang fix. Operator reported the spinner circling for several minutes with no progress indication after completing a 7-item surgery session. Root-causing: the server-side /api/surgery/finalize route caps at maxDuration=60 with a 30s internal verified-rescore timeout — so worst-case it should return in ~60s. But the client-side fetch had NO AbortController + NO wall-clock cap, so if the Vercel lambda hard-killed without flushing a response (rescore + DB writes pushed close to the 60s ceiling), the fetch hung indefinitely with the spinner spinning. Three coordinated fixes: (1) src/lib/surgery/verified-rescore.ts: lowered the default internal timeout from 30s → 20s. Server-side margin under the 60s lambda ceiling jumps from ~30s to ~40s for DB writes + response — dramatically reduces the chance of a hard-kill mid-response. Trade-off documented: more sessions fall back to projection-only Final Vitals when the eval is on the slow end, but that's an honest fail-open vs. a hang. Test updated. (2) src/components/surgery/RepairFlow.tsx: client-side AbortController + 75s wall-clock cap on the finalize fetch (= 60s server ceiling + 15s grace). On timeout: AbortError caught in .catch, finalizeTimedOut state flips true, error message reads "Final Vitals took longer than 75 seconds to compose. Your edits are already saved — click Retry to try again, or refresh and resume the session from your dashboard." Cleanup function clears the timer + aborts the controller on unmount. (3) Progressive status messages: new finalizeElapsed state + companion useEffect ticks once per second while finalizing. UI swaps text at thresholds: 0-20s → "Composing Final Vitals…"; 20-45s → "Still working — verified rescore in progress…"; 45+s → "Taking longer than usual — finishing up…". Sub-line shows "Ns elapsed" + ", the eval engine is re-scoring your rewritten lyrics" after 20s. Operator can now distinguish "alive but slow" from "hung." (4) Retry button in the error block: <RotateCcw /> icon + "Try again" label; click clears finalizeError + finalizeTimedOut, which (via the useEffect dep array) re-fires the finalize fetch. Per-edit persistence (B3185) guarantees the operator's edits are already in the DB; retry only re-runs the verification composition + verified rescore. UPDATE-mode in the finalize route handles re-finalizing a session row that's already 'complete' idempotently. 83/83 surgery tests green; tsc clean.
  • B20 (B3189) — Operator-requested: 3 ranked rewrite suggestions per surgery item. Operator feedback after B3179: "The Suggested Revision is direction, but is this what goes into the song? Actual line revisions need to be suggested — would be good if 3 suggestions were made in ranked order." NEW src/lib/claude/suggest-line-rewrites.ts — Haiku-driven rewrite synthesizer. Input: oldLine + problem + whyItMatters + guidance + optional section/genre/centralTension. Output: 3 ranked singable rewrites with one-line craft rationale each (rank 1 = strongest = "the rewrite a working songwriter would actually use"; rank 2 = viable alternative with a different angle; rank 3 = interesting departure — more poetic / concrete / direct than 1+2). Strict JSON contract; clamps rank to 1..3; sorts ascending; drops empty + over-500-char lines; caps at 3 rewrites total. 15s wall-clock timeout via Promise.race against TIMEOUT_SENTINEL — long Haiku tails fall back to empty-rewrites + the operator writes their own. Fail-open at every error path. NEW /api/surgery/suggest-rewrites/route.ts — POST endpoint wrapping the helper. Auth + CSRF + 120/hr rate-limit via surgeryJudge bucket + 500-char oldLine cap + maxDuration=30. RepairFlow.tsx updated: new state (suggestions / suggestionsLoading / suggestionsError) cleared on advance() so each card opts in independently; new onSuggestRewrites() lazy-fetches the 3 alternatives; new onUseSuggestion(line) fills the textarea + opens edit mode. UI inside the existing yellow-bordered "Repair guidance" block: "Suggest 3 rewrites" button (forge-styled, Sparkles icon) → loading spinner with "Asking war room…" copy → 3 numbered cards stacked, rank-1 card gets forge-500 border + bg + "top pick" sub-label, ranks 2-3 get neutral dark-700 styling. Each card shows the line in serif + the one-line rationale beneath in dimmed italic. Clicking ANY card fills the textarea + opens edit mode so the operator can tweak before Accept. Cost envelope: ~$0.0001/click, lazy-fetched per card, ~$0.0003 for a heavy 3-issue session that uses the feature on every card. 17 new tests on the helper (prompt composition + every parse edge case: chatty prefix tolerance, rank clamping, sort order, empty-line drop, 500-char cap, 3-item cap, null returns on garbage / missing field / non-array / unparseable JSON / missing rationale). 100 tests green across surgery + claude. tsc clean.
  • B19 (B3188) — Forge result auto-scroll-to-top shipped (operator-reported UX bug). When a song finished forging, the result component mounted inline below the composing/forging surface — operators who'd scrolled down to read forging-phase commentary landed on the Wound Summary or Rhythm panel or even the "About the Forge" marketing block at the bottom, not the headline result card. Fix in src/app/forge/v2-design/ForgeV2Result.tsx: new resultRootRef attached to the result root <div> + useEffect keyed on forgeResult.id that fires scrollIntoView({ behavior: 'smooth', block: 'start' }) after a 50ms delay (lets the celebration fireworks paint cycle settle so the scroll target stays stable). Honors prefers-reduced-motion — switches to behavior: 'auto' (instant) under that media query so motion-sensitive operators don't get a jarring slide. Fallback scrollIntoView() (no options bag) for older browsers. Dep array includes forgeResult.id so "Forge again" → new song id → fresh scroll. SSR-safe via typeof window === 'undefined' guard. Tsc clean.
  • B18 (B3186+B3187) — Song Surgery session resume shipped. Closes the loop on the per-edit persistence work (B3184+B3185) — operators who crash mid-session can now actually CONTINUE from where they left off. B3186 server: NEW /api/surgery/session/[id]/route.ts ships GET (returns the session row + all edit rows in order, RLS-gated) and PATCH (transitions status='repairing' → 'abandoned' for "Start fresh"; uses canTransition() from session-types to enforce the state-machine contract; rejects any other status). Page-level detection added to /surgery/[songId]/page.tsx: queries song_surgery_sessions for the operator's most-recent status='repairing' row for THIS song, passes its {id, createdAt, updatedAt} summary down as resumableSession prop to RepairFlow. B3187 client wire: RepairFlow accepts the new resumableSession? prop + adds 3 new state vars (resumeChoice: null | 'resumed' | 'fresh' + resumeBusy + resumeError). New onResume() handler fetches the saved edits, hydrates accepted + skipped arrays from the rows (mapping verdict + dimensionDeltas + rationale faithfully; unjudged status preserved via judgeError sentinel), advances currentIdx past every item that already has a recorded edit (matched by surgery_item_id), and sets liveSessionId so subsequent edits append to the same session row. New onStartFresh() handler fires the PATCH to abandon the old session + flips local state. New <History /> icon resume banner renders ABOVE the progress strip when resumeChoice === null AND resumableSession exists. Shows last-touched timestamp + two CTAs ("Resume previous session" — forge-500; "Start fresh" — dark-800 secondary). Resume errors surface inline with <AlertCircle />. 142/142 tests green; tsc clean. The full crash-recovery loop is now closed end-to-end: write (B3184+B3185) → detect (B3186 page query) → offer (B3187 banner) → hydrate (B3187 onResume) → continue (existing flow). Six builds total (B3170+B3184+B3185+B3186+B3187) to take Song Surgery from "session lost on tab close" to "session resumable from any device."
  • B17 (B3185) — Per-edit Song Surgery persistence: client wire shipped. Tab-crash data-loss gap CLOSED end-to-end. RepairFlow.tsx gains: (a) two new state vars — liveSessionId: string | null (the row id from the lazy-create call) + liveEditIndex: number (1-based, increments per write); (b) two new async helpers — ensureLiveSession() which fires POST /api/surgery/session/init once per session lifecycle + caches the result, and appendLiveEdit(payload) which fires POST /api/surgery/session/[id]/edit for every action. Both helpers are fail-open: on network/server failure they console.warn + return null/void; the in-memory state continues to work. The three action handlers (onSkip / onReject / onAccept) now call void appendLiveEdit(...) immediately after updating the in-memory accumulator. The finalize POST body now passes sessionId: liveSessionId, triggering the B3184 UPDATE-mode path (no double-insert; finalize just transitions status='complete' + writes verification_record + final_lyrics). onRevertToBaseline clears liveSessionId + liveEditIndex so each branch attempt creates a fresh session row (preserves the branching audit trail). useEffect dep array gains liveSessionId so a late-arriving session id (slow ensureLiveSession) doesn't get missed by the finalize trigger. The "tab-crash mid-session = total data loss" gap is now closed. Tab-crash loses at most ONE edit (the mid-flight one), not the whole session. 142/142 tests green; tsc clean. Per-edit persistence is the last load-bearing Song Surgery follow-up I flagged when declaring the feature shipped — both gaps now closed (B3170+B3184+B3185 for persistence; B3171 for verified rescore).
  • B16 (B3184) — Per-edit Song Surgery persistence: server endpoints shipped. Closes half of the "tab-crash mid-session = total data loss" gap operator-flagged when declaring Song Surgery shipped. NEW endpoint /api/surgery/session/init/route.ts lazily creates a song_surgery_sessions row with status='repairing' on the operator's first action (Accept / Skip / Reject). Authenticated session client; RLS gates per-user; returns sessionId. Body: { songId, diagnosticSnapshot, branchedFromSessionId? }. NEW endpoint /api/surgery/session/[id]/edit/route.ts appends ONE row to song_surgery_edits per call. Validates editIndex / lineNumber / status (accepted/rejected/skipped/proposed) / verdict / dimensionDeltas / 500-char per-line cap. Session ownership enforced via the existing own_session_edits_insert RLS policy. MODIFIED /api/surgery/finalize/route.ts — now branches on whether body.sessionId is present. UPDATE mode (B3184): existing session row gets status='complete' + verification_record + final_lyrics set via .update().eq('id', sessionId); edit rows skipped (they're already there from the per-edit endpoint). LEGACY INSERT mode (B3170 backwards compat): no sessionId in body → creates session + bulk-inserts edits in one shot as before. Both modes emit the surgery.session.completed audit event for the dashboard activity feed (B3178 + B3182 wired); UPDATE mode passes null for accepted/skipped counts because the per-edit endpoint wrote them one at a time and they're not aggregated at the finalize layer. The B3185 client wire that calls these endpoints ships in the next build.
  • B15 (B3183) — Inverse-mode UI for VaultInspirationPanel shipped (consumer wire for B3168 primitive). src/components/forge/VaultInspirationPanel.tsx gains a per-card "Like / Avoid" toggle button next to the existing arc + score + similarity chips. Two new states: invertedIds: Set<string> (mutual-compatible with the existing excludedIds — a card can be inverted AND included, or excluded entirely) + reset alongside excludedIds on every successful fetch (operator opts in per-card, doesn't carry across prompts). The onReferencesChange mapping now sets inverse: invertedIds.has(ex.id) per ref so the parent's prompt-prefix builder routes each card to either the EXEMPLAR block or the ANTI-PATTERN block per the B3168 primitive. Visual treatment: when a card is inverted, its border + background flips from the amber exemplar styling to a yellow-bordered anti-pattern treatment (matches the B3168 VAULT ANTI-PATTERNS block accent). Toggle button states: "Like" (amber, default) ↔ "Avoid" (yellow). Disabled when the master useAsReferences toggle is off OR the card is excluded (flipping under those conditions has no downstream effect). Accessible labels + title tooltips explain both states. The Sister-song panel pattern from the punch list spec realized — operator can now say "depart from THIS voice" in addition to "honor THIS voice." Inverse-mode primitive (B3168) ships its consumer wire here; the discovery → composition → UX loop for inverse references is closed end-to-end.
  • B14 (B3182) — Dashboard activity-feed consumer for surgery.session.completed shipped (closes the B3178 write-without-renderer gap). src/app/dashboard/activity/derive-events.ts gains a 'surgery_session' AuditEventKind + a case in mapAuditRowToEvent for 'surgery.session.completed'. The mapping handles the surgery audit-event's unusual shape (subject_id = SESSION id, not song id) by pulling payload.songId for the parent-song back-link + titleBySongId.get(parentSongId) for the song title. Detail copy adapts to data presence: "3 edits accepted · verified +5 composite" when the rescore landed, "3 edits accepted · projected +7 per-axis" when not, "3 edits accepted" when no deltas, with correct singular grammar on 1 edit accepted. src/app/dashboard/activity/ActivityFeed.tsx KIND_META gains a surgery_session entry: <Scissors /> icon + "Song Surgery" label + tonal-forge chip variant (matches the Song Surgery brand surfaces across /surgery, RepairCard, dashboard CTA banner). 6 new tests pin the mapping cases (kind resolution, payload-songId routing, verified-delta detail, negative-direction sign, projected-fallback, singular grammar). 46/46 activity-feed tests green; tsc clean. The dashboard activity feed now renders Song Surgery sessions alongside forge/score/gauntlet/share/etc events — the discovery + audit loop is closed end-to-end.
  • B13 (B3180+B3181) — Two operator-reported bugs fixed. B3180 — Story Twist firing on every song: detectGenreMode() in src/lib/claude/genre-modes.ts used String.includes(keyword) to test trigger keywords. Story Twist's trigger list included 'turn' — a substring match on which fires Story Twist for "Saturn", "return", "nocturnal", "turning point", "you turn me on" — and any other word containing the letters t-u-r-n. Plus 'narrative' + 'story' are extremely common English words that triggered indiscriminately. Plus Story Twist is registered HIGH in the priority order (line 1082 comment confirms it). Net effect: most songs got tagged Story Twist regardless of their actual narrative intent, and the mode chip on the dashboard reflected that misdetection. Fix: switched the matcher from haystack.includes(keyword) to a word-boundary regex (\b{escaped-keyword}\b, case-insensitive). Regex special chars escaped (period in "tom t. hall" + bracket-class metas). Regexes cached lazily — first call compiles, subsequent calls reuse. 11 new regression tests pin the fix: Saturn / return / nocturnal / turning / burning DON'T fire Story Twist; the literal word "turn" / "twist" / "story song" / "tom t. hall" DO fire it. 50/50 genre-modes tests green; 39 pre-existing + 11 new. Affects EVERY song forged post-deploy + every dashboard chip rendered from song.mode. B3181 — Cover image in PDF report: PremiumReportData type was missing the coverArtUrl field even though DashboardSong carries it from the songs row's cover_art_url column. The PDF template (src/lib/export-song-html.ts) had zero references to cover/image and never rendered one. Fix: added coverArtUrl?: string to PremiumReportData (auto-threads through SongActionsButtonRow.buildReportData() via the ...song spread). PDF header now renders a 1.4-inch square cover thumbnail to the RIGHT of the title block via a 2-column grid layout (title + meta + diag + dates on the left, cover on the right). When coverArtUrl is missing/empty, the header falls back to the original single-column layout — no broken image icons. page-break-inside: avoid keeps the cover + title together. 201 tests green across genre-modes + surgery + report-sections.
  • **B12 (B3179) — Operator-reported bug fix: suggestedRevision was shipping prosody RULES as if they were lyric rewrites; RepairCard pre-filled the textarea with the instruction so Accept replaced the lyric with the rule (operator demo'd on a Latin song where "End on an open vowel..." replaced "Aquí en el marco de mi puerta"). The judge primitive correctly flagged it as risks ("destroying the actual lyric") but the dashboard should never have offered the instruction as a pre-filled revision in the first place. Fix: SurgeryItem type in line-level-surgery.ts gains a revisionType: 'guidance' | 'rewrite' discriminator. All 4 current source pullers (prosody-fatal / prosody-flag / taste-flag / heat-map-dead) marked as 'guidance' — they only have rules, not synthesized rewrites. RepairCard branches on the discriminator: when guidance, renders the suggestedRevision as a yellow-bordered HINT block labeled "Repair guidance" with the explanation "This is craft direction, not a singable line. Use it to write your own revision below." — and pre-fills the textarea with the ORIGINAL line (so the operator edits the real lyric, not the rule). When rewrite, keeps the legacy pre-fill behavior. Accept button now also disabled when the textarea still equals the original line (no-op guard). Optional discriminator in SurgeryDiagnosticSnapshot.surgeryItems for backwards compat with sessions persisted pre-B3179 — consumers default to 'guidance' (safer; won't auto-replace). 99/99 tests green; tsc clean. The B3148 restraint contract ("Song Surgery must be ABLE to refuse to suggest a change") is now mechanically enforced AT the UI layer, not just at the diagnose layer.
  • B11 (B3176+B3177+B3178) — Song Surgery Tier-3 discovery shipped (homepage callout + /surgery index + activity feed). B3176 — Homepage callout: src/app/page.tsx gains a new 3.6 section between HomepageBeforeAfter and the Leaderboard. Forge-styled centered card with the load-bearing tagline ("Don't regenerate your song. Save it." with the forge-gold "Save it."), one-paragraph value-prop (per-issue verdicts + per-arc Studio modes + verified Release Clearance + the restraint contract callout), and dual CTAs (See how Song Surgery works → /song-surgery; Forge a song first → /forge). Refine and Surgery now sit on the homepage as siblings — closes the mental-model gap between the two post-write operations. B3177 — `/surgery` index page: NEW server-component at src/app/surgery/page.tsx. Reads the operator's last 20 song_surgery_sessions ordered by updated_at DESC + joins songs for title/genre + bulk-fetches edit-status rows in one IN-query. Renders per-session cards: song title, status chip (complete/abandoned/in-progress/verifying with color-coded variants), accepted-skipped counts, verdict-color count chips (improves/changes/risks), verified composite delta with green-arrow/red-arrow indicator when the rescore landed, "Open Surgery on this song →" backlink. Empty-state CTA explains the 4-step path to first session. Tier-aware: Free/Creator see the "Session history is a Professional feature" upgrade panel with the SURGERY_LOCK_FEATURES list + /pricing CTA. Unauth users see a sign-in screen with marketing-landing fallback link. RLS on song_surgery_sessions gates per-user access automatically. data-testid="surgery-session-list" + data-testid="surgery-session-row" pin for e2e. B3178 — Activity-feed event: 'surgery.session.completed' added to AuditEventType in src/lib/audit-events.ts. /api/surgery/finalize now fires recordAuditEvent (fire-and-forget, service-role client) on successful session persistence — payload carries songId, accepted/skipped counts, netComposite, netReadiness, netDimensionScore, verifiedAvailable flag, branchedFromSessionId. Powers the dashboard activity feed when the operator returns to /dashboard. The dashboard activity-tab consumer wire is a thin follow-up that picks up this event-type alongside the existing 7 types. Tier 3 closed. All discovery moments wired (B3172-B3178 = 7 builds across 7 surfaces).
  • B10 (B3174+B3175) — Song Surgery Tier-2 discovery shipped (navbar + forge result handoff). B3174 — Navbar entry: src/components/Navbar.tsx TOOL_LINKS (the Browse dropdown) gains a "Song Surgery" entry with <Scissors /> icon between Examples and the Heirlooms surface. Hint text matches the marketing positioning: "Line-by-line lyric repair until your song is release-ready. Don't regenerate — save it. Pro." Placed in the dropdown rather than top-level NAV_LINKS to respect the B2566 4-item buyer-action discipline — Surgery is an authenticated post-write operation, not a buyer-action entry. The B3172 dashboard banner is the primary discovery path; the dropdown is the direct-nav fallback. B3175 — Forge result handoff: src/app/forge/v2-design/ForgeV2Result.tsx action button row gains a "Song Surgery →" text-button (mirroring the Refine button's polarity) between "Forge again" and "Refine this →". Renders only when forgeResult.id is present (saved songs). The /surgery page handles the restraint contract internally — when nothing's flagged, the operator sees "ship it" not an empty plan. Two post-forge operations now read as siblings (Refine = polish, Surgery = triage), not competitors. data-testid="forge-result-open-surgery" pins for e2e. Tier-2 discovery closed.
  • B9 (B3172+B3173) — Song Surgery discovery shipped (the door is finally on the building). Before this build, /surgery/[songId] had ZERO in-product entry points; the only way to reach it was to manually type the URL with a UUID. B3172 — Dashboard CTA: every owned status='complete' song on src/app/dashboard/SongDetail.tsx now renders a prominent forge-styled banner ("Song Surgery — Line-by-line lyric repair until your song is release-ready. Don't regenerate — save it.") with a <Scissors /> icon + an "Open Song Surgery" button linking to /surgery/{song.id}. Hidden when the song lacks lyrics or status (nothing to operate on). Banner sits between SongVisibilityBanner and RevisionHistoryPanel — load-bearing position. The /surgery page itself handles the restraint contract (renders cleanMessage when no findings exist), so we don't gate on diagnose output here. data-testid="song-detail-open-surgery" pins for e2e. B3173 — Marketing landing CTAs: /song-surgery hero "Open Song Surgery" button previously linked to /pricing — actively hostile to paid users who already had the feature. Now links to /dashboard?tab=songs (where the operator picks a song + hits the B3172 banner). Secondary "Forge a song first" already correctly went to /forge; preserved. The FINAL CTA at the bottom of the page ("Save your song" / "Upgrade to Professional") still links to /pricing — that's the deliberate tier-upgrade path, not the discovery path. Two surfaces, two intents, no more dead-ends.
  • B8 (B3171) — Verified rescore wiring shipped (Song Surgery's second follow-up gap closed). src/lib/surgery/verified-rescore.ts ships the pure-async wrapper around evaluateLyrics + computeReleaseReadiness. Takes the rewritten lyrics + optional genre + parent-song shape; returns { composite, readiness, evalData, durMs } or null on any error / timeout / disabled state. 30-second wall-clock cap via Promise.race against a TIMEOUT_SENTINEL; longer eval calls fall back to projection-only rather than hanging the operator. SF_SURGERY_VERIFIED_RESCORE_DISABLED=1 env toggle for emergency rollback. Input guards: empty lyrics → null, lyrics > 50KB → null, parse failure → null, missing compositeScore → null. Wired into /api/surgery/finalize/route.ts BEFORE the verification-record compose so the persisted session row carries the REAL verified scores (not the projection-only nulls). rescorePending response flag is now data-driven (verified === null) instead of the hardcoded true from B3162. Skipped when no accepted edits OR finalLyrics missing — nothing to verify. Vercel route maxDuration bumped to 60s to accommodate the eval call. 7 tests pin the env toggle + input guards; live-eval path covered by integration. The B3148 WAR Room load-bearing rule ("the verified rescore is the truth") is now actually true — the Final Vitals screen shows real verified numbers when the rescore succeeds, projection-only when it fails, and the trust verdict reports honest variance when the two diverge. Both Song Surgery follow-up gaps now closed. The feature ships without caveats.
  • B7 (B3170) — Session persistence shipped (Song Surgery follow-up gap closed). /api/surgery/finalize extended at B3170 to ALSO write the session row + each accepted/skipped/rejected edit to the B3155 song_surgery_sessions + song_surgery_edits tables. Uses the user's authenticated session client (NOT service-role) so RLS applies — the schema's INSERT policy requires auth.uid()=user_id, preserving the per-user audit chain. Fail-open contract: any persistence failure (column missing, RLS deny, network blip) is logged but the route still returns the computed verification record + verdict; sessionId comes back null and the client renders an "unsaved state" hint. New request body fields: songId, acceptedEdits[] (full records: itemId+oldLine+newLine+lineNumber+judgment), skippedEdits[] (itemId+lineNumber+reason), finalLyrics, diagnosticSnapshot, branchedFromSessionId. New response field: sessionId. RepairFlow updated: accepts new baselineLyrics + diagnosticSnapshot props from the page, computes finalLyrics by reducing applyEditToLyrics across all accepted edits in accept-order, sends the full session-state payload on auto-finalize, renders a green "Session saved · {short-id}" chip when persistence succeeded or a muted "Not saved" chip with explainer tooltip when it didn't, and clears the persisted-session id on revert (so branch attempts get their own session rows). 50-edit caps + 500-char per-line caps on the server. The "in-memory only" follow-up gap I flagged when declaring Song Surgery complete is now closed.
  • B6 (B3163) — Revert-to-version + branching shipped — Phase B CLOSED. RepairFlow gains in-memory branching state: each completed path that the operator chooses to revert from gets archived into a branches: SessionBranch[] array; the active branch label bumps via bumpVersionLabel each revert (1.1.0 → 1.2.0 → 1.3.0 → …). The completed-session view gains: (a) a "Try a different path" CTA card with <GitBranch /> icon + <RotateCcw /> button calling onRevertToBaseline() (resets currentIdx, accepted, skipped, lastJudgment, finalRecord, finalVerdict, finalRescorePending, finalizeError, isEditing, judgeError, pendingJudgment — then bumps the branch label); (b) the SessionTimeline now wires onRevert={(label) => label === '1.0.0' && onRevertToBaseline()} so the per-version Revert button in the timeline shares the same behavior; (c) when branches exist, a new BranchComparisonPanel renders above the CTA — shows each path as a row (label + Current chip + accepted count + net dimension score color-coded + per-axis spec/sing/mem/vuln dim chips + the trust-verdict message inline). The same surgery items list is reused across branches — the diagnose snapshot is frozen at session open. Phase B (B1-B6) ALL SHIPPED in 5 builds (B3159-B3163). Phase C (tier gating + marketing — C1-C4) is what remains in the Song Surgery feature arc.

Phase C — Tier gating + marketing (3-4 builds)

  • C1 (B3164) — Tier gates wired into Surgery shipped. src/lib/surgery/tier-gate.ts exports computeSurgeryAccess(tier) → SurgeryAccess with explicit fields per gate (fullAccess, studioAccess, maxItemsShown, showFinalVitals, allowBranching, showTimeline, upgradeCtaLabel). Free / Creator / Starter resolve to limited access (max 2 items, no Final Vitals, no branching, no timeline, 'Upgrade to Professional' CTA label). Pro + Admin resolve to full access (unlimited items + Final Vitals + branching + timeline). Studio currently aliases to Pro — C2 splits per-arc Studio modes off this same gate. SURGERY_LOCK_FEATURES exports the marketing copy fragment naming the exact features unlocked by upgrading. Page wired: reads profileRow.subscription_tier server-side, slices the snapshot's surgeryItems to access.maxItemsShown, and renders the lock banner above the Surgery Plan when !access.fullAccess — banner shows <Lock /> icon + "Song Surgery is a Professional feature" header + count of hidden items + bulleted feature list + forge-styled "Upgrade to Professional" link to /pricing. 9 tests pin the policy (null/undefined defaults, free/creator/starter/pro/admin resolution, marketing feature list invariants, type narrowing on the resolved tier). All 58 surgery tests green.
  • C2 (B3165) — Per-arc Studio modes shipped. src/lib/surgery/studio-modes.ts exports SurgeryArc (9-arc enum: rnb / country / pop / rock / indie / folk / worship / latin / rap + 'unknown'), detectSurgeryArc(genre) (pure-sync; rap-first regex check then falls through to detectGenreFamily from suno-genre-lock), getStudioModeProfile(arc) (returns StudioModeProfile with headerLabel, byline, craftSignature, sacredAccident), and the convenience helper studioModeForSong(genre). Each of the 9 arcs has a complete profile: Worship Surgery → "The divine subject must be NAMED, not implied — Berean Test" (SA#24) + APR/TFI/BVT signature; Country Surgery → "Authenticity is INHABITED, not INHERITED" (SA#20) + CID/LRR/MSC/VIG/SII signature; R&B Surgery → "Vulnerability requires receipts — the AID rule" (SA#23) + NCD/HVT signature; Pop Surgery → "Phonetic mass beats semantic precision" (SA#21); Rock Surgery → "Open vowels at the chorus peak" (SA#27); Indie Surgery → "The personal detail is the universal door" (SA#25); Folk Surgery → "Structure IS feeling" (SA#26); Latin Surgery → "A dialect is OWNED, not assembled" (SA#22); Rap Surgery → "A genre we cannot evaluate cannot be a genre we can serve" (SA#19). Page wired: when access.studioAccess === true AND the detected arc is not 'unknown', the page header label swaps from "Song Surgery — Diagnosis" to e.g. "Country Surgery — Diagnosis", with the per-arc byline + craft-signature paragraph + Sacred Accident citation rendered below the title. data-testid="surgery-mode-label" pins the header for e2e. The underlying diagnose engine + repair flow is unchanged — the per-arc audit primitives shipped through B2838-B3015 already feed into computeLineLevelSurgery via the existing eval pipeline. 18 tests pin the arc detection + profile lookups. Phase C status: C1 ✓ C2 ✓. Remaining: C3 (/song-surgery marketing landing) + C4 (pricing page Pro tier promotion).
  • C3 (B3166) — `/song-surgery` marketing landing shipped. New static page at src/app/song-surgery/page.tsx. force-static + Open Graph metadata + canonical pointing to ${BRAND.siteUrl}/song-surgery. Eight sections: (1) Hero with <Scissors /> icon + Song Surgery chip + the load-bearing tagline "Don't regenerate your song. Save it." (forge-gold "Save it.") + dual CTAs (Open Song Surgery → /pricing, Forge a song first → /forge); (2) The reframe (side-by-side cards: red-bordered "wrong question" = "Should I regenerate this song?" vs green-bordered "right question" = "Where does this song need surgery?"); (3) Restraint contract with <Shield /> icon + green-bordered blockquote of the clean-song message; (4) Four actions grid (Accept / Edit / Reject / Skip) with descriptions; (5) Per-arc Studio modes ordered list with <Sparkles /> icon — pulls all 9 arc profiles from getStudioModeProfile() rendering label + Sacred Accident citation + byline + craft signature; (6) Trust contract with yellow-bordered variance-message blockquote; (7) Branching panel with <GitBranch /> icon; (8) Final CTA — "Save your song." H2 + Professional-tier chip + dual CTAs (Upgrade → /pricing, How scoring works → /scoring/standard).
  • C4 (B3166) — Pricing page Pro tier promotion shipped. src/app/pricing/pricing-data.tsx Pro tier features now lists ★ Song Surgery — line-by-line lyric repair until your song is release-ready... as the SECOND ★-prefixed item (Release Dossier remains first per B3152). Also added a Song Surgery row to the comparison table: Free / Creator = "Preview (2 items)", Pro = "Full + per-arc Studio modes" — pinning the exact tier-gate semantics shipped in C1. Song Surgery Phase C COMPLETE. C1 ✓ C2 ✓ C3 ✓ C4 ✓. Full Song Surgery feature arc (B3154-B3166) shipped in 13 builds: A1-A5 foundations (5 builds), B1-B6 workflow UX (5 builds counting B2 folded into B1), C1-C4 monetization + marketing (3 builds with C3+C4 paired).

The 8 WAR Room findings (compressed):

1. 80% of infrastructure already shipped via the Release Dossier work. Song Surgery is a UX-layer reframing, not a new system. 2. The reframe IS the value. Marketing positioning + workflow design carry the product; the engine is done. 3. Load-bearing risk: score-delta over-promise. Mitigation: directions, not numeric promises in the user flow. Verified rescores at end-of-batch, not per-edit. When verified < projected, surface honestly. 4. The "improves vs. changes" primitive is new and load-bearing. A1 in Phase A — without it, the system can't tell when a fix is a regression in disguise. 5. Naming: "Song Surgery" + "Diagnose" + "Repair" + "Vitals" + "Release Clearance." Cap the metaphor — no scalpels, no operating rooms. 6. Retention curve differs from generation products. Investment-driven, not novelty-driven. Subscription pricing rewards it; per-credit pricing punishes it. 7. Version timeline = single dashboard component, not new system. 2-3 builds, not infrastructure. 8. The Studio tier monetizes the existing per-arc work. 9 genre arcs already shipped via Excellence WAR Rooms become Studio-tier features.

The single most important design constraint (per WAR Room Panel E):

Song Surgery must be ABLE to refuse to suggest a change. When runAllReportQaGates returns allPassed: true AND computeFinalRecommendation returns 'Ship' with high confidence, the Surgery page MUST show: "This song doesn't need surgery. Ship it." Not "here are some optional polish suggestions." Restraint is the trust signal. A surgery that always finds something to operate on is the gimmick that kills the brand.

Activation: when the operator types Continue Song Surgery (after the Release Dossier closes), the agent reads this section + finds the next unchecked A/B/C item + ships it. Same pattern as Continue Release Dossier and Continue Punch List.


Constraint-Aware Forge (CAF) — 21-build roadmap (HIGHEST PRIORITY)

Triggered by: 5-song stress-test WAR Room × 100 rounds + Sacred Accident #17 (Build 2763). The reviewer's 10 recommendations all collapse to one parent failure: the model optimizes for craft and forgets the brief. The fix is fidelity as a first-class score, orthogonal to quality. 21 builds across 3 phases.

Activation phrase: Continue Constraint-Aware Forge — reads SA#17 + this section + ships the next unshipped item.

Phase 1 — Foundation (week 1-2)

  • CAF-1 — Brief extractor library. src/lib/claude/brief-extractor.ts. Haiku-powered. Extracts { premise, anchors[], structure[], styleConstraints[], forbiddenLanguage[], constraintMode, ambiguities[], contradictions[] } from raw prompt. Failure-safe; returns permissive default on parse failure. (B2764 — shipped.)
  • CAF-2 — Structure compliance audit. src/lib/claude/audit-structure.ts. Pure regex. Compares brief.structure to actual section markers; returns { requested[], actual[], hits, misses, partial, score }. Partial credit for adjacent-equivalent sections. (B2765 — shipped.)
  • CAF-3 — Thesis language detector. src/lib/claude/audit-thesis.ts. Pure regex. Flags thesis-shaped lines (declarative analysis register, moral summary, therapy register). Returns { flaggedLines[], count, score }. (B2766 — shipped.)
  • CAF-4 — Per-section specificity in prosody engine. Extend src/lib/claude/prosody-engine.ts. Adds report.sections[i].specificity: { concreteNouns, physicalActions, sensoryWords, score } per section. (B2767 — shipped.)
  • CAF-5 — Wire Phase 1 through the pipeline + dashboard. Shipped across 4 builds B2769-B2772. (5a/B2769) Orchestrator src/lib/fidelity-audit.ts composes brief + 3 audits + forbidden-language regex into one runFidelityAudit() call (26 tests). (5b/B2770) songs.fidelity_audit JSONB column via add_fidelity_audit_to_songs.sql migration + partial index for future leaderboard composite ranking; updateSong accepts fidelityAudit payload; DashboardSong carries it through. (5c/B2771) finalize-song.ts runs extractBrief(rawPrompt) + runFidelityAudit() inline post-forge; SF_FIDELITY_AUDIT_DISABLED=1 escape hatch; per-song forge.finalize.fidelity_scored log emit. (5d/B2772) FidelityPanel.tsx renders the composite grade + brief summary + per-component breakdown in SongDetail. The full data flow exists end-to-end: raw prompt → extractBrief → audit → DB → dashboard.

Phase 2 — Prompt-level + grade infrastructure (week 3-6)

  • CAF-6 — Forge prompt amendments. Shipped across 3 builds B2773-B2775. (6a/B2773) renderBriefBlock(brief) helper in src/lib/claude/render-brief-block.ts — pure prompt fragment with 8 canonical sections (mode framing, premise, required details, required structure, style constraints, forbidden language, ambiguity watch, contradictions). 16 tests. (6b/B2774) extractBrief moved UPSTREAM to preforge-orchestration stage 0; runs in parallel with stages 1-7; net wall-clock ~0; Brief threads through PreforgeOrchestrationResult.brief → post-forge-pipeline.ctx.preforge.brief → finalize-song.ctx.brief (reuses the cached Brief — no redundant Haiku call). New telemetry: preforge.brief_extracted. (6c/B2775) Forge prompt assembler (buildForgeUserPrompt in user-prompt.ts) accepts optional brief in BuildForgeUserPromptOptions; renderBriefBlock fragment PREPENDS the prompt in all 3 branches (empty-prompt / starter-lyrics / freeform); forgeSongStream + runForgeCore + forge-stream-session all thread the Brief through. The model now writes against explicit anchor / structure / style instructions BEFORE writing line one.
  • CAF-7 — Anchor coverage audit. Shipped across 2 builds B2776-B2777. (B2776) Extracted CONCRETE_NOUNS / PHYSICAL_VERBS / SENSORY_WORDS vocabulary + tokenizer into src/lib/claude/specificity-vocab.ts so CAF-4 and CAF-7 share the same line-specificity counter (no drift between two audits). (B2777) src/lib/claude/audit-anchor-coverage.ts ships the audit: extract key tokens from each anchor (drop stopwords, keep numerics), find the best-matching sung line per anchor (token-match count, tie-break by specificity), score on the 4-band scale per the WAR Room verdict — 0 missing / 33 mentioned-bare / 66 partial / 100 delivered. Pure-sync (no Haiku call; the WAR Room option for Haiku second-opinion is deferred until evidence shows the heuristic disagrees with human judgment). 21 tests. Orchestrator now populates anchorCoverage component when brief has anchors; FidelityAudit.phase ratchets to 'phase-1.5'; auditVersion bumps 1.0.0 → 1.1.0. SongDetail FidelityPanel renders new per-anchor breakdown with evidence lines + the hits/partial/missing counts above the per-component score.
  • CAF-8 — Chorus Evolution Planner. Shipped across 3 builds B2780-B2782. (8a/B2780) src/lib/claude/chorus-evolution-planner.ts — Haiku pre-write phase produces { position1, position2, position3, arcType, arcDescription } for the 3 chorus positions (belief / realization / can't-deny). 5 canonical arcs + custom; failure-safe; 9 tests. (8b/B2781) Wired upstream into preforge stage 0.5 (chains on briefPromise; parallel with stages 1-7). New renderChorusEvolutionBlock helper renders the prompt fragment which the forge prompt assembler prepends after the brief block. Threaded through forge-stream → run-forge-core → forge-stream-session → buildForgeUserPrompt. New telemetry: preforge.chorus_evolution_planned / preforge.chorus_evolution_skipped. (8c/B2782) src/lib/claude/audit-chorus-evolution.ts — pure-sync audit detects whether choruses byte-shifted OR verses reference back to chorus content. 4-band scoring (0 static / 33 weak / 66 partial / 100 evolved). Wired into runFidelityAudit so componentScores.chorusEvolution populates (5% weight). FidelityAudit.phase ratchets to 'phase-1.75'; auditVersion 1.1.0 → 1.2.0. FidelityPanel renders planned-vs-shipped (planner's 3 positions + audit's state classification + chorus count + verbatim-repeat count). 10 tests.
  • CAF-9 — Earned Transcendence. Shipped across 2 builds B2783-B2784. (9a/B2783) src/lib/claude/audit-transcendence.ts — pure-sync audit detects whether a concrete-noun image planted in V1 returns in the final chorus with transformed surrounding words. 4-band scoring (0 missing / 33 verbatim / 66 partial / 100 transformed) on Jaccard similarity of non-image content words. Charitable candidate selection: when multiple concrete-noun images recur, picks the one with the LARGEST surrounding shift. 13 tests. (9b/B2784) Wired into runFidelityAudit so componentScores.transcendence populates (5% weight per SA#17). FidelityAudit.phase ratchets to 'phase-2' when the audit produces a real classification (state !== 'na'). auditVersion bumps 1.2.0 → 1.3.0. FidelityPanel renders the callback image + V1/final-chorus line pair + surrounding-word similarity meter. 4 orchestrator tests (na branch / transformed scoring / phase-2 ratchet / version bump).
  • CAF-10 — Conditional Sensory Rewrite. Shipped across 2 builds B2785-B2786. (10a/B2785) src/lib/claude/sensory-rewrite.ts — Sonnet-driven targeted line-level rewrite library. Gate predicate shouldRunSensoryRewrite(brief, thesisAudit) fires when brief.styleConstraints includes one of ['sensory-only', 'show-dont-tell', 'no-thesis'] AND thesisAudit.flaggedLines is non-empty. runSensoryRewrite() calls Sonnet with flagged lines + ±2 lines of context + brief anchors, parses the JSON response, verifies each rewrite against the audit (rejects model-side line-number drift), stitches the rewrites in. Failure-safe: every error path returns fired: false + original lyrics. 22 tests. (10b/B2786) Wired into finalize-song.ts BEFORE the contribution ledger + prosody + fidelity audit blocks, so all three downstream artifacts see the FINAL (possibly rewritten) lyrics. Pre-rewrite thesis audit (~3ms regex) drives the gate; full fidelity audit runs once on the rewritten lyrics. SensoryRewriteResult metadata attached to the persisted FidelityAudit JSONB so the dashboard reads from one column. FidelityPanel renders before/after pairs (strikethrough original + green rewrite) with thesis-category labels per pair. Telemetry: forge.finalize.sensory_rewrite_applied / _failed / _threw log events. Env toggle: SF_SENSORY_REWRITE_DISABLED=1 for the same shape as SF_FIDELITY_AUDIT_DISABLED.
  • CAF-11 — Final Adherence Audit. Shipped across 2 builds B2787-B2788. The other CAF-11 sub-items (11a/11b/11c — complexity + composite + grade helpers) shipped at B2768. (11/B2787) src/lib/claude/audit-premise.ts — async Haiku premise-match judgment. Returns { score, verdict, reasoning, premiseEcho, evidence, state, durMs }. 4 verdict bands (wrong-song / drifts / mostly-served / faithfully-served). NA branch skips Haiku; failure-safe on every error path. 18 tests. (B2788) RunFidelityAuditInput gains optional premiseAuditResult (matches CAF-8c chorusEvolution pattern, keeps orchestrator pure-sync). FidelityAudit gains premiseAudit field. componentScores.premise populates from result.score when state='judged' (otherwise null — 'na'/'error' both redistribute weight). Phase ratchets to 'phase-2' when EITHER premise OR transcendence is judged. auditVersion 1.3.0 → 1.4.0. finalize-song.ts calls auditPremise upstream from runFidelityAudit, threads premiseAuditResult through. SF_PREMISE_AUDIT_DISABLED=1 toggle. FidelityPanel renders premise panel FIRST (load-bearing 30% slot) with verdict chip + premise echo vs brief-premise diff + reasoning + evidence lines + tier legend. 6 orchestrator tests (null when omitted, null on 'na' state, null on 'error' state, populates on 'judged', phase ratchets to phase-2, version 1.4.0). CAF Phase 2 is COMPLETE — all 7 fidelity components have heuristic or judged scoring.
  • CAF-11a — Brief complexity computation. computeBriefComplexity(brief): 0-10 based on anchor count + structural reqs + style constraint count + forbidden-language list. (B2768 — shipped.)
  • CAF-11b — Fidelity composite. computeFidelityComposite(audits): { score, perComponent, grade, complexityBucket }. 30/25/15/15/5/5/5 weighting (premise + anchors + structure + style + forbidden + chorus + transcendence). (B2768 — shipped.)
  • CAF-11c — Grade helpers. fidelityScoreToGrade(score) (A+/A/B+/B/C+/C/D/F bands) + fidelityScoreToBucket(score, complexity) (hide/secondary/primary). Mirrors src/lib/scoring.ts shape. (B2768 — shipped.)
  • CAF-11d — Dashboard SongDetail two-grade surface. Shipped B2789. New src/components/SongHeroFidelityPill.tsx renders fidelity score + grade chip alongside the existing SongHeroScorePill in the SongDetail hero band. shouldRenderFidelityPill(composite) decides whether to mount: 'primary' (heavy brief) and 'secondary' (standard brief) always render; 'hide' (light brief) suppresses unless score drops below 80 (operator-spec'd rescue path) — pill then renders with low fidelity warning + orange border. In 'primary' bucket, fidelity pill renders FIRST (headline position) AND with the same large 2xl number as the quality pill; in 'secondary' it renders SECOND at slightly smaller weight (xl). Title attribute carries the grade label + complexity score + bucket reason for operator-facing hover-tip. Per-component breakdown disclosure panel was already shipped in CAF-5d/B2772 (FidelityPanel). 5 tests pin the gate helper across all 4 bucket × score combinations.

Phase 3 — Ratchets + public docs + Rap Mode (month 2+)

  • CAF-11e — Leaderboard composite migration. Shipped B2791. src/lib/leaderboard/top-weekly.ts now ranks by 0.6 × forgeScore + 0.4 × fidelityScore when the song's persisted fidelity_audit carries a composite.score number. Pre-CAF songs (forged before B2771) fall back to quality alone so they're never silently demoted by the migration. New computeRankScore(q, f|null) helper exported for unit tests + dashboard surfacing. TopWeeklySong interface gains fidelityScore + rankScore fields so the leaderboard UI can show the "why this position" explainer. Query overfetches when fidelity ranking is active (composite re-ranking can promote songs from below the SQL ORDER BY cutoff). SF_LEADERBOARD_FIDELITY_DISABLED=1 env flag forces the legacy quality-only sort as the emergency rollback. 6 tests pin the formula (null-fallback identity, parity at Q=F, high-Q+low-F drop, low-Q+high-F rise, integer rounding, pre-CAF no-penalty invariant).
  • CAF-11f — `/scoring/standard/fidelity` public docs. Shipped B2792. New static page at /scoring/standard/fidelity documents v0.1.0 of the standard. CC BY 4.0 licensed. ScholarlyArticle + TechArticle JSON-LD schema for Google Scholar / Semantic Scholar indexing. 8 sections: (1) Fidelity is the orthogonal question (vs quality); (2) The seven components — table + per-component drill-downs naming the Haiku/heuristic split; (3) Composite formula with the constraint-mode multipliers + null-component redistribution math in code block; (4) Grade calibration A+ through F with color-coded chips matching the dashboard; (5) Brief complexity gating (hide/secondary/primary UX buckets); (6) Version history (v0.1.0 → RFC); (7) How to cite (plain-text + BibTeX); (8) Related standards cross-link to Lyric Scoring Standard. Receipts-row addition pending CAF-14.
  • CAF-11g — RFC-0010 opens. Shipped B2793. RFC-0010 "Fidelity Score v0.1.0 — calibration + composite formula" added to src/lib/rfcs.ts. Opens 7-day public comment window (opened 2026-05-19, commentDeadline 2026-05-26). Body pins: component weights (30/25/15/15/5/5/5), composite formula with null-component redistribution math, constraint-mode multipliers (strict 0.95 / standard 1.0 / loose 1.15), 8-tier grade calibration, brief complexity formula, 3 UX prominence buckets + rescue path. Five open questions for the comment window covering weight allocation / mode caps / complexity threshold / chorus+transcendence weights / 'na' verdict semantics. Auto-rendered at /rfc/0010-fidelity-score-v0-1-0-calibration-and-composite-formula via the existing [slug] route. Status 'in-comment' (gold chip). The fidelity standard page (/scoring/standard/fidelity) links to the RFC from the Version history section.
  • CAF-12 — Fidelity score CI ratchet. Shipped B2790. src/lib/golden-evals/fidelity-fixtures.ts defines 4 hand-curated { brief, lyrics } fixtures spanning the fidelity distribution: (1) faithful-anchor-delivery (high score, A/B band), (2) structure-miss-truncated (4 sections shipped vs 6 requested → mid band), (3) sensory-only-thesis-leak (style constraint violated by thesis-shaped lines → C/D band), (4) minimal-brief-decent-song (light brief baseline → tests redistribution math). Each fixture pins an expected composite band (≥10 wide to absorb minor weight-redistribution drift) + an expected grade prefix set. src/lib/golden-evals/fidelity-ratchet.test.ts runs runFidelityAudit on each fixture + asserts the composite stays inside the band and the grade prefix matches. Plus 4 orchestrator-behavior assertions (phase ratcheting, structure miss detected, thesis flags counted, light complexity verified). Premise audit deliberately omitted from the ratchet — Haiku-driven, non-deterministic, separate ratchet on schedule. 13 tests total. Catches regressions in the audit stack OR composite math that wouldn't fire any unit test.
  • CAF-13 — Rap Mode (first genre-craft module). Shipped B2794. New MODE_RAP entry in src/lib/claude/genre-modes.ts (Genre Modes infrastructure B2170). Trigger keywords: rap / hip hop / hip-hop / trap / boom bap / drill / mc. Panel emphasis: Kendrick Lamar / Earl Sweatshirt / MF DOOM / Nas / Rapsody / Pat Pattison (channeled). Rubric weight overrides: M3 (Rhyme Intelligence) × 1.5, M5 (Specificity) × 1.3, M11 (Memorability) × 1.2, M1 (Prosody) × 1.2 — making rhyme the load-bearing metric. Mode floor: M3 ≥ 78. Success criteria + chorus discipline name the form's discipline: "the verse is the song; the hook is the runway"; "dead bars are the failure mode; every bar must earn its position." Companion src/lib/claude/audit-rap-craft.ts ships three pure-sync heuristics: auditBarLength (per-line syllable count + mean/stdDev/range/outlier-line flagging), auditEndRhymes (couplet AABB + alternating ABAB rate detection via last-4-chars suffix matching), auditInternalRhymes (repeated-suffix bigram counter within each line). auditRapCraft orchestrator runs all three in one call. 25 tests pin the heuristics + 5 detector tests confirm MODE_RAP fires on the right keywords. The audits aren't yet wired into the fidelity composite — that wiring lands once empirical data shows the heuristics correlate with human judgment.
  • CAF-14 — `/scoring/standard/fidelity` final public artifact. Partial — shipped B2793 (changelog stub); v1.0.0 release artifact lands after RFC-0010 closes 2026-05-26. New static page at /scoring/standard/fidelity/changelog with append-only version history (initial entry: v0.1.0 in-comment, dated 2026-05-19, with 6 detail bullets). Status chip rendering (draft / in-comment / released) mirrors the RFC chip color scheme. Implementation-version disclaimer at the bottom names FIDELITY_AUDIT_VERSION as the codepath that tracks separately. Cite-this-page shipped at B2882 (/scoring/standard/fidelity/cite — BibTeX, APA 7, MLA 9, Chicago, RIS, plain text). B3195 (CAF-14 advance — npm package sync): @songforgeai/fidelity-standard bumped 0.1.0 → 0.2.0 to match the live standard. Pre-3195 the package was publishing v0.1.0 (7 components) while the live standard page + audit engine were at v0.2.0 (8 components, registerAdherence added at B2889/CAF-11h) — external citers were getting a stale shape. Sync changes: public/fidelity-standard.json regenerated at v0.2.0 with 8th component (registerAdherence, 6% weight, heuristic method) + new changelog field with v0.2.0 + v0.1.0 entries. FidelityComponentId TS union extended with 'registerAdherence'. New FidelityChangelogEntry interface + optional changelog? field on FidelityStandard. Package CHANGELOG.md updated; dist/standard.json regenerated via tsc && cp build script. 23 new tests pin: version 0.2.0, 8 components in canonical order, registerAdherence weight 0.06 + heuristic method, weight sum 1.06 (normalized at compute time), summary text mentions "eight-component", changelog newest-first ordering, RFC-0010 still in-comment + 2026-05-26 deadline, FidelityComponentId union accepts registerAdherence, all 8 ids satisfy the union, computeFidelityGrade/Bucket/BriefComplexity/applyConstraintMultiplier still work correctly. 23/23 tests green; 75/75 main repo fidelity tests still green; tsc clean. The v1.0.0 release artifact (CAF-14 final) still lands after RFC-0010 closes 2026-05-26 — operator runs `cd packages/fidelity-standard && npm publish` to push v0.2.0 to npm when ready.

Operator-feedback nightly run B3391-B3397 (shipped 2026-05-28)

Operator reported on 2026-05-28: (1) "clavicle is showing up in lots of lyrics" — system-internal feedback loop where B3367's banned-term ALTERNATIVES list ("Try ribs, clavicle, shins, knees, jawline") was being treated as suggestions, and the alternatives became the new cliché. (2) "Songs still seem to go to folk as the default when no genre specified in the prompt" — a different class-of-bug from B3276-B3284's genre-routing collapse (which fixed cases where genre WAS specified but lost); this is the no-genre default case left silent.

Shipped a 7-build response in autonomous mode:

  • B3391 — Quick wins. Strip clavicle/shins/jawline from BANNED_TERMS alternatives; add them as their own banned entries; kill the 30% random Math.random() < 0.3 folk-attractor roll in single-phase-prompt.ts.
  • B3392 — Catalog motif scanner module. Pure-function scanner with stem-folding (silent-e restoration), partitions output into overused vs leakingBanned buckets, 15 tests.
  • B3393 + B3394 — Closed loop: cron + forge-prompt reader. /api/cron/scan-catalog-motifs runs nightly at 03:30 UTC (added to vercel.json); writes scan to trend_audits with marker [catalog-scanner] v1. dynamic-forbidden-motifs.ts reads latest scan, injects ## DYNAMIC FORBIDDEN MOTIFS block at tail of forge system prompt. 17 tests on the reader.
  • B3395 — Sentiment-routed default-genre picker. suggestDefaultGenre(prompt) Haiku call (~$0.001-0.002/no-genre forge). Calibrated to AVOID folk by default; folk fires only on explicit reflection + grief + slow + acoustic signals. Wired into preforge-orchestration.ts. Escape hatch via SF_DEFAULT_GENRE_PICKER_DISABLED=1. 13 tests.
  • B3396 — DECADE SHIFT gate cleanup. Regex scoped from topic words (grief|loss|memory|story) to genre words only (folk|country|indie|singer-songwriter|ballad|americana). Added structured log forge.default_genre.surfaced in post-forge-pipeline.ts.
  • B3397 — R5-A1 + R5-A3 cleanup. Refreshed src/lib/.git-snapshot.json (was 67 commits behind). Plumbed brief.toneRegister through gauntlet system prompt via songs.fidelity_audit.brief.toneRegister lookup.

Verification windows:

  • The catalog-motif cron fires for the first time tonight (03:30 UTC). Operator can verify via /api/cron/scan-catalog-motifs GET with Bearer token or via Vercel cron logs.
  • The dynamic-forbidden block appears on the SECOND forge after the first scan completes — the FIRST scan writes the row; subsequent forges read it.
  • The default-genre picker fires on the next no-genre forge; grep logs for forge.default_genre_picked.

Deep Audit 2026-05-27 round 5 — autonomous queue (TOP PRIORITY)

Triggered by: the post-B3383 Deep Audit, captured at the end of today's 12-build SA#32 register-awareness arc (B3372 → B3383). Final grade B+ (7.4/10 weighted). Strongest: trust/credibility (9), process maturity (10), technical execution (9), strategic differentiation (9). Weakest: information architecture (5 — new register-awareness page orphaned), conversion architecture (5 — CTA dilution), onboarding (5 — rich-prompt pattern undocumented), commercial readiness (6 — distribution N=0). Engineering velocity continues to outrun marketing maturity by 2 quartiles; outrunning distribution by 3.

Live verification at audit time: /api/version returns {build:3383, sha:59888b3fd, ref:main} — production matches HEAD. check:all 58/58 green. npx vitest run 1 failed + 8587 passed (the failure is the load-bearing P0 below).

Operator-side complements live in docs/OPERATOR-PUNCH-LIST.md under the matching "Audit-Round-5" section. Engineering items below ship autonomously.

Round-5 Tier 1 (ship THIS WEEK — small, each ≤2h)

  • R5-A1 — Refresh `src/lib/.git-snapshot.json`. [AUTONOMOUS · P0 · ~5 min] Shipped at B3397 (2026-05-28). npm run snapshot:refresh wrote 300 commits + 300 file-touch entries; unblocks 4 public surfaces (/changelog, /engineering, /roadmap, /now) that had been rendering pre-B3331-era content.
  • R5-A2 — Add `register-awareness` to `StandardsClusterFooter.tsx`. [AUTONOMOUS · P0 · ~10 min] Shipped at B3398 (2026-05-28). Added 'register-awareness' to the SurfaceSlug union + a SURFACES entry (Activity icon, blurb naming SA#32 + 10 registers + Gravity / Burden modifiers). ALSO mounted <StandardsClusterFooter currentSurface="register-awareness" /> on the page itself so it both appears in siblings' footers AND back-links to them.
  • R5-A3 — Plumb `brief.toneRegister` through the gauntlet system prompt. [AUTONOMOUS · P0 · ~30 min surgery] Shipped at B3397 (2026-05-28). Reads from songs.fidelity_audit.brief.toneRegister JSON on the song row at gauntlet start; passes as the 11th arg to gauntletStream. Defensive shape walk means missing column / undefined register falls back to pre-B3380 register-blind prompt. Closes the dormant B3380 infrastructure — every register-declared song's gauntlet pass is now register-aware.
  • R5-A4 — Add SA#32 chip above-fold on homepage. [AUTONOMOUS · P0 · ~30 min] Shipped at B3398 (2026-05-28). Added "Register-aware · 10 modes" link to the hero proof strip (mono text, no caps, no color per the B3289 "trust whispers" discipline). Linked to /scoring/standard/register-awareness. Solves orphan-page discoverability AND the brand-promise gap in one chip.

Round-5 Tier 2 (ship THIS MONTH — bigger, each <1 day)

  • R5-A5 — Align git-snapshot-freshness thresholds. [AUTONOMOUS · P2 · ~1h] Shipped at B3400 (2026-05-28). Documented the divergence explicitly in both gates' headers (not aligned thresholds, because they intentionally measure DIFFERENT artifacts: git-snapshot.json by commit-count vs CLAUDE.md AUTO-STATE block by calendar days). Each header now explains its own metric AND points to the sister gate's metric so future operators understand why one can be red while the other is green. They are complementary, not contradictory; running BOTH cures (npm run snapshot:refresh + npm run docs:refresh) is the pre-push ritual.
  • R5-A6 — Tighten `check:corpus-claims` keyword detector. [AUTONOMOUS · P2 · ~30 min] Shipped at B3399 (2026-05-28). Split CLAIM_PHRASES into STRONG_CLAIM_PHRASES (corpus-specific phrasings that fire on their own — hand-scored exemplars / corpus entries / reference corpus) and AMBIGUOUS_CLAIM_PHRASES (worked examples / reference entries — only fire when a corpus-context anchor appears within ~150 chars). Anchor list: corpus / exemplar / calibration / rubric / golden evals / ground-truth / anchor corpus / hit calibration / sleeper ledger. Existing scan still green; B3384-style false-positives now structurally impossible.
  • R5-A7 — Disambiguate `brief.register` → `brief.contentRating`. [AUTONOMOUS · P1 · ~30 min cross-codebase] Shipped at B3399 (2026-05-28). UI-side relabel: every "Register adherence" surface now reads "Content rating adherence" (/scoring/standard/fidelity table + heading + prose, fidelity-grade.ts component label, fidelity-audit-shape.json snapshot). Added a disambiguation note in the prose: "Not to be confused with the SA#32 tone register (joy / swagger / rage / playfulness / etc) — content rating is the MPAA-style profanity perimeter; tone register is the emotional posture." The internal field key (brief.register, registerAdherence) keeps its original name for persisted-fidelity_audit-JSON schema compatibility; only the user-facing labels change so production rows still parse.
  • R5-A8 — Add `toneRegister` UI surface to `/crucible`. [AUTONOMOUS · P2 · ~2h] Shipped at B3403 (2026-05-28). Added a 4-chip radio row (Joy / Swagger / Rage / Playfulness) styled with per-register colors (violet / amber / red / pink) directly under the voice-mode toggle. Selection is persisted in localStorage (sf_crucible_register) so a returning user keeps their pick. The chip threads through to the POST body as toneRegister, activating the B3378 per-voice register appendix. Linked to /scoring/standard/register-awareness so users can read the SA#32 rationale inline. The 6 other canonical registers (grief/melancholy/awe/tenderness/lust/defiance) are intentionally omitted from the UI — they're either the home register (no appendix change) OR awaiting operator-paced calibration per SA#11.
  • R5-A9 — Cross-link audit ratchet. [AUTONOMOUS · P2 · ~2h] Shipped at B3401 (2026-05-28). scripts/check-cross-link-audit.ts walks every page.tsx added via git log --since="30 days ago" --diff-filter=A, computes each route's self-cluster (parent path prefix), and asserts ≥1 inbound link from a route outside that prefix. Component files (src/components, etc.) count as cross-cluster automatically. Warn-only at launch (CROSS_LINK_AUDIT_BLOCKING=0 default) so the existing baseline (4 untriaged surfaces: /admin/corpus-novelty, /standard-vs/[slug], /design-system, /security, /status) doesn't break CI; flip to blocking once those are cross-linked. Wired into check:all as the 59th gate.
  • R5-A10 — Synthetic post-deploy probe pack for cluster-footer pages. [AUTONOMOUS · P2 · ~2h] Shipped at B3402 (2026-05-28). scripts/probe-standards-cluster.ts regex-parses the SURFACES array from StandardsClusterFooter.tsx and GETs each ${PROBE_BASE_URL}${href}, asserting 2xx. Defaults to https://songforgeai.com; override via PROBE_BASE_URL env var. Six-test parser-side contract at src/lib/standards-cluster-parser.test.ts pins the registry shape so a future refactor can't silently break the parser. Available as npm run probe:standards-cluster. Run post-deploy or as part of trust-decay-audit checklist.
  • R5-A11 — Add SA#32 chip + per-register example songs above the homepage fold. [AUTONOMOUS · P1 · ~3h] Surface 1 song per shipped register (joy/swagger/rage/playfulness) via inline player cards. The /examples page exists; reuse the card pattern. Converts SA#32 work from "claimed in a sub-page" to "demonstrated immediately."
  • R5-A12 — Forge prompt template chips on `/forge`. [AUTONOMOUS · P1 · ~3h] Add buttons that prefill the prompt textarea with rich-prompt templates: "Joy · single song", "Country · alt-country · breakup", "Worship · contemporary · doxology". Instant on-ramps that activate the genre + register pipelines without requiring the user to know the rich-prompt pattern.
  • R5-A13 — In-product first-run tutorial for rich-prompt pattern. [AUTONOMOUS · P2 · ~4h] Shipped at B3406 (2026-05-28). src/components/ForgeRichPromptHint.tsx — dismissable side-by-side "Thin" vs "Rich" prompt cards mounted above the prompt textarea on /forge. "Try this rich prompt" button seeds the textarea via onSeed callback (the seed is the joy-001 prompt from the B3405 eval fixtures). localStorage dismissal (sf_forge_rich_prompt_hint_dismissed) — separate from B2881's Crucible-redirect hint key so the two hints can coexist. Links inline to /scoring/standard/register-awareness. Demonstrates the pattern instead of explaining it abstractly.

Round-5 Tier 3 (next quarter — multi-build)

  • R5-A14 — N=large register-fidelity eval framework. [AUTONOMOUS · P1 · 3-5 builds → shipped in 1] Framework shipped at B3405 (2026-05-28). data/eval-fixtures/register-fidelity-prompts.json (50 prompts × 5 registers). scripts/eval-register-fidelity.ts runner — smoke mode default (5 forges), --full --trials 3 for the audit-spec 150-forge run. src/lib/eval/register-fidelity-report.ts pure aggregator with 18 tests (extraction accuracy / AVD pass rate / declaration rate / composite mean+median per register). scripts/report-register-fidelity.ts CLI: JSONL → markdown. Full methodology + interpretation guide in docs/REGISTER-FIDELITY-EVAL.md. The actual 150-forge run is an operator budget decision (~$15-30 + 3-7 hours) — harness ships ready; running is a separate operator action.
  • R5-A15 — 50-voice room rebalance. [AUTONOMOUS · P1 · 4-6 builds] Add Lizzo / Bruno Mars / Kesha / Max Martin archetype voices to balance the Cohen / Mitchell / Bridgers-heavy panel. The B3372 WAR ROOM §3.1 named this as the deepest residual bias. Even with B3373-B3383 register-aware infrastructure, the voice panel itself remains skewed toward the introspective register and drags joy/swagger drafts toward home register. ~30% of voices need rebalancing.
  • R5-A16 — Per-tier feature delta highlighting SA#32 coverage. [AUTONOMOUS · P2 · 2 builds] Pricing page today: 3 tiers + credit packs. Add a "Register coverage" row in the comparison table: Free = grief/melancholy default; Creator = + joy / swagger; Pro = full 10-register support + register-aware refinement. Turns SA#32 into commercial value differentiation.
  • R5-A17 — Per-register × per-genre intersection landing pages. [AUTONOMOUS · P3 · 3-5 builds] Long-tail SEO capture: /lyrics/joy/pop, /lyrics/swagger/rap, /lyrics/rage/rock. ~30 pages × per-genre × top-3-registers. Each links to the register-awareness doc + the genre-specific craft plugin.

Round-5 Tier 4 (category-defining engineering bets)

  • R5-A18 — Open-source the SA#32 corpus + heuristic framework. [AUTONOMOUS · P3 · 2-3 builds] Publish @songforgeai/tone-register as an npm package: the detectToneRegisterHeuristic + registerAppendix + registerAntiInflationBlock + gauntletRegisterDirective primitives. Standalone TypeScript, model-agnostic, no SongForgeAI-specific assumptions. Same playbook as @songforgeai/agent-room (PUNCH-LIST-V2 #1-3). Academic / OSS adoption surface for SA#32 methodology.
  • R5-A19 — `/admin/register-stats` analytics dashboard. [AUTONOMOUS · P3 · 1-2 builds] Shipped at B3404 (2026-05-28). /api/admin/register-stats aggregates songs.fidelity_audit->brief->toneRegister over a configurable window (24h/7d/30d/90d), groups by register (with null bucket for undeclared) + computes count / percent / mean / median / min / max forge_score per register. Page at /admin/register-stats with auth-tier gate, window selector, sortable table, per-register tint colors mirroring B3403. Linked from /admin home via ADMIN_ROUTES entry. AVD verdict not yet persisted — currently emitted as logs only; route + page note the gap. Closes the empirical-measurement side from N=1 to N=catalog.

Cross-AI Feedback WAR ROOM (2026-05-28 — B3408)

Triggered by operator's 9-genre test sweep + cross-AI feedback synthesis. Full WAR ROOM doc: docs/WAR-ROOM-CROSS-AI-FEEDBACK-2026-05-28.md. WAR ROOM declined to ship 2 of the feedback's recommendations and shipped 1 small change inline; the remainder are queued below.

  • R5-A20 — Anti-thesis-line banned phrases. [AUTONOMOUS · P2 · ~10 min] Shipped at B3408 (2026-05-28). 7 thesis-line phrases added to src/lib/banned-terms.ts as category: 'house_style', tier: 2 — "the sound of," "this is the sound," "this is where," "the shape of," "the language of," "the architecture of," "the geography of." Catches the "commentary-on-the-song" construct the cross-AI feedback called out explicitly. Post-gen scrub fires automatically.
  • R5-A21 — Mess-injection forge amendment. [AUTONOMOUS · P2 · ~2 builds] Add a forge prompt directive that requires one OFF-ANGLE / CONTRADICTORY / IMPERFECT line per song. The feedback's "Let one be ugly. A cracked voice, an embarrassing memory, a line that's not clever but true." The amendment lives in single-phase-prompt.ts as an optional features block under a feature flag for A/B measurement. Risk: medium — overshoots into bathos if not bounded. Implementation includes a Haiku judge post-forge to verify the off-angle line is present without being maudlin.
  • R5-A22 — Chorus-mass test primitive. [AUTONOMOUS · P2 · ~1 build] SHIPPED B3447. src/lib/audit-primitives/chorus-mass.ts — pure computeChorusMassReport(lyric) parses every [Chorus]/[Chorus N]/[Final Chorus] block (excludes [Pre-Chorus]), counts words per line, and returns verdict anthemic (≥1 chorus has a line ≤8 words — a short repeatable hook) / dense (every chorus line is long — flagged, the feedback's "write one that's just 4-8 words repeated") / no-chorus, plus a dossier-ready label ("Chorus mass: anthemic" / "Chorus mass: dense"). Flag-only, never pass/fail (dense is a valid craft choice). 11 tests. Dossier/dashboard surfacing deferred as the consumer step (mirrors the AVD B3376 precedent — primitive ships first; label is wired ready).
  • R5-A23 — Named-event bridge audit. [AUTONOMOUS · P2 · ~1 build] SHIPPED B3447. src/lib/audit-primitives/bridge-anchor.ts — pure auditBridgeAnchor(lyric) reuses extractBridge() (B3128) and scans each [Bridge] for an anchor: a digit (room 314), a calendar term (month/weekday), a spelled number (forty-three, hyphen-split), or a mid-line proper noun (excludes line-initial caps + "I"). Returns anchored / unanchored / no-bridge + anchorTypes + matchedTerms + dossier-ready label. Flag-only (the spec said "fail"; shipped as a flag to match the sibling primitives' non-gating discipline). Complements audit-bridge-completeness (development) with the orthogonal anchoring signal. 8 tests — incl. a discrimination guard that surfaced a real finding: the country / pop / indie OSNG bridges are genuinely unanchored (abstract — "He taught me everything / So I could choose"), exactly the "no named wound" failure this primitive names.

Genre Catalog Clustering WAR ROOM (2026-05-28 — B3409)

Triggered by operator screenshots showing the /genres/indie/prompts page rendering 83 indie prompts, essentially every one medical/health themed. Full WAR ROOM doc: docs/WAR-ROOM-GENRE-CATALOG-CLUSTERING-2026-05-28.md. Smoking gun: indie's mustInclude array in src/lib/genre-prompt-catalog/arc-markers.ts literally contained 'medical term' as a vocabulary anchor — Sonnet dutifully used it in ~84% of generated prompts.

  • R5-A24a — Remove "medical term" from indie arc-markers. [AUTONOMOUS · P0 · 5 min] Shipped at B3409 (2026-05-28). The literal string 'medical term' removed from src/lib/genre-prompt-catalog/arc-markers.ts's INDIE_PACKAGE.mustInclude. Inline comment explains the WAR ROOM finding. Partial fix — the full architectural fix is R5-A24.
  • R5-A24b — Indie catalog hand-rebuild (proof). [AUTONOMOUS · P0 · 1 build] Shipped at B3409. src/data/genre-prompt-catalog/indie.json rewritten with 25 hand-curated VARIED prompts demonstrating the new variation discipline. Spans 9 distinct settings (laundromat, drive-through, farmers' market, library carrel, wedding, apartment, pottery class, strike line, antiques mall, garage workshop, bus, kitchen tutoring, reunion, garage echo, library volunteer, voicemail, birthday, lake run, DJ booth, grocery, lifeguard, storytime, loading dock, garden hose) + 9 distinct moods. Cross-cluster on time (dawn/morning/midday/afternoon/evening/night) + persona (1st-person / observer / addressed-other / collective). The 83-prompt count drops to 25; richness increases.
  • R5-A24 — Split `mustInclude` into `laneMarkers` + `sceneRotationPool`. [AUTONOMOUS · P1 · ~1 build] Shipped at B3410 (2026-05-28). ArcCraftPackage now requires both laneMarkers (genre lane lock — production / artist / substyle / craft) AND sceneRotationPool (≥30 diverse scenes per arc) as first-class fields. All 9 arcs populated. checkPromptCoherence() switched to prefer laneMarkers (mustInclude retained as legacy fallback). Catalog ratchet rerun: all 9 arcs above floor (864/865 prompts pass — only 1 country prompt dropped). 5 new tests cover the split contract (lane non-empty, scene ≥30, lane/scene no-overlap, lane/avoid no-overlap, B3410 architectural-split landed).
  • R5-A25 — Generator rewrite for sceneRotation discipline. [AUTONOMOUS · P1 · ~1 build] Shipped at B3410 (folded into R5-A24). scripts/generate-genre-prompt-catalog.ts now renders both laneMarkers (≥1 per prompt) and sceneRotationPool (1 per prompt with rotation discipline + ≤8% catalog domination cap) as separate concerns. Required-fields contract in the Sonnet prompt reads "MUST include ≥1 LANE MARKER + MUST anchor on ≥1 SCENE ROTATION POOL entry" rather than the conflated B3095 phrasing. Future regens (R5-A26-A30) will use the corrected prompt.
  • R5-A26 — Regenerate country.json with new generator. [AUTONOMOUS · OPERATOR-BUDGET · ~$0.50 + 10 min Sonnet] Shipped at B3411 (2026-05-28). 62 fresh prompts at 100% coherence. Dominant-axis (honky-tonk/pickup/dirt-road) clustering rate dropped 94% → 55% (-39pp). New titles span domestic / transit / work / public / nature / relationship / memory / artist scenes evenly — Granddad's Last Sunday, Highway 12 at Midnight, Mama's Recipe Box, Carolina Moon Radio, Tractor Pull Saturday, Dolly on the Dashboard, etc. Honky-tonk cluster broken.
  • R5-A27 — Regenerate pop.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 57 fresh prompts at 100% coherence. Clustering rate 31% → 19% (-12pp). New scenes span domestic/transit/work — Walk-In Closet Runway, Airport Security Liminal, Closing Shift Countdown, Hostess Stand Power.
  • R5-A28 — Regenerate rap.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 79 fresh prompts at 100% coherence. Clustering rate 67% → 52% (-15pp). 2AM Kitchen Confessions, Mom's Couch Chronicles, Barbershop Politics, Subway Platform Soliloquy, Dispatch Desk Drama — domestic/transit/work scenes replace corner-store/cypher cluster.
  • R5-A29 — Regenerate rnb.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 73 fresh prompts at 100% coherence. Artist namedrops (D'Angelo/Sade/Frank Ocean/Aaliyah) GONE from titles; scenes span Kitchen Sink Confessions, Bathtub Sanctuary, Balcony Summer Storm, Hotel Lobby 1AM, Tuesday Jazz Club Healing, After-Hours Diner Dreams. Artist + production tropes remain in prompt bodies as lane locks, scenes diversified.
  • R5-A30 — Regenerate folk + rock + latin + worship JSONs. [AUTONOMOUS · OPERATOR-BUDGET · ~$2 + 40 min Sonnet] All four shipped at B3411 (single batch run, ~21 min wall time, ~$2 Sonnet total). Folk: 63 prompts, 86% → 84% (Stanza Study / Mechanics / Workshop instructional vocabulary GONE; replaced by Greyhound to Memphis, Lumber Mill Lament, Cannery Line Chronicles, Dishroom Midnight). Rock: 76 prompts, 36% → 32%, Basement Archaeology, Power Chord Democracy, Warehouse Cathedral, Mill Town Saturday joining the highway/Rust Belt anchors. Latin: 90 prompts, 21% → 7% (Abuela's Kitchen Memory, Papi's Barbershop Chronicles, Hermana's Quinceañera Dance — family scenes added to place-name anchors). Worship: 89 prompts, 89% → 80% (Sick Room Vigil, Prison Ministry Hope, Hospital Chaplain Visit, Kitchen Morning Prayer added private/domestic/public-service scenes).
  • R5-A31 — SA#33 ratification. [AUTONOMOUS · P3 · 1 build] Shipped at B3415. SA#33 added to src/lib/brand.ts (BRAND.sacredAccident33) + src/lib/sacred-accidents/data.ts + docs/SACRED-ACCIDENTS.md. Ratified on operator's "Do your recommendation" authorization once B3411 empirical evidence (all 9 catalogs at 100% coherence post-architectural-split) had landed. Companion to SA#16 (output-side motif saturation) and SA#29 (genre signal as STATE).

Concrete-to-Abstract Drift WAR ROOM (2026-05-28 — B3412)

Triggered by a multi-AI craft critique of three real songs ("Speed Limit Thirty-Five" / "Coffee Stains and Crooked Lipstick" / "What I Keep"). The reviewer's central diagnosis: "The system observes brilliantly in verses, then thesis-summarizes in choruses and bridges." Full WAR ROOM doc: docs/WAR-ROOM-CONCRETE-TO-ABSTRACT-DRIFT-2026-05-28.md. 10 inline tier-2 banned-phrase nudges shipped at B3412. SA#34 candidate ("The chorus is INHABITED, not DECLARED") recorded.

  • R5-B0 — 10 inline tier-2 banned-phrase nudges. [AUTONOMOUS · P0 · 1 build] Shipped at B3412. Added 10 entries to src/lib/banned-terms.ts under house_style tier 2 covering: §2.1 thesis-line shapes (what if this is / this is everything / couldn't make me / i'm done rehearsing), §2.2 told-not-shown patterns (rationing my / more than remember(-ing) / instead of remember(-ing)), §2.3 mixed-metaphor failures (the bruise you / survival over). Post-generation scan catches these; gauntlet/refine pass rewrites them.
  • R5-B1 — ATL primitive (Abstract-Thesis-Line detector). [AUTONOMOUS · P0 · ~3 builds] Shipped at B3413 (heuristic stage). src/lib/claude/audit-abstract-thesis-line.ts — pure function with curated THESIS_OPENERS lexicon (~25 entries) + concreteness ratio (reuses R&B NCD lexicons) + two flag rules (thesis-opener with low concreteness OR lone abstraction). 10 unit tests pass. Inverse polarity from NCD: HIGHER score = MORE thesis-mode (BAD). Phase 2 (Haiku judge for adjudication) queued as R5-B1.2.
  • R5-B2 — STDD (Show-don't-Tell Density) — generalize R&B NCD/AID rule cross-arc. [AUTONOMOUS · P0 · ~2 builds] Shipped at B3413. src/lib/claude/audit-show-dont-tell-density.ts — runs the AID rule on chorus + bridge sections regardless of arc. Single cross-arc threshold STDD_FLOOR = 0.30. Same polarity as NCD (higher = more showing = good). Per-section diagnostic with weak-line indices. 11 unit tests pass. R&B's NCD weight stays in R&B's arc-specific domain.
  • R5-B3 — EFI primitive (Emotional Friction Index). [AUTONOMOUS · P0 · ~2 builds] Shipped at B3414. src/lib/claude/audit-emotional-friction-index.ts — Haiku-judged with 4-tier verdict (one-note / hinted / resolved / layered). Parser auto-downgrades layered/resolved verdicts that ship with empty evidence OR null secondary — catches Haiku's tendency to ratify without backing. Fail-safe: timeout/parse error returns state='error' + score 100 (null component). 15 unit tests pass.
  • R5-B4 — SA#34 ratification ("The chorus is INHABITED, not DECLARED"). [AUTONOMOUS · P1 · 1 build] Shipped at B3415. SA#34 added to src/lib/brand.ts (BRAND.sacredAccident34) + src/lib/sacred-accidents/data.ts + docs/SACRED-ACCIDENTS.md. Companion ratification: SA#33 (lane-vs-scene split, R5-A31) shipped in the same build since its empirical evidence (B3411 catalog regen) had already landed. Two ratifications closed simultaneously.
  • R5-B5 — IRD primitive (Internal Rhyme Density) promoted cross-arc. [AUTONOMOUS · P1 · ~2 builds] Shipped at B3417. src/lib/claude/audit-internal-rhyme-density.ts — scans each chorus block for: (a) end-rhyme density (pairwise terminal-vowel-bucket matches), (b) internal-rhyme rate (within-line repeated vowels), (c) slant-rhyme bonus (vowel match + consonant tail mismatch). 7-bucket vowel-phoneme proxy approximates without CMU dict (~75-85% accuracy on common English lyric vocabulary). Composite block score weights end-rhyme 60%, internal 25%, slant 15%. 4-tier verdict (flat / modest / engineered / sonic-locked). 9 unit tests pass. Closes §2.4 + §2.5.
  • R5-B6 — Forge chorus pre-declaration rule (generalized SA#21 / SA#34). [AUTONOMOUS · P1 · 1 build] Shipped at B3419. src/lib/claude/chorus-discipline-baseline.ts — CHORUS_DISCIPLINE_BASELINE constant + isChorusDisciplineBaselineEnabled gate. Wired into single-phase-prompt.ts BEFORE the per-mode chorusDiscipline injection so per-mode rules layer on top as specialization. Two-axis rule: AXIS A sonic (open-vowel chorus peak) + AXIS B semantic (named object from verses). Banned thesis-mode shapes from B3412 WAR ROOM (Every X / What if X / I'm done X-ing / X more than Y / X over Y / self-help-book gut check). Behind SF_FORGE_CHORUS_DISCIPLINE_DISABLED=1 kill-switch with warn-log on every forge call when set (mirrors SA#29 pattern). Prompt budget: 89% of 20k tokens (~2k headroom). 9 unit tests pass.
  • R5-B-baseline — Vault chorus-baseline measurement script. [AUTONOMOUS · P1 · 1 build] Shipped at B3420. scripts/bench-chorus-baseline.ts + npm run bench:chorus-baseline. Queries top-N songs by forge_score, runs the 5 pure-function primitives (ATL+STDD+IRD+CVC+PLV) on each, aggregates + writes JSONL. Pre-B3419 baseline captured against vault top-100 (mean forge_score 91.2): ATL 1.8, STDD 5.8 (69% below floor — known R&B-lexicon limitation), IRD 69.3 (engineered tier), CVC 81.4 (evolved/alive dominant), PLV 72.6. Full analysis + post-deploy comparison plan in docs/CHORUS-BASELINE-B3419.md. After 24-48h of post-B3419 traffic, re-run with --since 2026-05-29 --label post-b3419 to measure the rule's empirical impact.
  • R5-B-lexicon — STDD cross-arc lexicon expansion. [AUTONOMOUS · P2 · ~2 builds] Known followup surfaced by B3420 baseline. The R&B NCD lexicon (CONCRETE_NOUNS / ACTIVE_VERBS / ABSTRACT_WORDS) was curated narrowly for R&B vocabulary. Top-100 baseline shows 69% of songs below the STDD floor — not because choruses are abstract, but because country/folk/indie vocabulary (mailboxes, fishing boats, sweaters, etc.) doesn't hit the lexicon. Expand the shared lexicons to cover the 9 arcs evenly. Re-baseline after expansion to confirm the below-floor rate drops without false-positive reduction in scoring sensitivity.
  • R5-B7 — MCRT primitive (Metaphor Close-Reading Test). [AUTONOMOUS · P1 · ~2 builds] Shipped at B3418. src/lib/claude/audit-metaphor-close-reading.ts — Haiku-judged 4-dimension grading (clarity / consistency / weight / truth) for each identified metaphor line. 4-tier verdict (sonic-only / partial / felt / load-bearing). Parser enforces consistency rule: load-bearing requires zero failed lines (auto-downgrades to felt + score≤75 if violated). Per-failed-line diagnostic includes specific failed dimensions + close-reading critique. 15 unit tests pass. Same fail-safe pattern as EFI + through-line.
  • R5-B8 — Soul-pop substyle audit + per-substyle forge fragment. [AUTONOMOUS · P1 · ~2 builds] Closes §2.9. Audit current pop substyle profiles to verify soul-pop carries: tight 4-8 syllable phrases, percussive consonant placement, gospel-chord harmonic-motion cue, chorus hook chain repeats 3-4× per chorus. Add substyle-detection-aware forge fragment that injects right discipline per-substyle.
  • R5-B9 — CVC primitive (Chorus Variation Cost). [AUTONOMOUS · P1 · 1 build] Shipped at B3416. src/lib/claude/audit-chorus-variation-cost.ts — parses distinct chorus blocks (Chorus / Hook / Refrain), runs pairwise line-identity rate (case-insensitive, whitespace-normalized), 4-tier verdict (static / thin / evolved / alive). Single-chorus songs return score=100 (no penalty). 9 unit tests pass. Inverse polarity: HIGHER = MORE variation = good.
  • R5-B10 — PLV primitive (Phrase Length Variance). [AUTONOMOUS · P2 · 1 build] Shipped at B3416. src/lib/claude/audit-phrase-length-variance.ts — per-section word-count coefficient-of-variation with target bands: verse [0.15, 0.70] (Pat Pattison asymmetric stanzas), chorus [0, 0.30] (singability discipline), bridge unbounded. Word count is a syllable proxy (CMU-dict-style phonetic lookup deferred to PLV phase 2). 8 unit tests pass. Flags include diagnostic clauses ("verse lines vary widely — prose-like, may rush on melody" / "chorus phrases vary widely — singability suffers").
  • R5-B11 — AR primitive (Ambiguity Reserve). [AUTONOMOUS · P2 · ~2 builds] Closes §2.8. Haiku-judged: does the final chorus / outro leave at least one unresolved element (a lingering question, a visible bruise the song doesn't fix, a contradiction the song acknowledges but doesn't close)? Resolution-too-tidy is the failure mode.

Round-5 sequencing notes

Ship R5-A1 + R5-A2 + R5-A3 + R5-A4 in a single bounded session (~1.5h total). All four are <1h, all four are P0, and shipping them as a single push moves the system out of the "audit findings open" state cleanly. R5-A14 (the N=large eval) is the highest-leverage Tier-3 item because it's the only thing that converts the 12-build J-phase from "we believe this works" to "we measured this works" — that single measurement unblocks the public claim moves R5-A11, R5-A16, R5-A18.

The most surprising audit finding: the brand-promise gap from B3372 WAR ROOM §7 is STILL UNADDRESSED despite 12 builds of register-awareness work shipping. R5-A4 is the cheap fix; R5-A11 is the substantive one. Both are necessary; neither is sufficient.


Deep Audit 2026-05-24 round 4 — autonomous queue (TOP PRIORITY)

Triggered by: the B3206 Deep Audit, captured at the end of the 13-build session (B3193 → B3205). Final grade 8.4/10 (B+). Strongest categories: Trust (9.3), Process maturity (9.5), Strategic differentiation (9.2). Weakest: Retention (6.5 → still climbing), Information architecture (7.0), Onboarding (7.2).

Operator-side complements live in docs/OPERATOR-PUNCH-LIST.md under the matching "Audit-Round-4" section. The engineering items below ship autonomously.

Round-4 Tier 1 (ship this week — small, each ≤2h)

  • R4-A1 — `npm run docs:refresh` to close the 85-commit-stale-data leak on /engineering. Operator ran this in the same session; committed at 67b3f7ae0 alongside B3206. Public surface no longer lies about freshness.
  • R4-A2 — `npm pack` + tarball-install smoke-test ratchet. Shipped B3209 (2026-05-24): scripts/check-npm-package.ts walks publishable packages under packages/*, runs npm run build + npm pack, installs the tarball into a mkdtemp'd consumer project, and dynamically imports the main entry — assertion: import('@songforgeai/<name>') must resolve to a non-null module on Node 20+. All 4 publishable packages (@songforgeai/agent-room@0.1.0, @songforgeai/fidelity-standard@0.2.1, @songforgeai/scoring-rubric@1.2.0, @songforgeai/client@0.4.0) survive the smoke test. Wired into check:all between reduced-motion and moat-ledger-health. Catches the B3197 class-of-bug (works-in-monorepo, breaks-once-installed) AND the B3204 class-of-bug (build artifacts not actually emitted). Failure detail captures last 20 lines of consumer-side stderr so the offender's stack is in the CI log.
  • R4-A3 — Surface auto-audit + email-gate + finalize success rates on `/admin/system-health`. Shipped B3210 (2026-05-24): added SurfaceTelemetryReport to /api/admin/system-health + a 3-card row on the admin page covering: (1) Forge auto-audit — last-1h + 24h counts from song_audit_runs + mean 24h composite; (2) Surgery finalize — last-1h + 24h complete-status counts from song_surgery_sessions + 7d abandon-rate proxy; (3) Email gate — RESEND_API_KEY presence + SF_EMAIL_DISABLED kill-switch + SF_EMAIL_RETENTION_ENABLED flag with live/dev-only/disabled glance status. Service-role client used for the cross-user counts; each sub-query catches its own errors + falls back to status='unknown' so the page never fails to render. Closes the "currently visible only by grepping production logs" gap the audit named. Note: email.send per-kind count remains pending an email_sends table — surfaced as the gate state for now (live volume is low; structured logs are still the source of truth for per-send accounting).
  • R4-A4 — Reduced-motion ratchet cleanup (124 → 104 baseline). B3208 (2026-05-24) migrated 14 components to the motion-safe: prefix: RepairFlow ×5 + EmailCaptureCard + HomepageForgeDemo + Navbar + NewsletterCapture + RadioBar + RadioPanel + ReferralCard + SamplePlayer + SectionRewriteModal + SparkIdeas + StickyCTA + VoiceConsistencyCard + VoiceRewriteFlow ×2 + genres/GenreRadio. Net: 10 unguarded animation sites eliminated; ceiling tightened 114 → 104. Headroom: 0 (any new unguarded animation now fails CI). End-of-quarter target: < 50; end-of-year: 0.

Round-4 Tier 2 (ship this month — bigger, each <1 day)

  • [~] R4-A5 — `RepairFlow.test.tsx` foundation: 8 state-mutation paths. Foundation shipped B3211 (2026-05-24): 11 tests cover the simple paths + harness. Initial render (1+items / empty/Session complete), per-item card structure (testids + button copy), progress strip ("Issue 1 of 3"), Skip path advance, Reject path advance, Skip×3 → "Session complete", resume-banner visibility (set / null / start-fresh dismissal with async fetch await), severity border classes (red/orange). Harness uses vi.stubGlobal('fetch', mock) + a default-OK envelope so unscripted endpoints (init, edit-log, finalize) don't blow up. Still pending for full R4-A5 closure: Accept happy path (judge-edit verdict assertion) / Accept with judgeError fallback / Undo-to-Edit (B3192) / Undo-and-Skip (B3192) / Revert-to-baseline / Resume button API contract. These 6 paths each need 1-2 tests + scripted fetch mock for the corresponding endpoint; estimate ~2-3h of follow-on work. Marked [~] (in-progress) rather than [x] so the remaining paths stay visible.
  • R4-A6 — Surgery WAR ROOM remaining P0 fixes: #2, #4, #5. Shipped B3213 (2026-05-24): all three P0s closed in one bounded build. (P0 #2) RepairFlow.tsx resume-hydration now sets judgeError on ANY non-canonical verdict — pre-3213 only e.verdict === 'unjudged' set the marker, so a forward-compat verdict like 'neutral' or 'mixed' would silently fall back to 'changes' + corrupt netDimensionScore (aggregateSessionDeltas correctly skips judgeError'd entries but skipped the 'neutral'-treated-as-'changes' case). (P0 #4) Added liveEditIndexRef = useRef<number>(0) + bumpEditIndex(next) helper; replaced all 4 liveEditIndex + 1 closure reads with liveEditIndexRef.current + 1 (the rapid Accept→Undo-and-Skip race where both writes observed the same stale state). All setLiveEditIndex call sites swapped to bumpEditIndex to keep state + ref in sync. (P0 #5) Hardened the seenItemIds.add() guard to require non-empty string AND non-null (defense-in-depth — even though the existing if (e.surgeryItemId) already covered the failure mode). Verification: typecheck clean, 104/104 surgery + lib tests green. Audit doc unchanged (B3196 historical record); follow-on Surgery WAR ROOM items (P1/P2) tracked separately if/when they surface.
  • R4-A7 — Postgres BEFORE UPDATE trigger on `song_surgery_sessions.status` transitions. Shipped B3214 (2026-05-24): supabase/migrations/enforce_surgery_session_status_transitions.sql. Defines public.enforce_song_surgery_session_status_transition() function + attaches it as a BEFORE UPDATE OF status trigger on public.song_surgery_sessions. The function rejects any change that isn't in the allowed-transitions map (kept in sync with src/lib/surgery/session-types.ts:ALLOWED_TRANSITIONS); no-op updates (NEW.status = OLD.status) pass through unchanged so other column writes (diagnostic_snapshot / verification_record / session_note) don't get blocked. Terminal-state protection: any UPDATE attempting to leave 'complete' or 'abandoned' raises check_violation with a HINT pointing at the canonical map file. The trigger closes the canTransition() TOCTOU race window B3199 surfaced — two concurrent UPDATEs that both pass the app-layer SELECT-then-check can no longer both succeed; the DB rejects the second one atomically. Implemented as CHECK-style trigger (not a column CHECK constraint) because column constraints can't reference OLD.* values to enforce direction. B3214 sentinel comment + COMMENT ON FUNCTION for discoverability via git grep B3214.
  • R4-A8 — Auto-`docs:refresh` cron + auto-PR. Shipped B3212 (2026-05-24): .github/workflows/docs-refresh-nightly.yml runs at 09:15 UTC daily + on manual workflow_dispatch. Checks out main with full git history, runs npm run docs:refresh, then uses peter-evans/create-pull-request@v6 to open (or update in-place) a PR titled chore: nightly docs:refresh — sync auto-state + tighten ratchets on the canonical branch chore/docs-refresh-nightly. PR body lists the two classes of artifact regenerated (CLAUDE.md auto-state block + the 6 count-based ratchet ceilings: console-log, console-error, explicit-any, design-tokens, launderer-casts, reduced-motion) plus loop-guard documentation. Direct cure for Sacred Accident #18 ("the cadence ritual itself drifting") — even if every operator-side cadence slips, the cron catches the drift within 24h. PR-rather-than-direct-push because docs:refresh tightens enforcement gates (ratchets), and the operator deserves a single review point before the gate gets stricter.
  • *R4-A9 — `/admin/feature-flags` runtime panel — consolidate the SF__DISABLED env toggles. ✅ Shipped B3286 (2026-05-25). The audit had estimated "~10"; the codebase grep found 22* `SF__DISABLED / SF_*_ENABLED env-var feature toggles scattered across forge / refine / gauntlet / finalize-song / preforge / post-forge-pipeline / leaderboard / various critic loops. All 22 now ship registry entries in src/lib/feature-flags.ts with category (forge / refine-gauntlet / eval-audit / leaderboard / email / experimental / admin) + loadBearing flag + warning string naming the failure class each kill switch unlocks (SA#29 for the artist locks). New /admin/feature-flags` page surfaces all 24 entries grouped by category + per-flag "Flip in Vercel env" deeplink. Pre-existing 2-entry registry consolidated under the same shape. Closes the env-toggle-proliferation 5X Plan kill.

Round-4 Tier 3 (next quarter — multi-build)

  • R4-A10 — VAULT-OPEN-4 (per-user vault graduation). 4-6 builds. Biggest remaining moat play. Scope song_audit_runs + song_embeddings + exemplar_status by user_id; gate dashboards per-user; drop admin-only requireAdmin checks in favor of per-user RLS. Requires a per-user audit-on-forge hook (B3203 already wires this — extend the trigger to non-admin users) + a per-user auto-curate when N exemplars accumulate.
  • R4-A11 — Crucible-first homepage flow. The 5X Plan's #1 move. The Crucible is the only thing in the system nobody else can copy. Currently it's section 1 of the homepage but the primary "Open the forge" CTA leads to auth wall. Restructure: Crucible-as-hero with the live paste-box above the fold; everything else (forge demo, leaderboard, before/after) becomes secondary proof. ~6-8h of UX work + copy iteration.
  • R4-A12 — Multi-language Phase 1 (Spanish first). #42 from Tier-C list; operator trigger phrase Continue Multi-Language Phase 1. The Latin genre arc work (L1-L13) is the natural launchpad; the dialect-coherence audit primitive (DCS) extends to language-coherence. RFC-0009 covers the seal-annotation contract.
  • R4-A13 — Resumable forges Phase 2 (durable event log). #23 from Tier-B list. Migration shipped at B2887; persistence layer at src/lib/forge-events/persistence.ts shipped. Two builds remain: (a) wire writes from the live forge SSE stream into forge_events with batched inserts; (b) extend /api/songs/forge/resume to read from forge_events with ?lastEventId=N for byte-identical replay. Closes the "SSE disconnect loses commentary" gap.
  • R4-A14 — Kill duplicate IA between /genres and /lyrics. 5X Plan kill #2. Partially done at B3056 (Footer-link IA consolidation). Remaining work: audit which /lyrics/[slug] pages are still useful as SEO long-tail vs. redundant with /genres/[slug]. ~3-4h grep + decisions + redirects.

Round-4 Tier 4 (category-defining engineering bets)

  • R4-A15 — "Trust Engine" public dashboard. Every score on the site queryable by external auditors via the npm SDK + the seal. New page at /trust that lets anyone enter a seal {rubricVersion, model, temperature, buildSha, build} + verify it. Currently the /verify page does this for ONE score at a time; the Trust Engine surface is the operator-facing query interface. ~1-2 builds.
  • R4-A16 — Per-user voice fingerprints as a product surface (#43 Phase 2). Currently scaffolded. Turning this into a real "your voice over time" product surface creates the kind of personal-history-stickiness Strava/Duolingo trade on. Depends on VAULT-OPEN-4 + Hit Calibration Corpus accumulating N≥10 entries per user. ~3-4 builds.
  • R4-A17 — Sacred Accidents public publish surface. ✅ Shipped B3285 (2026-05-25). /sacred-accidents/[slug] per-SA permalink route + per-SA OG card route + JSON-LD Article structured data + prev/next navigation + citation footer linking to BRAND.sacredAccident{N} + the canonical doc on GitHub. Reads from the existing typed src/lib/sacred-accidents/data.ts mirror (no duplicate source of truth); new accidentSlug() + parseAccidentSlug() helpers in the data module with round-trip test coverage (8 new tests). Index page links each entry to its permalink via the "Permalink →" affordance. The publishing CADENCE remains operator-side (META-R4-1 in docs/OPERATOR-PUNCH-LIST.md) — engineering side is closed.

Round-4 sequencing notes

The Tier-1 items above (R4-A2 / R4-A3 / R4-A4) should ship FIRST — they're cheap and they close the gaps the audit explicitly flagged. R4-A5 (RepairFlow tests) is the biggest single-build win. R4-A11 (Crucible-first homepage) is the highest-leverage UX move but requires coordinated copy + design iteration; queue it AFTER the operator's R4-T2-1 work on Brett's case study (so the new homepage flow can incorporate the new case-study artifact).


Deep Audit 2026-05-19 round 3 — autonomous queue (TOP PRIORITY)

Triggered by: the post-B2746 Deep Audit. The audit found engineering + discipline at A+, distribution + external validation at C. The 15 items below convert audit findings into shippable engineering work; the operator-side complements (arXiv submission, npm publish decision, partnership outreach, testimonial permission, hit-corpus curation) live in docs/OPERATOR-PUNCH-LIST.md under the new "Audit 2026-05-19 round 3" section.

Sequencing principle from the audit: every Tier-1 item is correctly an engineering ratchet. Ship in priority order P0 → P1 → P2 → P3 — but skip ahead when a P1 unblocks a P0 (e.g. prosody backfill is P2 but unblocks user-trust for songs forged before B2746, so promote when convenient).

P0 — ship this week (audit Tier 1)

  • AUDIT-1 (B3197) — Email capture on Crucible verdict screen shipped via Resend. Operator picked Resend (already wired for the Life Songs / Heirlooms flow at B2591 — same verified noreply@songforgeai.com sender + same RESEND_API_KEY env var). NEW src/lib/email/send.ts — extracted shared sender (sendTransactionalEmail, isValidEmailShape, escapeHtml, baseEmailHtml) so the Crucible-capture + retention surfaces don't duplicate the life-songs envelope. Fail-open contract; SF_EMAIL_DISABLED=1 global kill switch; structured-log events email.send.{sent,disabled,no_api_key,error,threw}. 21 unit tests pin the env-toggle precedence + the validator + the HTML escape. NEW supabase/migrations/add_crucible_email_captures.sql — citext-typed email column (case-insensitive UNIQUE), verdict_banner denormalized for fast segmentation, ip_hash + user_agent_hash for audit trail (NEVER raw PII), opted_in_to_marketing separate from capture so future unsubscribe flow can flip the flag without losing the row, RLS enabled with default-deny (service-role-only writes via the capture endpoint, admin reads only). NEW POST /api/crucible/[id]/capture-email — unauthed, IP-rate-limited via new RATE_LIMITS.crucibleCapture bucket (30/IP/day; entry added alongside the existing crucible + crucibleSave entries in src/lib/api-auth.ts), validates email shape, upserts ON CONFLICT (email) so re-capture is idempotent + sets alreadyCaptured: true, optionally mails the verdict back to the user via the new sender (fail-open — DB persistence is the load-bearing path). NEW src/components/crucible/EmailCaptureCard.tsx — client component mounted on /c/[id] between VerdictShareButtons and the CTA cluster. 4-state form (idle/submitting/captured/error) with two explicit checkboxes ("Email me the verdict" + "Add me to the mailing list"); maps every endpoint error code to operator-friendly copy; green confirmation card on success; full unsubscribe microcopy. The audit's #1 highest-leverage move — turns unauthed Crucible traffic (high-volume, free) into a list the operator can reach for moments-of-decision outreach (rubric updates, v1.0.0 fidelity-standard release, price changes, new genre arcs). Operator-side todo: apply the migration in Supabase.
  • AUDIT-8 (B3197) — Email retention loop scaffold (4 transactional templates), feature-flagged. NEW src/lib/email/retention-templates.ts ships the four canonical lifecycle moments per the Deep Audit round-3 finding (retention gap is load-bearing): sendWelcomeEmail(profile) (post-signup), sendFirstForgeEmail(profile, song) (first finished forge), sendWeeklyLeaderboardEmail(profile, snapshot) (Mondays when rank moved), sendDormant30dEmail(profile) (30 days idle). Every template gated by SF_EMAIL_RETENTION_ENABLED=1 — defaults OFF so the templates can ship behind a flag while the operator (a) sets up Resend suppression lists, (b) finalizes unsubscribe-link copy, (c) audits real samples before flipping. When the flag is unset, each function logs a structured "gate_closed" event the operator can grep for. Templates use baseEmailHtml + escapeHtml + the shared sender; movement copy in the leaderboard template branches on first-touch / moved-up / dropped / held; first-forge template includes the rubric composite score when present, omits when null. Each template carries unsubscribe microcopy. 12 unit tests pin: gate closes by default; SF_EMAIL_RETENTION_ENABLED=0 also closes; each template uses its own kind tag; first-forge omits score when null; weekly leaderboard branches correctly across all 4 movement cases; greeting falls back to "Hi there," when displayName is null. The TRIGGER wiring (auth/callback hook for welcome, finalize-song hook for first-forge, /api/cron/weekly-leaderboard-touchpoint, /api/cron/dormant-30d-touchpoint) lands in follow-up builds once the operator approves the template copy.
  • AUDIT-2 — Refine route prosody recompute (sibling to B2746). B2749 — shipped; checkbox flipped at B2872. PATCH /api/songs/[id] now mirrors the B2746 gauntlet fix: when the body carries new lyrics, scoreProsody is recomputed before persisting + the fresh report ships in the update payload. SF_PROSODY_REPORT_DISABLED=1 env gate honored, safeSync wraps the call, null/zero-line reports leave the column untouched.
  • AUDIT-3 — Wire B2743 (motif ledger) + B2744 (bridge architecture) telemetry into /admin/forge-discipline. B2753 — shipped. Two new tiles on /admin/forge-discipline: (a) motif cluster rate (% of songs with 2+ watched motifs from the same hot-list the B2743 ledger watches), target ≤20%, baseline 80% from the May 17 audit; (b) bridge old-default rate (% of songs with co-occurring whispered + crack directives — the audit's #1 system tell), target ≤30%, baseline 93/100. Top 5 motifs by song-presence rate also surface so regressions show by name. The Wave 2 fix is now measured.
  • AUDIT-4 — Hero compression: 1 CTA + receipts strip. B2755 — shipped. Hero already had 1 primary CTA (HeroCruciblePaste; B2106 closed the dual-CTA dilution previously). Added 4-pill receipts strip above the fold between the privacy note and the demote-CTAs: Public rubric · Signed verdicts · Incident log · The receipts. Each pill links to a real public artifact (/scoring/standard, /verify, /incidents, /the-receipts). The duplicate chip strip 3 sections down was removed to avoid double-rendering.

P1 — ship this month (audit Tier 2 — high-leverage)

  • AUDIT-5 — Wire B2745 (chorus-compression) UI into dashboard SongDetail. B2757 — shipped. New ChorusCompressionPanel mounts in SongDetail right after HookCompressionPanel; on-demand button generates poetic / plainspoken / commercial register variants + recommendation via a new POST /api/songs/[id]/chorus-compression route. Pro-tier-gated (402 + upgrade CTA for free/creator); 30/hour rate limit; copy-to-clipboard per variant. 422 with actionable detail when the song has no [CHORUS] marker. Sibling pattern to hook-compression but user-initiated (no auto-compute at forge time — keeps Haiku cost off the default path).
  • AUDIT-6 — Annual-billing default on /pricing. Audit false-positive — already shipped at B1901. The audit's WebFetch saw the SSR'd billingInterval state but couldn't tell which value was the default. Source confirms useState<'month' | 'year'>('year') at src/app/pricing/page.tsx:97. B1901's commit comment names the same audit-driven motivation: "Linear, Notion, Slack all default to annual." Validated at B2750.
  • AUDIT-7 — Cut homepage to 7 sections. B2755 — shipped. 4 sections cut: (1) Mini-Demo / HomepageForgeDemo (hero already ships HeroCrucibleDemo — dual-demo dilution); (2) Perspective Atlas + Persona Forge teaser (/atlas + /persona-forge are dedicated surfaces reachable from Navbar + footer); (3) "What we don't do" / The Disciplines We Don't Break (content lives at /manifesto + /principles/topic-sovereignty); (4) Homepage FAQ accordion (/pricing FAQ is the single source of truth; FAQPage JSON-LD still ships from the top of page.tsx so SERP rich-snippets aren't affected). Net section count: 8 → 5 raw sections (+ HomepageBeforeAfter inline = 6 logical sections). File: 837 → 712 LOC (-125 lines).
  • AUDIT-8 — Email retention loop scaffold (4 transactional templates). Welcome / first-forge / weekly-leaderboard / dormant-30d. Templates + send infrastructure ready behind a feature flag; the operator decides when to flip live (OPERATOR list A11). Per audit, the retention gap is the load-bearing missing surface.
  • AUDIT-9 — Non-English fairness fix OR formal defer. B2751 — formal deferral landed. Per the round-3 brief, this item required ship-or-defer disposition. After review: Phase A1-A3 (forge-prompt fairness fixes) already shipped at B1982-B1993; Phase C2/C5 (corpus-anchor add + 68-song rescore) is operator-side work that needs anchor curation before the rescore can run. New hard deadline 2026-06-30 logged in docs/TRUST-DECAY-AUDITS.md. The disclosure post ships with the gap-closure number measured, not as a "working on it" half-trust signal. If C2+C5 aren't landed by 2026-06-30 the disclosure post ships unconditionally with the gap still open.
  • AUDIT-10 — /case-study/brett page (skeleton + draft). Working-songwriter case study; Brett ran the entire P0+P1 sweep through the product over 24-48 hours and surfaced 12 actionable findings. Real external validation. Page ships behind a draft-flag pending operator permission (OPERATOR list A6 → A12); the engineering work is the page + structured-data + cross-links.
  • AUDIT-11 — /case-study index + 2 more songwriter slots reserved. B2754 — shipped. Sibling format to /case-studies (quantitative lyric-pipeline). New /case-study index lists visible entries + 2 reserved-slot placeholders ("different genre / different cycle" and "international / multilingual"). Brett entry stays hidden from index until DRAFT flips false; index renders an empty-state explainer in the meantime so the surface doesn't look unfinished while operator A6/A7 outreach is in flight.

P2 — ship this quarter (audit Tier 2 — cleanup)

  • AUDIT-12 — /status page with 90-day uptime + recent incidents. B2756 — shipped. New IncidentTimeline section atop /status renders a 90-day calendar grid (green = no incident, gold = publicly logged incident), counts in window, days-since-last, list of incidents in window with severity chips, full-history link. Honest disclosure: we don't publish an uptime % because per-minute outage duration isn't captured — the receipt-shaped truth is the incident-count + per-day grid. Trust-positive per the audit's framing.
  • AUDIT-13 — Bundle-size ratchet: re-baseline OR retire. B2729 was the disposition (round-2 Tier-3 #19); same-day-as-audit timing meant the round-3 audit didn't catch the fix. B2752 confirmation: ratchet is green AND ratcheting downward. Current bundle 4,024,440 bytes (3.84 MB) shrunk -2779 B vs B2729 baseline; ratchet ceiling 4.43 MB, lastKnownTotal updated to the new measurement so future builds must hold-or-shrink against the tighter floor. Item closed.
  • AUDIT-14 — Prosody backfill script for pre-B2746 songs. B2758 — shipped. scripts/backfill-prosody-reports.ts walks the songs table; for every complete song with lyrics, recomputes scoreProsody and writes the fresh report. Idempotent; --dry-run default; --only-stale optional filter skips rows whose existing report already matches the fresh computation. npm run backfill:prosody-reports -- [--live] [--limit N] [--only-stale]. Per-song cost is microseconds (pure sync function — no Claude calls); ~700 songs take ~5 min serial dominated by Supabase round-trips. Failure-safe per-row; final summary prints updated / skipped / failed counts.
  • AUDIT-15 — Schema.org markup audit. B2885 — shipped. New scripts/check-schema-markup.ts ratchet enforces required JSON-LD @type values on /blog/[slug] (BlogPosting|Article + Person + Organization), /songwriting/[slug] (Article|HowTo|Course + Organization), /about (Person + Organization), and /scoring/standard (ScholarlyArticle|TechArticle|Article|WebPage + Organization). All 4 surfaces pass on first run. Wired into check:all (41st gate) + dedicated ci.yml step. The 0-violation ratchet the audit asked for, forward-only.
  • AUDIT-16 — Sacred Accidents framework as public artifact. B2570 — shipped; checkbox flipped at B2872. Public surface lives at /sacred-accidents (NOT /about/sacred-accidents — flat URL preferred). Renders the typed src/lib/sacred-accidents/data.ts source. OG image at /sacred-accidents/opengraph-image.tsx. scripts/check-sacred-accidents-sync.ts ratchet (B2699) enforces that every ## SA#N header in docs/SACRED-ACCIDENTS.md has a matching SACRED_ACCIDENTS entry in data.ts. Currently surfaces SA#11–#19 in full + stubs for #1–#10 awaiting historical reconstruction.
  • AUDIT-17 — Pricing FAQ cut to 6 items. B2872 — shipped. PRICING_FAQ trimmed 12 → 6 items: voice flattening, training data, ownership, free tier, refund, plan changes. The 6 cut items (music generation, credit packs, why upgrade, Creator→Pro trigger, ChatGPT comparison, Topic Sovereignty) moved to a new EXTENDED_FAQ const + rendered on the new public /faq page. /pricing FAQ ends with a "More questions answered on the canonical FAQ →" link to /faq. The FAQPage JSON-LD on /pricing now reflects only the 6 kept items (matches what renders).

P3 — defer until a deliberate sprint slot opens

  • AUDIT-18 — Forge V2 cold-visitor SSR shell. B2883 — shipped. New src/app/forge/ForgeColdShell.tsx server-renderable component replaces the pre-hydration blank <div> Suspense fallback with H1 + value prop + skeleton prompt textarea + skeleton genre/voice selectors + "Sign in to forge" CTA + Crucible escape hatch. Cold visitors + crawlers + first-paint humans now see meaningful content in the first 200ms instead of an empty viewport.
  • AUDIT-19 — Tool discovery card on /forge. B2881 — shipped. New src/components/ForgeFirstVisitHint.tsx renders above ForgeToolsExplainer on the idle composing surface: cyan-bordered callout, "First time here? Start with the Crucible — paste any lyric, get an 8-voice verdict in 30 seconds." localStorage-persisted dismissal so returning users never see it after closing it once. Pairs with the explainer below (which disambiguates the 3 tools); this hint adds the time-based first-visit nudge the audit specifically asked for.
  • AUDIT-20 — Mobile audit pass on /scoring/standard + /pricing. B2886 — shipped. New e2e/mobile-audit.spec.ts Playwright spec enforces two invariants at iPhone-13 viewport across 5 SEO-trust surfaces (/scoring/standard, /scoring/standard/whitepaper, /scoring/standard/fidelity, /scoring/standard/fidelity/cite, /pricing): (1) no horizontal scroll (document.scrollWidth ≤ viewport.clientWidth + 2px tolerance) — the "wide-table scroll-jail" the audit named; (2) every interactive element ≥ 44×44 px tap target per WCAG 2.5.5 + Apple HIG. Inline prose links exempted from tap-target rule. Future page additions get the invariant suite automatically by appending to MOBILE_PAGES.

Audit's meta-finding (not a build, but a steering rule)

The next 90 days needs ~70% of operator hours on outbound, ~30% on engineering — the inverse of recent history. The engineering velocity is 5x ahead of where it needs to be; distribution velocity is 1/5x. The autonomous queue above keeps engineering moving without operator input; the operator's hours should be re-routed to the OPERATOR-PUNCH-LIST.md audit-round-3 items (arXiv submission, partnership outreach, songwriter testimonial recruitment, hit-corpus curation, HN submission, Substack mirror).


Vault Intelligence V1 — Phase 3+ open threads

Phases 0-3D shipped B3024 → B3051 across the 2026-05-22 session. The 2,429-song admin corpus is audited (v1.2.0), embedded (1536-dim HNSW), classified (96 exemplars + 12 junk), surfaced via 5 admin dashboards, AND retrievable inline in /forge with arc-filtering + per-card selection + full-lyrics mode + localStorage persistence.

These five items are what's left for V1 → V2:

  • VAULT-OPEN-1 — Operator spot-check the 96 auto-exemplars (~8 min interactive). Walk /admin/vault/curate?status=exemplar&arc=indie (etc.) with keyboard hotkeys (E/X/J/R/C/N). Demote any that look wrong. The auto-pass picked top-12 per arc by composite_score; some are probably calibration artifacts (especially in pop/indie where scores skew high). Cheap sanity check that validates retrieval quality.
  • VAULT-OPEN-2 (B3169) — Cross-arc score normalization primitive shipped. src/lib/vault/score-normalization.ts ships the pure-compute layer: computeArcScoreStats(rows) (per-arc mean + sample stdDev + sorted scores), computePercentileRank(score, sorted) (Hazen-rule 0..100 percentile with tie half-credit), computeZScore(score, mean, stdDev) (standard z-score with null-safe stdDev=0 handling), and the one-shot normalizeScores(rows) that takes raw {arc, composite_score} rows and returns the same rows annotated with normalized_score + percentile_rank + z_score. Null-handling is uniform: null composite_score in → null fields out; arc with N<2 in → null fields out (no variance to normalize against). Companion percentileLabel(pr) produces operator-facing chip text ("Top 1%" / "Top 5%" / "Top 10%" / "Top 25%" / "Above median" / "Median" / "Below median" / "Bottom 25%" / "Bottom 10%" / "Bottom 5%" / "Unranked"). The motivating test case ("rock-60 = top of rock, pop-95 = top of pop, both percentile-equivalent") passes — the 35-point raw composite-score gap collapses to the same percentile rank under per-arc normalization. 34 tests pin every helper + edge case (empty input, single-row arc, ties, floating-point clamp, non-finite values). normalized_score == percentile_rank in this calibration; the separate field reserved for a future scaled-z calibration without breaking consumers. This ships the math; the dashboard ranking surfaces consuming it land as a follow-up.
  • VAULT-OPEN-3 (B3168) — Inverse-mode "avoid this voice" references shipped (prompt-side). src/lib/forge/vault-references.ts VaultReference interface gains an optional inverse?: boolean field (default false; backwards compatible). buildVaultReferenceBlock() splits the input refs into positive + inverse arrays preserving input order within each group, then renders TWO distinct blocks: (a) the original "VAULT REFERENCES — examples from the operator's corpus...Honor the established voice + craft choices" with EXEMPLAR labels; (b) a new "VAULT ANTI-PATTERNS — voices to depart from...Do NOT honor the voice, register, image vocabulary, or structural shape" block with ANTI-PATTERN labels. The closing directive adapts to which sets are present: positive-only → original "Use the references to anchor voice + craft" copy; inverse-only → "Depart from the ANTI-PATTERN voices — the new song should feel UNLIKE what this operator has written before"; mixed → "Honor the EXEMPLAR voices; depart from the ANTI-PATTERN voices." 18 tests pin: backwards compat (inverse=false renders as EXEMPLAR), polarity routing, mixed-mode ordering (positives always first), per-group independent numbering, input order preservation within each group, closing-directive variants, trailing-separator preservation. The UI side (operator-flagged inverse refs in VaultInspirationPanel.tsx per the punch list's "Sister-song panel pattern works as the UI ('Like this / Not like this' pair)") is a forthcoming build — this ships the prompt-side primitive so the consuming code can wire when the operator hits the inverse-toggle button.
  • VAULT-OPEN-4 — Per-user Vault graduation: every paid user gets their own corpus (4-6 builds). Currently only admin tier has corpus. To turn this from operator-only-tool into product-feature: scope song_audit_runs + song_embeddings + exemplar_status by user_id, gate the curation/anthropology dashboards per-user, drop the admin-only requireAdmin checks in favor of per-user RLS. Biggest moat play but most work. Requires a per-user audit-on-forge hook + a per-user auto-curate when N exemplars accumulate.
  • VAULT-OPEN-5 (B3167) — Contribution Ledger auto-write hooks shipped. src/lib/contributions/log.ts exports logContribution(input) — fire-and-forget service-role insert into song_contributions. Fail-open (never throws, never blocks the song save); env-toggle disabled via SF_CONTRIBUTION_LEDGER_DISABLED=1; column-missing degradation matches the skillFingerprint / prosodyReport / fidelityAudit defensive patterns. Three call sites wired: (1) finalize-song.ts writes ai_generation + whole_song + forge-v2.<BUILD_NUMBER> + 'substantial' on every successful fresh forge; (2) refine/route.ts writes human_refine + whole_song + refine-v1.<BUILD_NUMBER> + 'substantial' + contributor_id=user.id on every saved user-driven refinement, with the parent songId + version number captured in artifact_locator; (3) gauntlet/route.ts writes ai_revision + whole_song + gauntlet-v1.<BUILD_NUMBER> + 'de_minimis' on every standalone gauntlet completion, with the linesReplaced count in artifact_locator. 13 tests pin the row-builder + env toggle. The Moat 4 ledger is no longer empty — every forge / refine / gauntlet writes a row from B3167 forward.

Recommended order: OPEN-1 (cheap), then OPEN-4 (the moat), then OPEN-2 + OPEN-5 in parallel, then OPEN-3.


Genre Pages Excellence — B3053 WAR Room

Full findings doc: docs/GENRE-PAGES-WAR-ROOM.md. The operator asked for a 100-expert / 100-round audit of /genres, /genres/ [slug], /genres/[slug]/[substyle] with the white-text-readability complaint as the spark + four new features to integrate: 100-prompt catalog per genre, Genre Radio (admin songs with audio + best-lines + sort), batch auto-generate by genre with per-prompt reroll. The audit found 13 presentation issues (top: per-arc color promise unused — every accent is violet — and text-dark-500 BPM caption AA-FAIL) and laid out a 7-build sequence.

Tier 1 — B3054 (smallest defensible build)

  • GENRE-1 — Contrast pass. B3054 — shipped. Every text-dark-400 body-copy instance raised to text-dark-300 across /genres, /genres/[slug], /genres/[slug]/[substyle], /genres/audit-in-action. The text-dark-500 BPM caption (AA-FAIL 3.5:1) raised to text-dark-400. SA italic quote lifted to text-dark-50. Two text-violet-100 gradient-body instances (AA-FAIL ~3.1:1) raised to text-violet-50.
  • GENRE-2 — Mobile pass. B3054 — shipped. Audit-primitives table wrapped in overflow-x-auto with min-w-[640px] floor (no silent overflow at 375px). 4-pill stat row uses grid grid-cols-2 at xs → sm:flex sm:flex-wrap at 640px+ to prevent uneven 2×2 wrap.
  • GENRE-3 — Hierarchy reorder. B3054 — shipped. New order: Hero → Try-It (was position 4) → Substyles (was 5) → Load-Bearing Primitive (was 2) → Audit Primitives collapsed in <details> (was 4) → Banned → Corpus → CTA. CTAs no longer buried under two slabs of expository copy.
  • GENRE-4 — JSON-LD polish. B3054 — shipped. TechArticle schema gains author: { Organization } and mainEntityOfPage: { @id: url } for SERP rich-result eligibility.
  • GENRE-5 — `/lyrics/[genre]` cross-link. B3054 — shipped. New "lyrics guide →" button added to the corpus block on /genres/[slug] — the two overlapping-audience surfaces are no longer strangers.

Tier 2 — B3055 + B3056

  • GENRE-6 — Per-arc accent system. B3058 — shipped. New src/lib/genre-arc-theme.ts defines an ArcAccent token bundle (16 slots × 9 arcs = 144 static class strings, all JIT-safe). 9 arcs picked per WAR Room: folk emerald-500, rock orange-500, pop violet-500 (canonical), rap amber-500, indie cyan-500, rnb fuchsia-500, worship yellow-500, country red-400 (red-500 was AA-borderline), latin sky-500. getArcAccent(slug) resolves the bundle; unknown slugs fall back to pop. Wired through all violet hardcodes on /genres, /genres/[slug], /genres/[slug]/[substyle]. The 9-card grid on /genres now blooms 9 distinct colors; folk and rap and pop no longer look identical. 6-test invariant suite ratchets the bundle shape + per-arc color uniqueness + -50 gradient-text WCAG floor.
  • GENRE-7 — Per-arc hero glyph + gradient. B3059 — shipped. ArcAccent interface gained 3 slots (heroIcon lucide name, heroBg warm/cool token, iconLarge -400 shade). Per-arc icons: folk Trees, rock Flame, pop Heart, rap Mic2, indie Sparkles, rnb Music3, worship Sun, country Sunset, latin Music2. Warm bg for folk/rock/rap/worship/country; cool for pop/indie/rnb/latin. Hero block on /genres/[slug] now wraps in the radial backdrop + renders a 64px tinted glyph above the H1.
  • GENRE-8 — Failure-mode disclosures. B3060 + B3076 — shipped. New src/lib/failure-mode-definitions.ts registry with getFailureDefinition(arcSlug, modeName). B3060 seeded the rap arc with 10 inline definitions from docs/RAP-FORBIDDEN-ARCHIVE.md (5 with examples). B3076 expanded coverage from 10 → 90 — all 9 arcs now have all 10 canonical mode names registered with operational 1-sentence definitions. Country / folk / rock pull verbatim from <arc>-eval-context.ts (canonical names match the eval taxonomy). Pop / latin / rnb / worship use GENRE_ARCS canonical names (which diverge from the eval-context "Diane Warren By Numbers"-style register) and were authored against the SA#21-#27 paradigms + per-arc Forbidden Archive docs. Indie blends 7 eval-context matches + 3 newly-authored (Voice Shift Mid-Song / Detail Without Stakes / Universalist Shortcut). Cross-check script confirms 90/90 canonical → registered with zero orphans. UI on /genres/[slug] wraps each red chip in a <details> with the mode name as <summary> + chevron; expansion reveals the definition + optional example.
  • GENRE-9 — 100-prompt catalog data file (infrastructure + seed). B3061 — shipped. New src/lib/genre-prompt-catalog.ts ships the CatalogPrompt interface (extends GenreShowcasePrompt with id + mood/tags), getCatalogPrompts(slug) loader (falls back to the 3-per-arc showcase seed when no expanded catalog is registered), and getRandomCatalogPrompt(slug, { excludeIds, substyle }) for reroll flows. 8 unit tests cover seed presence + unique ids + stable refs + filter behavior. scripts/generate-genre-prompt-catalog.ts ships the Haiku batch generator (per-arc + --all flags; outputs to tmp/genre-catalog-<arc>.json for operator curation). Operator runs the script + curates the output to expand from the 27-prompt seed to 900. UI consumes via the loader, agnostic to whether the catalog is at seed (27) or expanded (~900).
  • GENRE-10 — Featured-prompt hero + scroll strip + Shuffle. B3062 — shipped. New src/components/genres/TryItSection.tsx client component renders: (a) one LARGE featured-prompt card with per-arc gradient + prominent Forge button + Shuffle, (b) horizontal scroll strip of next 12 catalog prompts as chips (scroll-snap mandatory), (c) "Browse all N →" link to the future /genres/[slug]/prompts page. Shuffle uses session-scoped rolled-set to avoid repetition until the pool exhausts. SSR-safe: server renders the first catalog entry as the initial featured; client takes over for reroll. Replaces the 3-static-card grid. The Try-It section is now the page's primary conversion engine, not wallpaper.

Tier 3 — B3063 (all 5 items in one vertical ship)

  • GENRE-11 — Genre Radio data layer. B3063 — shipped. New /api/genres/[slug]/radio route. 2-step query: pulls song_audit_runs rows at v1.2.0 matching the arc, intersects with admin-published songs that have audio_url. Filters by subscription_tier = 'admin', is_public = true, status = 'complete'. Cached s-maxage=120, stale-while-revalidate=600.
  • GENRE-12 — Audio player integration. B3063 — shipped. Inline native <audio controls autoPlay> in each row. Clicking "Play" promotes the row to autoplay; another row's Play replaces. Pause sync via onPause/onEnded.
  • GENRE-13 — Best-line extraction. B3063 — shipped. Server-side join into evaluations.transcendent_lines; surfaces the first line per song under the title with a Sparkles icon + italic quote treatment.
  • GENRE-14 — Sort controls. B3063 — shipped. 3-button pill row: top score (default) / newest / most played. Toggle pauses any playing audio (list reorders, target row may scroll out).
  • GENRE-15 — Loading skeleton. B3063 — shipped. 5-row pulse skeleton with staggered animationDelay for initial fetch + sort changes. Errors degrade silently to "Radio temporarily unavailable" copy.

Tier 4 — B3064 (5 items shipped, full /forge batch-state integration deferred)

  • GENRE-16 — Batch Genre-Catalog mode. B3064 — shipped as standalone batch picker. New /genres/[slug]/prompts page + <PromptBatchPicker> client island. Operator picks N (1-10) prompts via per-card toggle; selected prompts pool into a sticky-bottom batch tray. The handoff to /forge is per-prompt deep-link (each "Open #N" button opens the forge in a fresh tab with ?prompt=...&genre=...). Full integration into the existing use-batch-state.ts state machine deferred to a follow-up — would require batch-state to accept a multi-prompt URL contract; mechanical but tangled, not worth tying up this round.
  • GENRE-17 — Preview state with reroll/remove. B3064 — shipped. Each batch-tray row has Reroll + Remove icons. Reroll swaps the slot in-place using getRandomCatalogPrompt(slug, { excludeIds }) so the new pick respects the rest of the selection. Visual flash (ring-amber-400) on the rerolled card for 800ms.
  • GENRE-18 — Anti-dupe weighting. B3064 — shipped via excludeIds. The reroll passes the current selection's ids as the excludeIds set — the picker can never duplicate within a batch unless the operator manually adds a prompt twice (which the toggle blocks anyway).
  • GENRE-19 — `/genres/[slug]/prompts` browse page. B3064 — shipped. SSR'd /genres/[slug]/prompts route. Static-generated per arc. Renders the full catalog with substyle-filter pills (count per substyle), per-card Forge + Add-to-batch toggle, and the picker tray below. Sparse when catalog is at seed (3/arc); fills out after Haiku expansion.
  • GENRE-20 — Telemetry. B3064 — shipped. Fire-and-forget posthog.capture('genre_prompt_clicked', { arc, prompt_id, substyle, exercise_axis, source }) on Forge-link clicks. source distinguishes card-click vs batch-tray click. Drops silently when PostHog isn't loaded yet — best-effort.

Commit footer for B3054–B3060: Moat: Distribution (presentation + funnel improvements aimed at converting browse traffic).


Genre Pages Audit Round 2 — B3065 (operator-reported "doesn't flow with the rest of the site")

Full findings doc: docs/GENRE-PAGES-AUDIT-B3065.md. After shipping the 11-build B3054–B3064 sequence, the operator photographed the country page rendering with cream/pink Forbidden Archive chips + cream/pink Load-Bearing Primitive box. The WAR Room audit confirmed: the B3058 per-arc accent system was authored with `dark:` Tailwind prefixes that DO NOT FIRE because `tailwind.config.ts` has no `darkMode:` setting (defaults to media, which gates on OS preference). Light-OS visitors see the LIGHT-mode classes win. The 144 light/dark class pairs in src/lib/genre-arc-theme.ts plus 9 inline sites across the genre pages are all affected. Smoking gun confirmed.

Tier P0 — B3070 (must — the dark-mode bug blocks correct rendering for ~half of visitors)

  • GENRE-22 — Drop light-mode classes from all 144 accent slots. B3070 — shipped. src/lib/genre-arc-theme.ts rewritten: 9 bundles × ~16 slots, every paired bg-X-50 dark:bg-X-950/40 collapsed to bg-X-950/40. 6 inline dark:-prefix sites on src/app/genres/[slug]/page.tsx (breadcrumb link + Haiku pill + Forbidden Archive disclosure chips) collapsed to dark-tone classes only. 8 inline sites on src/app/genres/page.tsx (hero pre-header + 4 stat numbers + 3 "Why this matters" section icons) collapsed via bulk replace. New scripts/check-no-dark-prefix-in-genre-pages.ts ratchet walks src/lib/genre-arc-theme.ts + src/app/genres/** + src/components/genres/** and fails the build on any dark: prefix reappearing. Currently green: 0 violations across 9 gated files.

Tier P1 — B3067 (brand parity)

  • GENRE-23 — Replace solid-fill gradient blocks with site-canonical pattern. B3077 — shipped. ArcAccent gained a ctaBorder: 'border-{color}-500/30' slot (all 9 arcs). 5 surfaces converted: /genres/[slug] CTA, /genres/[slug]/[substyle] CTA, /genres/audit-in-action CTA (now per-arc instead of hardcoded violet), TryItSection featured card, PromptBatchPicker selected-prompt pills. New pattern: rounded-2xl border ${accent.ctaBorder} bg-dark-900/60, heading uses accent.boxLabel for tint, body uses text-dark-200. CTA buttons route through <Button intent="forge">. Per-arc identity now lives on tinted border + heading; surface is site-canonical dark elevated card. 7-test invariant suite ratchets the new slot's shape (/^border-[a-z]+-500\/30$/).
  • GENRE-24 — Define `font-display` alias. B3077 — shipped. tailwind.config.ts gained display: ['Sora', 'Inter', 'system-ui', 'sans-serif'] matching the canonical heading family. Pre-3077, 58 font-display headings across 19 files (admin/vault dashboards, /genres/* surfaces, components, homepage genre grid) silently inherited Inter because the alias was undefined. One-line config fix; affects rendering site-wide.
  • GENRE-25 — Standardize breadcrumb pattern across 3 genre routes. B3074 — shipped. All 3 sibling routes now use one idiom: segments separated by /, current page as a non-link in text-dark-300, parent segments as accent-colored Link components, aria-label="Breadcrumb" on the nav element. /genres/[slug] = All genres / {Genre} (2-segment); /genres/[slug]/[substyle] = All genres / {Genre} / {Substyle} (already 3-segment, unchanged); /genres/[slug]/prompts = All genres / {Genre} / Prompts (was a single back-link with arrow icon).

Tier P2 — B3068 (flow + density)

  • GENRE-26 — Compress `/genres/[slug]` 8 sections → 5. B3075 — shipped (partial). Load-Bearing Primitive section folded into the hero as an inline accent-bordered chip; the standalone rounded-xl section is gone. GenreRadio moved from position 4 (after Substyles) to position 6 (just before Corpus) — proof cluster now adjacent. Corpus rounded-xl card flattened into a border-t pt-6 block with shortened button labels. Final order: Hero (with LB chip) → Try-It → Substyles → Audit Primitives (collapsed) → Forbidden Archive → Genre Radio → Corpus → CTA. The LB section absorbed entirely; the Corpus section's vertical real estate roughly halved. "Other substyles" deletion from the substyle deep page deferred to a follow-up.
  • *GENRE-27 — Focus-visible rings on custom buttons in `src/components/genres/.** *B3074 — shipped.* Canonical focus-visible:outline-none focus-visible:ring-2 focus-visible:ring-cyan-400/70 focus-visible:ring-offset-2 focus-visible:ring-offset-dark-950` added to: TryItSection Shuffle button; GenreRadio 3 sort pills (top-score / newest / most-played); PromptBatchPicker substyle filter pills (× ~5-17 per arc) + per-card batch toggle + Reroll + Remove icon buttons. Keyboard users tabbing through the genre surfaces now see the canonical cyan ring that matches the rest of the site.
  • GENRE-28 — `motion-reduce:` modifiers. B3074 — shipped. Every animated/transition surface in src/components/genres/* now honors prefers-reduced-motion: TryItSection Shuffle button, GenreRadio sort pills, all PromptBatchPicker buttons get motion-reduce:transition-none. PromptBatchPicker reroll flash (the 800ms ring-2 ring-amber-400/60) downgrades to ring-1 under reduced-motion — still a visible cue, lower intensity.

Tier P3 — Backlog

  • GENRE-29 — Per-arc accent propagation (Option A: propagate). B3080 — shipped. Operator chose A: extend arc accents beyond /genres/*. Three surfaces threaded:

1. `/lyrics/[genre]` — hub badge, "Why the profile is tuned" heading + check icons, leaderboard heading + score numbers all pick up the resolved arc accent. Hub-only slugs (hip-hop) translate to arc slugs (rap) via the new resolveArcSlug() helper. 2. Navbar — when the route matches /genres/X or /lyrics/X, a 2px absolute-positioned bar tinted in the arc's accent renders at the navbar's lower edge. Uses usePathname only (no useSearchParams) so global static-rendering stays intact. New navTopBar ArcAccent slot (bg-{color}-500/40) all 9 arcs. 3. `/forge?genre=X` — "Tuned for {Label}" chip with per-arc tint surfaces above the composing pad when the deep-link arrives with a recognized arc. Mirrors the persona-locked chip idiom; clicking "what this means →" deep-links to /genres/{arc}. Free-text genres still pass through unaccented (pre-3080 byte-identical behavior).

New library exports: resolveArcSlug(input) handles arc slugs, hub slugs (hip-hop → rap), labels (Hip-Hop, R&B), and punctuation-stripped fallbacks; getArcAccentForGenreInput(input) returns the bundle or null (NOT the violet default — lets the Navbar avoid misleading tint on non-genre routes). 7 new tests cover the resolver edge cases.

  • GENRE-30 — Custom audio player for GenreRadio rows. B3077 — shipped. New inline RadioAudioPlayer component replaces the native <audio controls autoPlay> UA chrome. Layout: per-arc-tinted play/pause toggle button (uses accent.pillBg + accent.pillText), monospace tabular-nums elapsed/total time, native <input type=range> seek bar styled with accent-current, surface wrapped in border ${accent.ctaBorder} bg-dark-900/60. Honors canonical focus-visible ring + motion-reduce; autoplay-on-mount with graceful Safari-block fallback (button visible, click → play). Same onEnded contract as the old native player. ~75 LOC inline (kept local to GenreRadio since no other surface needs it).
  • GENRE-31 — Reroll flash arc-accent threading. B3078 — shipped. ArcAccent gained two new slots: flashRing: 'ring-{color}-400/60' + flashBorder: 'border-{color}-500/40' for all 9 arcs. PromptBatchPicker's 2 hardcoded ring-amber-400/60 usages (the prompt-card flash + the batch-tray-row flash) now resolve through accent.flashRing + accent.flashBorder. Country / worship / indie / etc. each flash in their own arc color on reroll instead of all 9 flashing with rap's amber. Test suite gained matching invariants (/^ring-[a-z]+-400\/60$/ + /^border-[a-z]+-500\/40$/).
  • GENRE-32 — Sticky tray + Footer collision. B3078 — shipped. /genres/[slug]/prompts main element gained pb-32 (split from the symmetric py-12 sm:py-16 into pt-12 sm:pt-16 pb-32). The sticky batch tray inside PromptBatchPicker (sticky bottom-3 z-10) was colliding with the global Footer on short pages + at scroll bottom. ~128px of bottom-padding gives the tray clear room above the Footer.
  • GENRE-33 — Mobile filter-pill compression. B3078 — shipped. On a 375px viewport, 17-substyle Latin was filling 4-5 rows of pills because every pill ran at px-3 py-1 text-xs with a · N count suffix. Post-3078: tighter mobile sizing (px-2 py-0.5 text-[10px] mobile, sm:px-3 sm:py-1 sm:text-xs desktop) + count suffix hidden on mobile via <span className="hidden sm:inline">. Same 17 pills now fit in 2 rows on mobile + 1 row on desktop. Accessibility preserved via aria-label="Filter to {s} ({n} prompts)" so screen readers still read the count.

Tier 1.5 — Operator-spotted IA mismatch (small standalone)

  • GENRE-21 — Footer `/lyrics/` hubs don't match Navbar `/genres`. B3056 — shipped (Option B). The 8-link "Lyrics by genre" row was retired from the Product column. A single "Genres · 9 deep" link sits alongside Forge / Refine / Crucible at the top of the Product column, pointing at /genres. The /lyrics/[slug] SEO long-tail surfaces remain reachable from /genres/[slug] corpus blocks (B3054 GENRE-5) + from the /lyrics index. Two competing IAs (one with "hip-hop" + missing latin, one with all 9) collapsed into one canonical surface.

Today's top 5 (Tier A first wave)

These are the foundational discipline items. Each is one focused day's work. All five = one work week. Each one ratchets a CI gate so regression is impossible.

  • #1 — Kill the `explicit-any` allowlist down to ZERO. Every : any, every as any, every <any> becomes a precise type or typed unknown + guard. Ratchet ceiling at 0; CI now fails on a single new occurrence. (B1178 partial 194 → 110, B1182 closeout 110 → 0. Total: 194 anys eliminated across the codebase. Manhattan + data-intelligence supabase params typed via SupabaseClient bulk pass; gauntlet/route.ts gained a ParsedGauntletOutput interface; episodic-trace gained EpisodicTraceDbRow; deep-consolidation, dreaming, plateau-detection all typed against EpisodicTrace. extractLearnings widened to `unknown` to remove launderer casts at call sites.)
  • #2 — Wire jsdom + write the FIRST real `renderHook` test. Add jsdom devDep, configure vitest with a *.behavior.test.ts glob in jsdom env. Convert useSongDetailEdit test from structural-source-text to real renderHook + act() + state assertions. Path opens for converting all extracted hooks. (B1174)
  • #3 — Stand up `src/lib/design-tokens.ts`. Semantic color names (tokens.score.high/mid/low, tokens.status.error/warning/success). Migrate /s/[slug] score-color logic + ScoreBadge as proof. Lint rule: forbidden literal hex/rgb in JSX. (B1175 stand-up; B2098 milestone — ratchet hit ZERO. 414 → 478 → 0 across the B1175→B2098 arc. Every JSX hex/rgb/rgba in `src/app` + `src/components` now lives in `design-tokens.ts`. CI ceiling fixed at 0; any new literal fails immediately.)
  • #4 — Bundle-size CI gate flips from snapshot to blocking. bundle-size-ceiling.json exists; not enforced. Write scripts/check-bundle-size.ts, gate CI. Shrink-or-stay rule. (B1176)
  • #5 — Golden-eval regression gate ships. src/lib/golden-evals/ exists but invisible. Pick 3 prompts (country/pop/hip-hop), forge once, snapshot lyrics + scoring output, commit. CI re-forges with same model+prompt+temp and asserts byte-identical output. Locks the engine against silent prompt drift. (B1177 — shipped as 4 deterministic system-prompt snapshots: evaluator, refine@50, gauntlet@high, gauntlet@low. The forge prompt is non-deterministic by design (Math.random feature gates) so the byte-identical gate runs against the eval/refine/gauntlet prompts where prompt-drift actually matters.)

Net by end of week 1: +2 CI gates, 1 new test pattern, 1 new design primitive, 194 fewer anys.


Tonight\u2019s run (B1571\u2014B1584, 2026-04-27)

A 14-build sprint triggered by one operator question ("How are we doing to be able to make an Italian opera piece or Gregorian Chant?"). What started as a capability lift exposed the deeper "capability without surface" bug class, triggered a Deep Audit, and produced:

  • 2 new CI ratchets installed (design-canon at 662, capability-

surface at 0 gaps \u2014 both compounding-down).

  • 8 new one-click Tradition presets on V2 forge (Operatic,

Sacred Chant, Delta Blues, Murder Ballad, Sea Shanty, Spiritual, Hymn, Spoken Word). Each pairs a dedicated ghost voice + DNA entry + pure-mode prompt path + unit-test contract.

  • 6 new ghost voices in GHOST_PROFILES (verdi, hildegard,

crossroads, reaper, mariner, witness, hymnwright, speaker).

  • 8 new genre DNA entries in GENRE_DNA_DATABASE

(opera, chant, murderballad, shanty, spiritual, hymn, spokenword + 4 surfaced from existing DNA: gospel/punk/ reggaeton/dancehall).

  • Multilingual UI: language picker surfaced on V2 (was hidden

since B1420), enabling Latin chant + Italian opera one-click.

  • Conversion CTA: songs-remaining chip on V2 forge converts

the dashboard-only counter into a forge-page upgrade prompt.

  • First-time-user mode: zero-prior-songs detection reduces

decision overhead with a "Show all options" escape.

  • Console-log ceiling: 19 \u2192 10 (\~96% cumulative reduction

since B1037 baseline of 251).

  • Unchecked-index allowlist: 10 \u2192 8 (-20% this build).
  • 2 new evergreen guides at /songwriting (chant + opera craft).
  • 1 new blog post at /blog (the build-narrative).

The full list of builds + their punch-list mappings:

  • B1571: opera + chant capability lift (Verdi + Hildegard ghosts)
  • B1572: chant capability fixed (DNA + pure-mode prompt + V1 tiles)
  • B1573: design-canon CI ratchet
  • B1574: language picker on V2 (capability-surface gap closed)
  • B1575: capability-surface ratchet (4 genre gaps closed)
  • B1576: V2 Tradition presets (Sacred Chant + Operatic) + unit tests
  • B1577: blog post on the 6-build chant/opera narrative
  • B1578: songs-remaining chip on V2
  • B1579: console-log ceiling 19 \u2192 10 (9 calls migrated)
  • B1580: chant + opera evergreen guides at /songwriting
  • B1581: first-time-user simplified mode for V2
  • B1582: Tier-1 traditions batch 1 (Delta Blues + Murder Ballad + Sea Shanty)
  • B1583: Tier-1 traditions batch 2 (Spiritual + Hymn + Spoken Word)
  • B1584: unchecked-index allowlist 10 \u2192 8

The pattern worth keeping: every operator critique becomes either a fix OR a CI ratchet that prevents the bug class. Capability/ surface was the second ratchet installed this run; the first (design-canon) had been Quality Council #1 forward-looking item all week.


Tier A — Foundation (Linear parity)

  • #1 — Kill explicit-any allowlist (B1178 partial 194→110, B1182 closeout 110→0. Ratchet at 0; one new any fails CI.)
  • #2 — jsdom + first real renderHook test (B1174)
  • #3 — Design tokens module (B1175)
  • #4 — Bundle-size CI gate (B1176)
  • #5 — Golden-evals CI gate (B1177)
  • #6 — Kill the unchecked-index allowlist (B1597 CLOSEOUT: cumulative 166 → 0 (100%). B1595 cleaned components/CoverArt.tsx (15 \u2014 hexToRgb signature widened to absorb 12 + RGBA pixel-stride asserts for 3); B1596 cleaned lib/genre-profile.ts (15 \u2014 `!` on every GENRE_PROFILES static-key lookup since each key is guaranteed in the record); B1597 cleaned admin/lineage/page.tsx (30 \u2014 narrow-once locals in the Fruchterman-Reingold force-layout loops). Allowlist now empty. Tier-A item closed.)
  • #7 — Kill the strict-tsc allowlist (B1600 CLOSEOUT: cumulative 37 → 0 (100%). B1598 cleaned dashboard/page (4 dead-code orphans deleted: EvalDeleteButton + handleDeleteEval + upgrades pipeline). B1599 cleaned dashboard/SongDetail (10 violations: unused destructured hook returns + handleSunoExport orphan). B1600 cleaned forge/page (56 violations \u2014 ~40 dead imports from EXTRACT-1 factory pattern + 9 dead destructured hook fields + 2 orphan closures + cascading dead-setter cleanup). Allowlist now empty. Tier-A item closed.)
  • #8 — Kill the console-log allowlist (B1586 CLOSEOUT: 10 → 0. Cumulative since B1037: 251 → 0. Tier-A target hit. Final 8 calls migrated to logger.info: forge/criticalSend (2 in create-stream-plumbing), mycelial-pathways, data-intelligence/orchestrator, manhattan/creative-debt, manhattan/skill-frontier, side-effects/cover-art, side-effects/focus-group. Two legitimate exclusions documented in the script: observability/logger.ts (the canonical sink) and developer/page.tsx (a JS code snippet inside a multi-line backtick template). The exclusions are commented inline in scripts/check-console-log.ts.)
  • #9 — Visual regression with Playwright toHaveScreenshot() (every page, every breakpoint) (B1185: e2e/visual.spec.ts snapshots 8 public surfaces with dynamic-content masks (SongCounter, ShippingThisWeek, leaderboard, hero canvas). Per-platform baseline path (Linux is the CI source of truth). Tightened threshold 0.20 → 0.15. CI auto-uploads new baselines + diff PNGs as artifact for PR review. Single-breakpoint (1280x800) for now; mobile + tablet breakpoints land as follow-up when the desktop baseline is stable.)
  • #10 — A11y audit with axe-core in CI, block on serious violations (B1179: e2e/a11y.spec.ts scans 10 public surfaces; block on serious+critical, advisory warnings on moderate/minor)
  • #11 — Storybook for every primitive (Button, Card, Badge, Input, ScoreBadge, etc.) (B1602 deferred; B2358 stand-up + B2372 closeout. B2358 installed Storybook 8.6 + the first 3 stories (Button 12, Chip 10, ScoreBadge 9). B2372 shipped the remaining 7 (Disclosure, Surface, MetadataRow, NextActionPair, TrustBlock, FooterLadder, CrucibleVoiceIcons) in a single Audit 2026-05-14 A16 batch — every design-system primitive now has a story file with realistic args + variant coverage. Item flipped to closed at B2421 (the punch list lagged the actual shipment by ~50 builds; B2421 caught the drift during a punch-list audit).)
  • #12 — Real component library (replace ad-hoc Tailwind with <Button intent="forge" size="md">) (B1602 kickoff: Button primitive shipped \u2014 typed `intent="forge|ai|outline"` + `size="sm|md|lg"`, polymorphic <button> / <Link> via href prop, loading state with spinner, leadingIcon/trailingIcon slots. First migration: dashboard BillingTab upgrade CTA (next/link \u2192 Button href=...). 11 contract tests assert intent/size class composition + canon-compliance (no off-palette utilities, no rounded-3xl, no inline gradients). Closes when 30+ btn- call sites are migrated to Button + the design-canon gate adds a forbid-rule for raw <button className="btn-">.)
  • #13 — Type-generated Supabase schema (database.types.ts regenerated on migration; no hand-typed table shapes) (B1186 prep: workflow + npm script + placeholder + freshness check shipped. CLOSEOUT B1761: four-build chain landed the criteria — B1756 typed factory `createTypedServiceRoleClient()` scaffold + 6 contract tests; B1757 operator ran `npm run db:gen-types` + types regenerated against live schema (32 tables, 4 RPCs, schema 14.4); B1758 wave 1 migration (admin/regression-check + admin/costs, launderer ceiling 13 → 11); B1761 wave 2 migration (admin/forge-metrics + admin/cost-per-song, ceiling 11 → 8, including FULL/FALLBACK dynamic-select dead-code removal). Five admin routes now use the typed client; row types narrow against the live schema; the placeholder is gone; the freshness check would fail loudly if a migration drifted the types. Item closed.)
  • #14 — Zod runtime validation on every SSE event (no trusted casts) (B1180: PipelineEventSchema + CrucibleEventSchema; runtime validation in withSSEStream, forge create-stream-plumbing, and crucible route. 34 schema tests covering valid variants + drift catalog. Catches typo'd type discriminators, missing required fields, mixed envelope drift (text vs message). Critical events that fail validation get replaced with a generic error envelope so the client never sees a corrupt payload.)
  • #15 — Discriminated-union state for ForgeResult (Draft | Evaluated | Gauntleted, each with required fields) (B1181: ForgeSongState union + deriveSongState() pure derive function in src/app/forge/song-state.ts. useSongState() hook combines useForgeSessionState + useGauntletState through deriveSongState. 20 unit tests cover every transition path including the partial-Gauntleted guard. Source-of-truth state stays in the existing hooks; consumers migrate to the union as files are touched.)

Tier B — "Actually beat Linear"

  • #16 — xstate forge state machine (every transition explicit, every state typed) (B1045 hand-rolled state-machine reducer + B1532-B1536 V2 cutover delivered the same outcome xstate would have: explicit typed states, impossible-state representations rejected at compile time, forbidden transitions throw at the call site. Hand-rolled saved a 15-40KB bundle hit + the second mental model. The B1405 scaffold doc remains as the migration plan IF a future requirement makes the runtime visualizer worth the dependency. As of B1664 the forge state surface is tested at src/app/forge/forge-state-machine.test.ts + use-forge-state-machine.test.ts. Goal achieved without xstate.)
  • #17 — Forge page.tsx under 500 LOC (logic behind hooks/machines, page = pure composition) (B1626 V1 deletion brought page.tsx from ~1,909 LOC to 63 LOC. The page is now a Suspense + searchParams-keyed mount of <ForgeV2 />. All forge logic lives behind extracted hooks + the state machine. Item closed.)
  • #18 — Real-time presence for batch forging (multiple windows / collaborators see live progress)
  • #19 — Optimistic UI everywhere (public toggle, delete, rename — all instant + reconcile) (public-toggle B894, delete B1222, share-link B1254, collection-assign B1268, B1278 closeout: rename + collection-bulk-move + cover-art-regen. handleUpdateSong is now optimistic across the board (every PATCH lands locally before the round-trip; per-song snapshot restores on failure). handleBulkMoveToCollection clears selection + applies move IMMEDIATELY then reverts only the rows whose PATCH failed. regenerateCoverArt closes the prompt dialog + clears the input on submit so the loading spinner takes the thumbnail unobstructed; on failure the dialog reopens with the prompt re-primed and the error inline.)
  • #20 — Soft-delete with 30-day undo (no destructive action without recovery) (B1264 migration draft + B1275 API/UI wiring. DELETE /api/songs/[id] now soft-deletes (sets deleted_at + deleted_reason); falls back to hard-delete when the migration column is missing so the button never breaks. New POST /api/songs/[id]/restore + GET /api/songs/trash + /dashboard/trash page with one-click recover. Migration applies via supabase/migrations/add_soft_delete_to_songs.sql; trigger auto-fills the 30-day window. Reaper cron for past-window hard-delete is the follow-on.)
  • #21 — Per-user audit log (B1220 surface, B1228 SQL draft, B1269 read endpoint with migration-pending fallback, B1328 wired writes from finalize-song / restore / share / DELETE/PATCH on /api/songs/[id]. CLOSEOUT: operator applied the migration; B1727 swapped /dashboard/activity from songs-only derivation to user_audit_events as the authoritative source, with songs-derived events filling in for entries the audit table doesn't cover (older songs, pre-B1313 forges).)
  • #22 — Inngest async job queue for forge → score → gauntlet (currently inline) (B1407 scaffold + B1429 Phase 1 (cover-art) + B1550 Phase 2A (focus group) + B1726 Phase 2B CLOSEOUT (audio generation lift) — `forge/audio.requested` Inngest event + `generateAudioJob` handler in src/lib/queue/jobs.ts; src/app/api/songs/forge/post-forge-side-effects.ts dispatches the event before falling through to the inline path. 3-retry exponential backoff replaces the legacy fire-and-forget IIFE. Phase 2C (eval) + Phase 2D (gauntlet/superstyle) are UX-deferred indefinitely — they emit user-visible output during the SSE stream and breaking that apart for queue purity isn't worth the cost. Item flipped to closed at B2421 — the punch list lagged the actual shipment by ~700 builds; B2421 caught the drift during a punch-list audit. Status detail + Phase 2B work plan in docs/INNGEST-ASYNC-QUEUE-SCAFFOLD.md.)
  • #23 — Resumable forges (SSE disconnect → next request picks up where it left off) (B2357 Phase 1 SHIPPED: `GET /api/songs/forge/resume?id=<songId>` returns an SSE stream that polls the song row every 1.5s and emits phase/result/error events matching the fresh-forge envelope. Auth + ownership checks; hard cap at 285s; resumed events carry `resumed: true`. Client helper at `src/app/forge/resume-forge.ts` (5 tests). Phase 2 — durable event log + true event-replay — open; would require per-forge event persistence so a resumed stream can replay from a last-seen event id. Phase 1 gives the user a working resume path; Phase 2 makes resumed sessions byte-identical to the original.)
  • #24 — Public SDK on npm — @songforgeai/client, used by external tools (v0.2.0 published B1214; v0.3.0 published B1859 with verifySeal + voiceFingerprint. B3194 closeout: v0.4.0 adds `crucible()` — the 8-voice adversarial critique via POST `/api/v1/crucible`. SDK now covers every public `/api/v1/ endpoint surface. CrucibleVoiceResult + CrucibleVerdict + CrucibleRequest + CrucibleResponse interfaces; loose-string typing on voiceId / verdict / banner / panelAgreement so future server-side additions don't break consumers. 4 new tests pin POST shape + invalid_lyrics error mapping + rate_limit_exceeded handling + traditionScope preservation. 18/18 SDK tests green; tsc clean. Operator runs cd packages/sdk && npm publish` to push v0.4.0 to the registry when ready.)*
  • #25 — OpenAPI spec for /api/v1/score + auto-generated docs at /developer/api-reference (B988 — pre-existing.)
  • #26 — Distributed tracing with spans per pipeline phase (Honeycomb or DataDog) (B1226 Phase 1 — in-memory ring + tracing.ts; B1403 Phase 2 — forge SSE phases auto-instrumented via phase-logger bridge + /admin/trace UI with Gantt-style flame graphs. Vendor integration (Phase 3) deferred — in-process data informs the choice before commit.)
  • #27 — Sentry integration (every uncaught error categorized, source maps) (B1301-B1326 + B1379 shipped: Sentry runtime FULLY live \u2014 SDK init in 3 config files, captureException at 10+ surfaces, structured tags (build/sha/env/runtime/userId-hash) auto-attached. B1601 attempted to add withSentryConfig wrapper for build-time source-map upload but used the v8 API on a v10 SDK \u2014 the third options arg was silently dropped, removing the graceful-degradation guards, which broke SSE streaming + auth flow in production. B1617 reverted. B1725 CLOSEOUT: re-enabled withSentryConfig with the v10-correct 2-arg signature + sourcemaps.disable defaulting to !SENTRY_AUTH_TOKEN. The wrapper degrades to a no-op for every build environment that doesn't have the operator-supplied SENTRY_AUTH_TOKEN, so dev + contributors + token-less CI runs are unaffected. Activates source-map upload immediately when the operator sets SENTRY_AUTH_TOKEN + SENTRY_ORG + SENTRY_PROJECT in Vercel env. The B1618 OpenTelemetry-hang root cause was fixed via force-dynamic on the three affected pages, so the wrapper no longer triggers a static-gen Supabase deadlock.)
  • #28 — Per-route p99 latency dashboard at /admin/perf with regression alerts (B1204 — in-memory ring + /api/admin/perf snapshot + /admin/perf live table. /api/v1/score is the first wrapped route; expand by copying the recordLatency pattern.)
  • #29 — Feature flags via GrowthBook (every UX change ships behind a flag, A/B tested) (B1227 cohort-flag primitive + B1270 /admin/flags integration. Ready for vendor swap when GrowthBook decision is made; surface stays identical so consumer code doesn't change.)
  • #30 — Lighthouse score ≥ 95 on every public route, gated in CI (B1237 ratchet ceiling at .github/lighthouse-ceiling.json + scripts/check-lighthouse.ts gate; B1259 operator-run snapshot script; B1304 flipped the GitHub Actions workflow from soft-warn to BLOCKING. Current per-category floors enforced in CI: performance ≥ 80, accessibility ≥ 95, best-practices ≥ 90, SEO ≥ 95. Three of four categories meet the literal "≥ 95" target; performance pegged at 80 because real-world cold-start performance on a 9-route Next.js app with images + Vercel SSR is the realistic ceiling — bumping to 95 would either fail every push or require routes to ship without any non-trivial assets. Closeout B1414. Audit annotation if the literal-95-everywhere interpretation matters: the perf floor lives in `.github/workflows/lighthouse.yml` line 123 and `.github/lighthouse-ceiling.json`; raise both in lockstep with measured-improvement evidence, never speculatively.)

Tier C — "Past Linear" (our unique edge)

  • #31 — Public scoring SDK becomes the standard (third parties cite our npm package)
  • #32 — Per-metric API endpoints — /api/v1/metric/specificity (test individual signals) (B1210 — POST + GET on /api/v1/metric/[slug]. Same eval pipeline + auth + rate limits + reproducibility seal as /api/v1/score; projects out the requested metric.)
  • #33 — Scoring corpus opens — 1000+ human-scored lyrics, public, used to version the rubric
  • #34 — Public model card at /scoring/standard/model-card (prompts, temperatures, training rationale) (B1197 — pipeline tables + temperature rationale + reproducibility-fields table + 5 known limitations. Cross-linked from /scoring/standard.)
  • #35 — Quarterly rubric versioning (v1.1, v1.2 published with diff + migration notes) (B1211 partial: scaffold proven via v1.0.1 PATCH bump. Cadence policy formalized in RFC-0001. Closes when the first MINOR/MAJOR bump ships through the published process.)
  • #36 — Reproducibility seal (every score includes rubric version + model + temp; results reproducible from public docs) (B1199 — `seal: { rubricVersion, model, temperature, buildSha, build }` attached to every shapeScoreResponse output. Wires through both sync /api/v1/score and async/jobs routes.)
  • #37 — Internal CLI — sfai forge "prompt" from terminal, sfai score --file lyrics.txt (B1203 — bin/sfai.ts first cut. Verbs shipped: help / status / score / counts. `npm run sfai -- <verb>`.)
  • #38 — Crucible API (third parties hit the 8-voice critique programmatically) (B1248 — /api/v1/crucible synchronous JSON endpoint with Bearer auth, 60/hr per key, reproducibility seal.)
  • #39 — OG card per metric page (/scoring/metrics/voice/opengraph-image rendered programmatically) (B1202 — generateImageMetadata fans out 12 PNGs; tier-colored accent, giant metric number, gradient name, definition + URL.)
  • #40 — Embedded score widget (third-party blogs/sites embed a "scored by SongForgeAI" card with verifiable hash) (B1209 — /api/v1/badge SVG endpoint + /developer/embed snippet generator. Verifiable hash deferred to a future RFC; v1 trusts URL.)
  • #41 — Webhook outbound — song.forged, song.scored, gauntlet.completed events to user-supplied URLs (B1238 song.scored from /api/v1/score; B1274 song.forged from forge finalize-song + gauntlet.completed from PATCH /api/songs/[id]. Shared dispatch-helpers.loadSubscribersForUser plus per-event subscription filter on `subscribedTo`. All three cohort-gated on `webhook_dispatch_live` (default 0%); fire-and-forget delivery; failures captured to Sentry. Env-var subscriber list for v1; per-user table swap is a future build.)
  • #42 — Multi-language scoring (Spanish, French, Japanese rubric pages + scoring support) (B1404 prep: RFC-0009 opened in public comment through 2026-05-03. Pins seal-annotation contract + 3-phase methodology BEFORE implementation lands so the first non-English score carries honest disclosure. Phase 1 implementation (language param + UI dropdown) deferred to "Continue Multi-Language Phase 1" trigger phrase.)
  • #43 — Voice fingerprint analysis (given a corpus of one writer's lyrics, score Voice consistency over time) (B1413 scaffold doc + B1422 Phase-1-part-1 pure function: computeVoiceConsistencyIndex shipped in src/lib/voice-fingerprint.ts. B1685 Phase-1 endpoint shipped at /api/v1/writer/me/voice-consistency. B1687 Phase-1 dashboard surface shipped as VoiceConsistencyCard mounted in /dashboard?tab=stats — loading / insufficient-samples / error / ok all rendered, behavior tests cover all four states. Phase 1 (a) dashboard surface = DONE. Item stays open until Phase 2: multi-axis VCI with per-song fingerprint vectors persisted (RFC-0009 deferred until calibration corpus reaches threshold).)
  • #44 — Versioned songs UI (every revision on a timeline; restore any version) (B1255 gauntlet revision card + B1276 full timeline closeout. SongDetail.tsx now shows the complete version history (newest first) ungated from versionNumber > 1, so the feature is discoverable on day-one songs. Restore button snapshots the CURRENT state as a new version BEFORE overwriting, so any revert is itself reversible — round-trip safe. Empty state explains the contract; "Most recent" badge marks the freshest snapshot. Existing /api/songs/version POST + GET endpoints unchanged.)
  • #45 — Public weekly engineering report at /engineering (auto-generated from git log + changelog + telemetry) (B1201 — built from `git log` at request time. Highlights / ratchets / area heatmap / all commits. Force-dynamic; no telemetry plumbed in so the report is reproducible from outside.)
  • #51 — Hit-Song Calibration Corpus (Tier 3 of Brett review). Verified-hit lyrics scored against the rubric, used to detect calibration misses (rubric marks #1 hits as "Collapsed") and trigger Gravity-Rule recalibration. Closes when (a) corpus has ≥15 scored entries per genre, (b) at least one genre clears ≥90% pass-rate against its floor, and (c) /scoring/standard/hits ships as a public artifact with the reproducibility seal. Trigger: Brett (Nashville pro) found the rubric scored a verified hit "Collapsed" — credibility blocker for any commercial writer.

- Infrastructure shipped (B2078–B2083): src/lib/calibration/hit-corpus.ts schema + helpers (findCalibrationMisses, summarizeCorpusHealth, CALIBRATION_FLOOR_BY_GENRE), 11 vitest tests, scripts/calibration-init.ts + scripts/calibration-report.ts + scripts/calibration-set-lyric.ts, private-corpus/ gitignored, docs/HIT-CORPUS-CURATION.md (policy) + docs/HIT-CORPUS-WALKTHROUGH.md (literal step-by-step). - Operator pause point (2026-05-05): operator started populating with "Last Night" by Morgan Wallen as the first country entry. Got blocked on JSON escaping, fixed via B2083 set-lyric helper. Recovery recipe is in chat history; the workflow is: save lyric to lyric.txt → npm run calibration:set-lyric -- <id> ./lyric.txt → score on songforgeai.com → record scores block manually → npm run calibration:report. - To resume: trigger phrase Continue Hit-Song Calibration should re-load docs/HIT-CORPUS-WALKTHROUGH.md and pick up at "score the first entry" step. Or operator follows the walkthrough independently and pings when ready to ship /scoring/standard/hits. - Definition of done: country pass-rate ≥90% AND /scoring/standard/hits page live AND reproducibility seal showing build number + scored-at timestamp on every entry. At that point the rubric publicly stands on chart-agreement, not curator opinion. Brett gets a magic link to that page for the week-10 re-review. - Why this is in Tier C (and not Tier D or a separate tier): it's our deepest standards moat. Linear has nothing equivalent because they don't ship a craft rubric. The closest peer move is the Lyric Scoring Standard whitepaper (B1093). Hit-Corpus is what makes the Standard credible to commercial writers, not just literary ones.

Tier D — "Why-would-Linear-do-this"

  • #46 — Open the operating principles as a doc (how decisions get made, what gets prioritized, what gets cut) (B1200 — /about/principles. 10 principles, each paired with the inversion we reject. 4 decision gates. Linked from /about + footer Company column.)
  • #47 — Public RFC process (every major change ships after a 7-day public RFC at /rfc/<slug>) (B1208 — /rfc index + /rfc/[slug] detail + RFC-0001 (rubric versioning policy) opened in-comment through 2026-05-02.)
  • #48 — Reproducible deploys (git rev in footer matches what's running, every time, signed) (B1198 — src/lib/build-info.ts reads VERCEL_GIT_COMMIT_; footer renders "build N · shortSha" linked to GitHub commit; /api/version JSON endpoint for uptime checks.)*
  • #49 — Deterministic CI (same commit produces same artifact bytes, every time) (B1272 scaffold doc + B2099 Phase 1 + B2356 Phase 2 + Phase 4 infra shipped. `scripts/check-deterministic-build.ts` walks two `.next/` directories, hashes every artifact, categorizes diffs. `.github/workflows/deterministic-build-check.yml` runs weekly + on workflow_dispatch. B2356 added: SOURCE_DATE_EPOCH + TZ=UTC + LC_ALL=C.UTF-8 pinned in the build env (Phase 2); `--strict` flag + `.github/deterministic-build-allowlist.txt` allowlist loader + workflow_dispatch `strict` input for Phase 4 dry-runs. 10 new unit tests lock the allowlist parser + partitioner. Item stays open until the default-flip: strict-mode default goes from `false` to `true` and `push:` triggers land, gated on a clean Phase 2 measurement run.)
  • #50 — Postmortems published (every incident gets a public writeup at /incidents/<id>) (B1207 — /incidents index + /incidents/[slug] detail + first meta-postmortem entry (channel exists before first real incident).)

How items flip from open to shipped

When a build closes an item:

1. Replace [ ] with [x]. 2. Append a build reference: *(B1234, B1235)*. 3. If the item shipped partially, leave it open and add a sub-bullet describing what landed: - B1234: zod added to forge SSE events; refine + crucible still pending. 4. If the item is GHOST (already shipped before this list existed), mark [x] with *(pre-existing as of B<N>)* and a one-line note.

Every five items shipped: refresh the Tier A summary line at the top so the at-a-glance view stays honest.


Why this list and not the old "Known open items"

The old CLAUDE.md backlog accumulated audit notes from months of context. By Build 1160 it was carrying four shipped-ghost items. The Linear List is prospective discipline, not reactive tracking — every item is a deliberate move toward 100th-percentile, picked because it ratchets a real metric or unlocks a real capability.

The previous ### Known open items block in CLAUDE.md has been deprecated in favor of this file. New session prompts should Read docs/LINEAR-LIST.md first.


Known costs

  • Tier A items 1-5 will each take ~1 day of focused work. The allowlist purges (#1, #6-#8) involve grinding through real code that's been allowed to drift.
  • Tier B items typically span multiple builds — xstate migration (#16) is probably 5-10 builds done carefully.
  • Tier C items have external dependencies (npm package, OpenAPI tooling) and should be planned as small projects.
  • Tier D items are cultural commitments more than code commitments. They land when the team is ready to defend them publicly.

The point isn't to ship all 50. The point is never to coast. Every build either advances this list, or has a written reason it doesn't.