The Punch List
Renamed from "The Linear List" at Build 1216. The activation phraseContinue Linear Liststill works for backwards-compat with old session prompts; the canonical phrase going forward isContinue Punch List. Internal commit subjects from B1173-B1215 retain the "Linear List #N" tag; they're matched by the engineering-report regex alongside "Punch List #N" and "Excellence #N".
Mission: 100th-percentile engineering discipline as a public-facing artifact. The discipline is the moat; this list ratchets it deliberately.
Status: Started Build 1173 (post-extraction-batch). Roll new sessions with this file as the reference.
Activation
When the user types `Continue Punch List` (or the legacy Continue Linear List), the agent should:
1. Read this file. 2. Find the next unchecked item (highest leverage given current state). 3. Scope it as one or more builds. 4. Ship it (typecheck → tests → commit → push). 5. Mark the checkbox here, reference the build number(s).
Order is not strictly sequential. Tier A first, but skip ahead when:
- A Tier B/C/D item is obviously cheaper and unblocks others.
- Surface an explanation when reordering.
Each item carries a one-line scope. If a scope expands beyond ~3 builds, split it into sub-items here first.
Song Surgery — 12-18 build flagship feature (DEFERRED to post-Release-Dossier)
Triggered by: operator-proposed concept + B3148 100-expert / 100-round WAR Room (2026-05-23). Verdict: "Asymmetric upside; bounded downside. Build it" — but NOT until the Release Dossier closes through Phase 8.
The reframe: "Put this song through Song Surgery" replaces "regenerate this song." Same machinery (the Release Dossier engines from Phases 2-6 already do 80% of the work — runAllReportQaGates + computeRevisionRoi + computeLineLevelSurgery + Refine flow); different UX wrapper. The marketing value is high (category-creating positioning: "Don't regenerate your song. Save it.") and the engineering cost is low (the engines exist; what's needed is the workflow + one new Haiku primitive + tier gating).
Tagline candidates:
- "Song Surgery: line-by-line lyric repair until your song is release-ready."
- "Don't regenerate your song. Save it." (the WAR Room favorite)
The 12-18 build sequence (3 phases):
Phase A — Foundation (4-5 builds)
- A1 (B3154) — `judgeRevisionDelta` Haiku primitive shipped.
src/lib/claude/judge-revision-delta.ts+ 18 tests injudge-revision-delta.test.ts. Returns{ verdict: 'improves' | 'changes' | 'risks', dimensions: { specificity, singability, memorability, vulnerability } on -2/+2 scale, rationale, judgeError? }. Fail-open contract: when the Haiku call fails (network / parse / timeout), returnsverdict='changes'+ zero dimensions +judgeErrorset; caller treatsjudgeErroras "unjudged." Short-circuits the Haiku call on identical-input and deletion cases. IncludesaggregateSessionDeltas()helper for end-of-session per-axis totals + net composite. Uses the samegetAnthropicClient()+getModel('focus-group')(Haiku) +temperature: 0pattern as the existing compliance-check + witness-details primitives. The four dimensions map intentionally to the load-bearing Release Dossier axes (specificity → Lyric Quality, singability → Singability, memorability → Hook Clarity, vulnerability → SA#23 AID rule). - A2 (B3155) — `SurgerySession` data model + Supabase table shipped. Two-table schema in
supabase/migrations/create_song_surgery_sessions.sql:song_surgery_sessions(status state-machine + frozen diagnostic snapshot + verification record + final_lyrics + branched_from_session_id for alternate-path branching) +song_surgery_edits(per-edit log with verdict + dimensionDeltas + operator status). Full RLS — operators see only their own sessions. updated_at trigger so timeline UI shows accurate last-touched timestamps. TypeScript types + state-machine helpers insrc/lib/surgery/session-types.ts:SurgeryStatusunion (6 values matching SQL CHECK),SurgeryEditStatusunion (4 values),SurgeryDiagnosticSnapshot,SurgeryVerificationRecord,SurgerySession,SurgeryEdit.canTransition()+isTerminal()enforce the state-machine contract (diagnosing→planning→repairing→verifying→complete, any non-terminal→abandoned).buildEditRecord()converts a judgeRevisionDelta judgment into a DB-row-shaped record;applyEditToLyrics()does 1-based line-number replacement preserving all other lines. 17 tests pin the state machine + the edit-record builder + the lyric-apply helper. - A3 (B3156) — Surgery diagnose builder shipped.
src/lib/surgery/diagnose.tsprovidesbuildDiagnosticSnapshot(song)reading three engines:runAllReportQaGates(B3143) +runTasteSensitivityScan(B3138) +computeLineLevelSurgery(B3147). Returns{ hasFindings, snapshot, cleanMessage }. The load-bearing restraint contract enforced here: when ALL three signals come up clean,hasFindings=falseandcleanMessagecarries the operator-facing copy ("This song doesn't need surgery. Composite X/100, Release Readiness Y/100. Ship it.") instead of an empty Surgery Plan. Three flavors of clean message: high-scoring → "Ship it"; mid-scoring with no flags → "Surgery only triages flagged issues"; unscored → "Score first, then return."prioritizeSurgeryItems()helper enforces the severity-ladder + line-number sort contract callers can rely on. 10 tests pin the restraint contract + the snapshot shape + the prioritization order. The HTTP route layer comes in a forthcoming build — this ships the pure-compute orchestrator that the route will call. - A4 (B3157) — Lyric-diff UI primitive shipped.
src/components/surgery/LyricDiff.tsx. Two render variants: 'card' (vertical, for the per-issue repair surface) + 'inline' (compact one-row, for the timeline). Per-verdict color treatment (improves=green / changes=yellow / risks=red / unjudged=neutral). Shows: Before (strike-through italic), After (bold), verdict chip (top-right), per-axis dimension delta chips (when non-zero), one-line rationale (italic, dimmed). Reusable across the future workflow page (B1) + repair card (B3) + session timeline (B4) + Final Vitals screen (B5). - A5 (B3158) — Verified-rescore trigger + Final Vitals math + trust verdict shipped.
src/lib/surgery/verify.ts. Three functions:shouldVerifyNow()fires the verified rescore whenacceptedSinceLastVerify >= 3OR explicit operator verify OR session closing.buildVerificationRecord()composes the diagnostic snapshot baseline + the session's accepted-edit judgments + the verified post-rescore scores into theSurgeryVerificationRecord(per-axis aggregate dimensions + net composite delta + net readiness delta + verifiedAt).computeTrustVerdict()produces operator-facing copy comparing projected (per-axis sum) vs. verified (composite delta): clean improvement, projection variance flag ("Projected +8; verified +2 — review which axes the edits actually improved"), negative-direction movement honesty ("wrong direction — review the edits"), no-change reporting. VARIANCE_THRESHOLD = 4 (conservative: 2-point gap is noise, 4-point is signal). 14 tests pin trigger rules + Final Vitals math + every trust-verdict path. Phase A CLOSED — A1+A2+A3+A4+A5 all shipped. 41 tests green across all 4 Phase A modules.
Phase B — Workflow UX (CLOSED; shipped 31 builds B3159-B3203)
Build 3272 re-scope (Deep Audit Tier 3 #21). Original header said "5-6 builds." Actual: B1-B6 shipped the core workflow UX (B3159-B3163) — the 5-6 estimate was honest for that scope. Then B21-B31 (the AUDIT-1 through AUDIT-12 follow-on items, B3172-B3203) piggybacked under the same Phase B header rather than opening Phase B2. The 2026-05-24 Deep Audit Section 5 agent flagged this as artifact-drift: the header stopped tracking ground truth, which eroded the planning artifact's load-bearing function. All 31 items are done. This re-scope is a honest-relabel (not a re-open). Future Song-Surgery work opens Phase B2 with a fresh header that matches its scope from day one. The B21-B31 wave's lesson: when a follow-on item set exceeds the original scope by ~3x, open a new phase rather than appending under the old one. The B2435 "one inferred fix per build, max" rule has a sibling here: "one phase per scope estimate, max."
- B1 (B3159) — `/surgery/[songId]` page shipped. Server-component shell at
src/app/surgery/[songId]/page.tsx. Auth viacreateServerSupabase()+ owner-check on thesongsrow; unauth → sign-in CTA; unconfigured-supabase → graceful fallback page; non-owned song →notFound(). ComposesPremiumReportDatafrom the songs row JSONB columns (eval_data / first_listen / prosody_report / focus_group / forge_settings) and runsbuildDiagnosticSnapshot(song)(B3156). Two branches enforce the load-bearing restraint contract: when!hasFindings, renders ONLY the green CheckCircle banner withdiagnosis.cleanMessage+ "Song Surgery only triages flagged issues — restraint is the trust signal" footer (no plan, no list, no polish suggestions). When findings exist, renders the Surgery Plan: counts row (critical/major/minor), QA gate failures section, ordered surgery items with severity color borders (red/orange/yellow), and a "Begin Surgery" footer noting the interactive flow ships in B3.data-testid="surgery-clean-message"pins the restraint branch for e2e tests. Force-dynamic + noindex metadata. - B2 (folded into B3159) — Diagnosis screen shipped as the read-only branch of the `/surgery/[songId]` page. The page renders the prioritized counts row (critical / major / minor), the QA-gate failures section, and the ordered Surgery Plan as part of the same shell. No separate component —
buildDiagnosticSnapshotoutput rendered inline. The "Begin Surgery" placeholder was retired in B3160 when the RepairFlow took over the items list. - B3 (B3160) — Per-issue repair card + RepairFlow shipped.
src/components/surgery/RepairFlow.tsxis the client-component interaction primitive. Single-issue card: severity left-border (red/orange/yellow), Problem header, Current line (italic), Why it matters, Suggested revision (with inline Edit button → editable textarea, max 500 chars, font-serif lyric typography), and four actions: Accept (green; calls /api/surgery/judge-edit, applies the possibly-edited revision, shows loading spinner during judgment), Reject suggestion (the fix is wrong), Skip line (leave the line alone), Continue to next issue (after judgment renders). On accept, the LyricDiff component (B3157) renders inline showing Before/After + verdict chip (improves / changes / risks / unjudged) + per-axis dimension delta chips + the one-line rationale. Session summary view fires when all items exhausted: 4-tile counts grid (Accepted / Skipped / Rejected / Improves) + verdict-timeline of all accepted edits as inline LyricDiffs + footer noting verified rescore + Final Vitals ship in B5. The companion API route shipped alongside:src/app/api/surgery/judge-edit/route.ts— POST, auth-gated, CSRF-validated, 120/hour rate limited via the newsurgeryJudgebucket, 500-char per-line cap. Wires through tojudgeRevisionDelta(B3154) with{ oldLine, newLine, genre? }. Page now ends with the RepairFlow instead of the placeholder block; the diagnose-clean branch is unchanged (restraint contract preserved). - B4 (B3161) — Session timeline component shipped.
src/components/surgery/SessionTimeline.tsxis the reusable version-history surface. Pure presentation (no state, no fetch — parent owns the data + revert callback). Takesversions: TimelineVersion[](each carrying{ label, createdAt, caption?, edits, verifiedComposite?, verified? }) + optionalcurrentLabel+ optionalonRevert(label)callback +compactmode for the dashboard strip. Renders a vertical stack with a left spine line + bullet dots (forge-gold ring on the current version, dimmed circle on the rest). Each version card shows: label + caption + Current/Verified chips + edit count + per-verdict-color count chips (improves/changes/risks/unjudged) + verified composite when present + per-editLyricDiff(inline variant, hidden in compact mode) + Revert button whenonRevertis passed AND the version is not current. CompanionbumpVersionLabel('1.x.y') → '1.(x+1).0'semver helper (Song Surgery sessions bump minor, roll patch to 0). The RepairFlow session-complete view now uses SessionTimeline (via thebuildTimelineVersions(accepted)helper that yields v1.0.0 baseline + v1.1.0 post-session) instead of the ad-hoc accepted-edits list. 8 tests pin the bump-version semantics (undefined → 1.0.0, unparseable → 1.0.0, patch-rollover, major-preservation). - B31 (B3203) — Operator-reported Genre Radio gap closeout: auto-audit on forge-finalize. Operator created a Latin song, uploaded audio, shared it — and it didn't appear on the Latin Genre Radio. Root cause:
/api/genres/[slug]/radiorequires an entry insong_audit_runsataudit_version='1.2.0'with the matchingdetected_arc_slug. That table was populated ONLY by the manualscripts/backfill-song-audits.tsscript. Fresh forges never appeared until the operator re-ran the backfill. Fix in `src/app/api/songs/forge/finalize-song.ts`: inline auto-audit IIFE inside the DB-update success branch. CallsauditSong({id, lyrics, genre, sunoStyleString})(pure-compute; no LLM), inserts the resultingSongAuditRunDraftintosong_audit_runsvia service-role client (RLS on the table is admin-only). Fire-and-forget; load-bearing forge path unaffected if the audit or insert fails.SF_FORGE_AUTO_AUDIT_DISABLED=1env toggle is the emergency rollback. Structured-log events:forge.auto_audit.{disabled_via_env,skipped_no_audit,no_service_role,insert_failed,persisted,threw}. Every new forge now appears on its arc's Genre Radio within 2 minutes (thes-maxage=120cache TTL). Pre-existing songs that never got audited still need the operator's one-time backfill:npx tsx scripts/backfill-song-audits.ts --admin-only. tsc clean.
- B30 (B3202) — AUDIT-8 follow-up: first-forge email trigger in finalize-song. Fires when the user's complete-song count BEFORE the just-finalized one is 0 (= this is their first finished forge). Fire-and-forget; gated by
SF_EMAIL_RETENTION_ENABLEDinside the template. Includes the song's composite score in the email body when present. Structured-log events on count failure / no-email / dispatch / catch.
- B29 (B3201) — AUDIT-8 follow-up: welcome email trigger on onboarding completion. New
POST /api/onboarding/completeendpoint firessendWelcomeEmailfire-and-forget. Client posts to it fromsrc/app/onboarding/page.tsxafterupdateProfile()resolves. Idempotent in practice (onboarding form redirects away once username is set). Gated bySF_EMAIL_RETENTION_ENABLED— until flipped, every trigger no-ops withretention.welcome.gate_closedlogs the operator can grep to verify wiring.
- B28 (B3200) — Surgery WAR ROOM P1 #9 closeout: rate-limit split into 3 lanes. Pre-3200 all 6 surgery endpoints shared one
surgeryJudgebucket (120/hr). A real session = ~1 init + 5-10 per-edit writes + 5-10 judge calls + 0-3 suggest-rewrites + 1 finalize = 12-25 calls EACH. 4-6 sessions/hour locked operators out mid-session AND silently shred the per-edit persistence path (fail-open → produces the "Session record: Not saved" surprise that B3198 now surfaces honestly via the failed-writes banner). Fix in `src/lib/api-auth.ts`: split into 3 RATE_LIMITS entries:surgerySession(30/hr; init + finalize + PATCH/GET session — once-per-session-lifecycle),surgeryEdit(600/hr; per-edit append — load-bearing persistence path, generous for burst-handling),surgeryJudge(200/hr; judge-edit + suggest-rewrites — Haiku calls, bumped from 120). Each endpoint routed to its matching lane:session/init→ surgerySession;finalize→ surgerySession;session/[id](GET + PATCH) → surgerySession;session/[id]/edit→ surgeryEdit;judge-edit→ surgeryJudge (unchanged);suggest-rewrites→ surgeryJudge (unchanged). Each routing change carries aBuild 3200comment naming the rationale per endpoint. tsc clean.
- B27 (B3199) — Surgery WAR ROOM P0 #3 closeout: canTransition() enforcement on finalize UPDATE. Pre-3199 the finalize route's UPDATE-mode at
route.ts:393-402wasupdate({status:'complete'})UNCONDITIONALLY. If the session row was already in a terminal state ('complete' from an idempotency replay, or 'abandoned' because the operator clicked "Start fresh" in another tab), the route silently re-wrote the verification_record + final_lyrics OVER the terminal state — a state-machine bypass that thecanTransition()helper insession-types.tswas supposed to prevent (but no caller enforced). Two-tab race + idempotency-replay both produced silent corruption. Fix in `src/app/api/surgery/finalize/route.ts`: importedcanTransition+SurgeryStatusat module level. UPDATE-mode now SELECTs the currentstatusfirst (with.eq('user_id', auth.userId)for RLS-equivalent owner gate), runscanTransition(currentStatus, 'complete'), and only proceeds when legal. On illegal-transition, logssurgery.finalize.illegal_transition(warn, with fromStatus/toStatus) + nulls sessionId so the UI's "Saved · X" chip surfaces "Not saved" honestly — the operator sees something didn't land. Success-log now carriesfromStatusso the audit trail shows the transition shape. SELECT-then-UPDATE has a TOCTOU race window (operator could abandon in another tab between SELECT and UPDATE) but it's at the audit/failure-mode tier, not the "single tab happy path" tier; the in-code comment names a future hardening pass that moves the check into a Postgres CHECK constraint on status transitions. 17/17 session-types tests still green (canTransition test surface unchanged); tsc clean.
- B26 (B3198) — Surgery WAR ROOM P1 #10 closeout: appendLiveEdit failure surfacing. The audit's #2 top-leverage fix. Pre-3198,
appendLiveEditfires across 5 call sites (onAccept / onSkip / onReject / onUndoJudgmentAndSkip + resume hydration) as fire-and-forget — failures only logged aconsole.warn. Operators walked through the whole session unaware their edits weren't persisting; "Session record: Not saved" surprised them 5+ minutes later at session-close, by which point the audit trail was already corrupted (UPDATE-mode finalize writes verification_record + final_lyrics but does NOT bulk-insert the missing edits). Fix in `src/components/surgery/RepairFlow.tsx`: newfailedEditIndicesstate queue captures every failed write (network error / HTTP non-2xx / session-init failed); newretryingFailedEditsboolean tracks the retry-in-flight state.appendLiveEditnow queues the payload on every failure path (including the session-init-failed case where we don't have an editIndex yet — queued with placeholder -1, retry assigns a real index on re-fire). NewretryFailedEdits()function walks the queue, re-fires each write, keeps failures in the queue + drops successes. Sticky banner UI renders ABOVE the resume banner + progress strip whenever the queue is non-empty: yellow border, AlertCircle icon, copy distinguishes 1 vs N edits, explains the contract honestly ("in-memory session is intact, contributes to Final Vitals, but won't be resumable from dashboard until X succeeds"), single primary "Retry failed writes" button (RotateCcw icon, loading state with Loader2 spin).role="status" aria-live="polite"so screen readers announce new failures. 93/93 surgery tests green; tsc clean. Closes the "Session record: Not saved" surprise — operators see failures at the moment they happen, not after a 10-edit walk.
- B25 (B3196) — Song Surgery WAR ROOM audit + P0 #6 fix shipped (rescorePending stuck-state). Operator screenshot showed a completed surgery session with the Release Clearance chip stuck on "VERIFYING" + the Verified pane chip stuck on "Rescoring" (spinning Loader2). Root-cause: the server-side
/api/surgery/finalizeroute atroute.ts:339setsrescorePending = verified === null— meaning "verified composite is null" (TERMINAL failure / no-edits-to-verify / env-disabled state). Pre-3196FinalVitals.tsxlines 110-117 + 215-220 treatedrescorePending=trueas in-progress — rendered spinning loaders + "Verifying" + "Rescoring" copy on a finished session. The route had already returned; the rescore was DONE; the UI lied about that. Operator's audit ask: WAR ROOM deep audit of all Song Surgery code. Dispatched a 100-expert-style Agent against ~5,400 LOC; produced 25 findings across 4 tiers documented indocs/SURGERY-WAR-ROOM-AUDIT-B3196.md. Highlights: P0 #1 (clearance chip falls to "No net change" when verified ≠ null but startingComposite = null); P0 #2 (resume hydration silently rewrites'unjudged'verdicts to'changes'+ fabricates zero dimensions); P0 #3 (finalize UPDATE-mode skipscanTransition()— two-tab race produces silent state-machine corruption); P0 #4 (liveEditIndexclosure-capture race produces duplicate edit_index rows); P0 #5 (resumecurrentIdxadvance is wrong whensurgeryItemIdis null — puts operator back at item 0); P1 #10 (theappendLiveEditfire-and-forget pattern is the load-bearing weak link producing the "Session record: Not saved" surprise — silent data loss). Also: ZERO test coverage on RepairFlow.tsx (1802 LOC), FinalVitals.tsx (334 LOC), LyricDiff.tsx, AND all 6 API routes. Architecture observations: state machine split across 3 layers with no single owner; 3 persistence modes routed by implicit client state; fail-open signals inconsistent across 5 fetch sites. Top 5 fixes by leverage identified for the next 6-8 builds. Ship in B3196: collapsed the FinalVitals.tsx clearance ladder + Verified-pane chip into the!verifiedAvailablebranch with "Projection only" copy + Sparkles icon + a tooltip explaining the terminal cause; removed the Loader2 import (now unused). Long inline comment names the rename for a future build + cites SURGERY-WAR-ROOM-AUDIT-B3196.md. NEWsrc/components/surgery/FinalVitals.behavior.test.tsx— 10 tests pin the B3196 fix: "Projection only" chip renders when rescorePending=true (with no spinner / no "Rescoring" copy / no "Verifying" copy); same chip when rescorePending=false; no.animate-spinelements at all; clearance chip is not "Verifying"; all 4 verified-available chip states render correctly; dimension scoreboard always shows the 4 axes + net delta. First test coverage for FinalVitals + first surgery-component behavior test ever. 10/10 green; tsc clean.
- B5 (B3162) — Final Vitals + Release Clearance screen shipped.
src/components/surgery/FinalVitals.tsxis the end-of-session presentation component. Reads aSurgeryVerificationRecord(B3155 type) +VerificationTrustVerdict(B3158 type) + the baseline composite + readiness scores, then renders three scoreboards: (1) Projected — per-axis dimension sum (specificity / singability / memorability / vulnerability) with up/down/neutral trend icons + net delta header; (2) Verified — composite + Release Readiness paired panes showing start → end → delta, with "Pending / Rescoring" chip when verifiedComposite is null; (3) Trust verdict callout — full operator-facing message fromcomputeTrustVerdict()with yellow-bordered variance treatment when projected >> verified; (4) Release Clearance chip computing the structural state (Cleared / Mixed result / Wrong direction / No net change / Verifying / Projection only). Companion API endpoint atsrc/app/api/surgery/finalize/route.ts— POST, auth-gated, CSRF-validated, rate-limited via the surgeryJudge bucket. Composes the verification record from accepted-judgment dimension deltas usingbuildVerificationRecord()+computeTrustVerdict(). Projection-only path: verifiedComposite + verifiedReadiness ship as null; the actual verified rescore (re-evaluating the rewritten lyrics via the eval engine — ~$0.10/call) wires through in a follow-up build. The Final Vitals screen handles the null case cleanly via the rescorePending flag. RepairFlow now: (a) auto-fires the finalize call when the session reaches completion via useEffect; (b) renders FinalVitals + SessionTimeline + summary tiles in the completed view; (c) shows loading + error states cleanly. Phase B B1+B2+B3+B4+B5 all shipped — only B6 (revert + branching) remains. - B24 (B3193) — Critical bug fix: Final Vitals auto-abort race. Operator-reported: spinner stuck at 196s+ with no Retry surfacing. B3190 thought it added a 75s client-side timeout. It DID add one, but never actually got to fire because of a useEffect dep-array race that's been latent since B3170. Trace: the finalize useEffect had
finalizingin its dep array. The effect body callssetFinalizing(true)near the top. React processes the state update → re-render → useEffect deps compared →finalizingchangedfalse → true→ CLEANUP function of the just-started run fires → cleanup callscontroller.abort()(B3190 addition) +clearTimeout(abortTimer)(B3190) +cancelled = true(B3170). The fetch then aborts within microseconds of starting..catchseesAbortErrorbutcancelled=truereturns early..finallychecks!cancelled→ false →setFinalizing(false)never runs. State is stuck:finalizing=trueforever, spinner spins forever, the 75s timer was killed in the same cycle, no path back. Operator saw 196s+ elapsed with no Retry button (becausefinalizeErrorwas never set — the.catchreturned early). Fix in `src/components/surgery/RepairFlow.tsx`: removefinalizingfrom the useEffect dep array (witheslint-disable-next-line react-hooks/exhaustive-deps+ a long-comment explaining why). The guard inside still readsfinalizingvia closure — which is fine because the closure captures the value at effect-scheduling time, and the only way the effect re-fires legitimately is when a different dep changes (e.g. operator clicks Retry →finalizeErrorchanges from non-null → null), at which pointfinalizingSHOULD befalseanyway (.finallyset it). NowsetFinalizing(true)doesn't trigger a re-fire, the cleanup doesn't run mid-fetch, the abort timer survives to fire at 75s, and the catch path properly surfaces the timeout error + Retry button. 83/83 surgery tests green; tsc clean. The pre-B3190 code had the same latent race but was harmless because the cleanup only setcancelled=truewithout callingcontroller.abort(); the fetch ran to completion (eventually) and.finally's!cancelledguard suppressed the state update. Bug surfaced when B3190 added the abort call — operators with slow finalize calls would have hit it. Sacred Accident candidate: "An effect's deps array should never include a state variable the effect mutates." Future builds should hold this line.
- B23 (B3192) — Operator-requested: post-judgment back-out paths in Song Surgery. Operator hit the case where they picked rewrite #1 (an R&B chorus line), got the judgment ("RISKS" with -1/-1/-2 dimension deltas + a rationale explaining the loss of emotional choreography), and had no way to undo. The only post-judgment action was "Continue to next issue," which commits the un-improved edit. They asked for "go BACK and change to another option or SKIP." Fix in `src/components/surgery/RepairFlow.tsx`: two new client-side functions + a refreshed post-judgment button row.
onUndoJudgmentToEdit()pops the last in-memoryacceptedentry, clearslastJudgment(returns the card to its action-row state), and flipsisEditing=trueso the textarea reopens populated with what they tried — operator can pick a different ranked suggestion (still rendered above) or edit further.onUndoJudgmentAndSkip()pops the accept + records a skip + firesappendLiveEditwithstatus='skipped'+ advances. The post-judgment button row is now aflex justify-betweenwith two back-out buttons on the left (Try different revision —RotateCcwicon + dark-800 styling; Skip this line instead —SkipForwardicon + transparent styling) + the existing primary "Continue to next issue" forge-500 button on the right. Visual hierarchy keeps the happy-path primary dominant; back-out is available but de-emphasized. DB persistence trade-off documented in-code: the B3185 per-edit endpoint already wrote the accept row when the judge call returned, and we leave that row in place — the session timeline + activity feed accumulate an honest "tried + reverted" history. The verification record composed at finalize() reads the in-memoryacceptedarray, so the metrics reflect only the operator's final decision, not their first try. A future build could add a DELETE endpoint that cleans the superseded edit row; the load-bearing UX (operator can undo) ships here without it. 83/83 surgery tests green; tsc clean.
- B22 (B3191) — Operator-reported follow-up: forge-result autoscroll didn't go "ALL the way up." B3188 shipped autoscroll via
resultRootRef.scrollIntoView({ block: 'start' })— aligns the result root's TOP with the viewport top. Operator reported they were still landing partway down on the analysis sections instead of at the title + score + lyrics. Likely causes: (a) lazy-mounting post-result cards (SamplePlayer, cover art, wound summary, heat density) shift the page mid-animation and the single 50ms scroll loses the race; (b)block: 'start'on a nested element can leave the page short of true top when there's a wrapper above it. Fix in `src/app/forge/v2-design/ForgeV2Result.tsx`: switched fromscrollIntoViewtowindow.scrollTo({ top: 0, left: 0, behavior: ... })— unambiguous "go to the very top of the page" — plus a safety re-fire at 350ms to catch layout shifts that happen after the first scroll lands (~300ms after the smooth-scroll completes). Smooth by default;'auto'(instant) underprefers-reduced-motion. Try/catch fallback towindow.scrollTo(0, 0)for older browsers without the options-bag API. Removed the now-unusedresultRootRefref + its attachment on the outer<div>. SSR-safe viatypeof window === 'undefined'guard. 11/11 ForgeV2Result behavior tests green; tsc clean.
- B21 (B3190) — Operator-reported: "Composing Final Vitals…" hang fix. Operator reported the spinner circling for several minutes with no progress indication after completing a 7-item surgery session. Root-causing: the server-side
/api/surgery/finalizeroute caps atmaxDuration=60with a 30s internal verified-rescore timeout — so worst-case it should return in ~60s. But the client-side fetch had NO AbortController + NO wall-clock cap, so if the Vercel lambda hard-killed without flushing a response (rescore + DB writes pushed close to the 60s ceiling), the fetch hung indefinitely with the spinner spinning. Three coordinated fixes: (1)src/lib/surgery/verified-rescore.ts: lowered the default internal timeout from 30s → 20s. Server-side margin under the 60s lambda ceiling jumps from ~30s to ~40s for DB writes + response — dramatically reduces the chance of a hard-kill mid-response. Trade-off documented: more sessions fall back to projection-only Final Vitals when the eval is on the slow end, but that's an honest fail-open vs. a hang. Test updated. (2)src/components/surgery/RepairFlow.tsx: client-sideAbortController+ 75s wall-clock cap on the finalize fetch (= 60s server ceiling + 15s grace). On timeout:AbortErrorcaught in.catch,finalizeTimedOutstate flips true, error message reads "Final Vitals took longer than 75 seconds to compose. Your edits are already saved — click Retry to try again, or refresh and resume the session from your dashboard." Cleanup function clears the timer + aborts the controller on unmount. (3) Progressive status messages: newfinalizeElapsedstate + companion useEffect ticks once per second while finalizing. UI swaps text at thresholds: 0-20s → "Composing Final Vitals…"; 20-45s → "Still working — verified rescore in progress…"; 45+s → "Taking longer than usual — finishing up…". Sub-line shows "Ns elapsed" + ", the eval engine is re-scoring your rewritten lyrics" after 20s. Operator can now distinguish "alive but slow" from "hung." (4) Retry button in the error block:<RotateCcw />icon + "Try again" label; click clearsfinalizeError+finalizeTimedOut, which (via the useEffect dep array) re-fires the finalize fetch. Per-edit persistence (B3185) guarantees the operator's edits are already in the DB; retry only re-runs the verification composition + verified rescore. UPDATE-mode in the finalize route handles re-finalizing a session row that's already 'complete' idempotently. 83/83 surgery tests green; tsc clean.
- B20 (B3189) — Operator-requested: 3 ranked rewrite suggestions per surgery item. Operator feedback after B3179: "The Suggested Revision is direction, but is this what goes into the song? Actual line revisions need to be suggested — would be good if 3 suggestions were made in ranked order." NEW
src/lib/claude/suggest-line-rewrites.ts— Haiku-driven rewrite synthesizer. Input: oldLine + problem + whyItMatters + guidance + optional section/genre/centralTension. Output: 3 ranked singable rewrites with one-line craft rationale each (rank 1 = strongest = "the rewrite a working songwriter would actually use"; rank 2 = viable alternative with a different angle; rank 3 = interesting departure — more poetic / concrete / direct than 1+2). Strict JSON contract; clamps rank to 1..3; sorts ascending; drops empty + over-500-char lines; caps at 3 rewrites total. 15s wall-clock timeout via Promise.race against TIMEOUT_SENTINEL — long Haiku tails fall back to empty-rewrites + the operator writes their own. Fail-open at every error path. NEW/api/surgery/suggest-rewrites/route.ts— POST endpoint wrapping the helper. Auth + CSRF + 120/hr rate-limit via surgeryJudge bucket + 500-char oldLine cap + maxDuration=30.RepairFlow.tsxupdated: new state (suggestions / suggestionsLoading / suggestionsError) cleared on advance() so each card opts in independently; newonSuggestRewrites()lazy-fetches the 3 alternatives; newonUseSuggestion(line)fills the textarea + opens edit mode. UI inside the existing yellow-bordered "Repair guidance" block: "Suggest 3 rewrites" button (forge-styled, Sparkles icon) → loading spinner with "Asking war room…" copy → 3 numbered cards stacked, rank-1 card gets forge-500 border + bg + "top pick" sub-label, ranks 2-3 get neutral dark-700 styling. Each card shows the line in serif + the one-line rationale beneath in dimmed italic. Clicking ANY card fills the textarea + opens edit mode so the operator can tweak before Accept. Cost envelope: ~$0.0001/click, lazy-fetched per card, ~$0.0003 for a heavy 3-issue session that uses the feature on every card. 17 new tests on the helper (prompt composition + every parse edge case: chatty prefix tolerance, rank clamping, sort order, empty-line drop, 500-char cap, 3-item cap, null returns on garbage / missing field / non-array / unparseable JSON / missing rationale). 100 tests green across surgery + claude. tsc clean. - B19 (B3188) — Forge result auto-scroll-to-top shipped (operator-reported UX bug). When a song finished forging, the result component mounted inline below the composing/forging surface — operators who'd scrolled down to read forging-phase commentary landed on the Wound Summary or Rhythm panel or even the "About the Forge" marketing block at the bottom, not the headline result card. Fix in
src/app/forge/v2-design/ForgeV2Result.tsx: newresultRootRefattached to the result root<div>+useEffectkeyed onforgeResult.idthat firesscrollIntoView({ behavior: 'smooth', block: 'start' })after a 50ms delay (lets the celebration fireworks paint cycle settle so the scroll target stays stable). Honorsprefers-reduced-motion— switches tobehavior: 'auto'(instant) under that media query so motion-sensitive operators don't get a jarring slide. FallbackscrollIntoView()(no options bag) for older browsers. Dep array includesforgeResult.idso "Forge again" → new song id → fresh scroll. SSR-safe viatypeof window === 'undefined'guard. Tsc clean. - B18 (B3186+B3187) — Song Surgery session resume shipped. Closes the loop on the per-edit persistence work (B3184+B3185) — operators who crash mid-session can now actually CONTINUE from where they left off. B3186 server: NEW
/api/surgery/session/[id]/route.tsships GET (returns the session row + all edit rows in order, RLS-gated) and PATCH (transitions status='repairing' → 'abandoned' for "Start fresh"; usescanTransition()from session-types to enforce the state-machine contract; rejects any other status). Page-level detection added to/surgery/[songId]/page.tsx: queriessong_surgery_sessionsfor the operator's most-recentstatus='repairing'row for THIS song, passes its{id, createdAt, updatedAt}summary down asresumableSessionprop to RepairFlow. B3187 client wire: RepairFlow accepts the newresumableSession?prop + adds 3 new state vars (resumeChoice: null | 'resumed' | 'fresh'+resumeBusy+resumeError). NewonResume()handler fetches the saved edits, hydratesaccepted+skippedarrays from the rows (mapping verdict + dimensionDeltas + rationale faithfully; unjudged status preserved via judgeError sentinel), advancescurrentIdxpast every item that already has a recorded edit (matched bysurgery_item_id), and setsliveSessionIdso subsequent edits append to the same session row. NewonStartFresh()handler fires the PATCH to abandon the old session + flips local state. New<History />icon resume banner renders ABOVE the progress strip whenresumeChoice === nullANDresumableSessionexists. Shows last-touched timestamp + two CTAs ("Resume previous session" — forge-500; "Start fresh" — dark-800 secondary). Resume errors surface inline with<AlertCircle />. 142/142 tests green; tsc clean. The full crash-recovery loop is now closed end-to-end: write (B3184+B3185) → detect (B3186 page query) → offer (B3187 banner) → hydrate (B3187 onResume) → continue (existing flow). Six builds total (B3170+B3184+B3185+B3186+B3187) to take Song Surgery from "session lost on tab close" to "session resumable from any device." - B17 (B3185) — Per-edit Song Surgery persistence: client wire shipped. Tab-crash data-loss gap CLOSED end-to-end.
RepairFlow.tsxgains: (a) two new state vars —liveSessionId: string | null(the row id from the lazy-create call) +liveEditIndex: number(1-based, increments per write); (b) two new async helpers —ensureLiveSession()which fires POST/api/surgery/session/initonce per session lifecycle + caches the result, andappendLiveEdit(payload)which fires POST/api/surgery/session/[id]/editfor every action. Both helpers are fail-open: on network/server failure theyconsole.warn+ return null/void; the in-memory state continues to work. The three action handlers (onSkip/onReject/onAccept) now callvoid appendLiveEdit(...)immediately after updating the in-memory accumulator. The finalize POST body now passessessionId: liveSessionId, triggering the B3184 UPDATE-mode path (no double-insert; finalize just transitions status='complete' + writes verification_record + final_lyrics).onRevertToBaselineclearsliveSessionId+liveEditIndexso each branch attempt creates a fresh session row (preserves the branching audit trail). useEffect dep array gainsliveSessionIdso a late-arriving session id (slow ensureLiveSession) doesn't get missed by the finalize trigger. The "tab-crash mid-session = total data loss" gap is now closed. Tab-crash loses at most ONE edit (the mid-flight one), not the whole session. 142/142 tests green; tsc clean. Per-edit persistence is the last load-bearing Song Surgery follow-up I flagged when declaring the feature shipped — both gaps now closed (B3170+B3184+B3185 for persistence; B3171 for verified rescore). - B16 (B3184) — Per-edit Song Surgery persistence: server endpoints shipped. Closes half of the "tab-crash mid-session = total data loss" gap operator-flagged when declaring Song Surgery shipped. NEW endpoint
/api/surgery/session/init/route.tslazily creates asong_surgery_sessionsrow with status='repairing' on the operator's first action (Accept / Skip / Reject). Authenticated session client; RLS gates per-user; returnssessionId. Body:{ songId, diagnosticSnapshot, branchedFromSessionId? }. NEW endpoint/api/surgery/session/[id]/edit/route.tsappends ONE row tosong_surgery_editsper call. Validates editIndex / lineNumber / status (accepted/rejected/skipped/proposed) / verdict / dimensionDeltas / 500-char per-line cap. Session ownership enforced via the existingown_session_edits_insertRLS policy. MODIFIED/api/surgery/finalize/route.ts— now branches on whetherbody.sessionIdis present. UPDATE mode (B3184): existing session row getsstatus='complete'+ verification_record + final_lyrics set via.update().eq('id', sessionId); edit rows skipped (they're already there from the per-edit endpoint). LEGACY INSERT mode (B3170 backwards compat): no sessionId in body → creates session + bulk-inserts edits in one shot as before. Both modes emit thesurgery.session.completedaudit event for the dashboard activity feed (B3178 + B3182 wired); UPDATE mode passes null for accepted/skipped counts because the per-edit endpoint wrote them one at a time and they're not aggregated at the finalize layer. The B3185 client wire that calls these endpoints ships in the next build. - B15 (B3183) — Inverse-mode UI for VaultInspirationPanel shipped (consumer wire for B3168 primitive).
src/components/forge/VaultInspirationPanel.tsxgains a per-card "Like / Avoid" toggle button next to the existing arc + score + similarity chips. Two new states:invertedIds: Set<string>(mutual-compatible with the existingexcludedIds— a card can be inverted AND included, or excluded entirely) + reset alongsideexcludedIdson every successful fetch (operator opts in per-card, doesn't carry across prompts). TheonReferencesChangemapping now setsinverse: invertedIds.has(ex.id)per ref so the parent's prompt-prefix builder routes each card to either the EXEMPLAR block or the ANTI-PATTERN block per the B3168 primitive. Visual treatment: when a card is inverted, its border + background flips from the amber exemplar styling to a yellow-bordered anti-pattern treatment (matches the B3168VAULT ANTI-PATTERNSblock accent). Toggle button states: "Like" (amber, default) ↔ "Avoid" (yellow). Disabled when the masteruseAsReferencestoggle is off OR the card is excluded (flipping under those conditions has no downstream effect). Accessible labels + title tooltips explain both states. The Sister-song panel pattern from the punch list spec realized — operator can now say "depart from THIS voice" in addition to "honor THIS voice." Inverse-mode primitive (B3168) ships its consumer wire here; the discovery → composition → UX loop for inverse references is closed end-to-end. - B14 (B3182) — Dashboard activity-feed consumer for surgery.session.completed shipped (closes the B3178 write-without-renderer gap).
src/app/dashboard/activity/derive-events.tsgains a'surgery_session'AuditEventKind + a case inmapAuditRowToEventfor'surgery.session.completed'. The mapping handles the surgery audit-event's unusual shape (subject_id = SESSION id, not song id) by pullingpayload.songIdfor the parent-song back-link +titleBySongId.get(parentSongId)for the song title. Detail copy adapts to data presence:"3 edits accepted · verified +5 composite"when the rescore landed,"3 edits accepted · projected +7 per-axis"when not,"3 edits accepted"when no deltas, with correct singular grammar on1 edit accepted.src/app/dashboard/activity/ActivityFeed.tsxKIND_METAgains asurgery_sessionentry:<Scissors />icon + "Song Surgery" label +tonal-forgechip variant (matches the Song Surgery brand surfaces across /surgery, RepairCard, dashboard CTA banner). 6 new tests pin the mapping cases (kind resolution, payload-songId routing, verified-delta detail, negative-direction sign, projected-fallback, singular grammar). 46/46 activity-feed tests green; tsc clean. The dashboard activity feed now renders Song Surgery sessions alongside forge/score/gauntlet/share/etc events — the discovery + audit loop is closed end-to-end. - B13 (B3180+B3181) — Two operator-reported bugs fixed. B3180 — Story Twist firing on every song:
detectGenreMode()insrc/lib/claude/genre-modes.tsusedString.includes(keyword)to test trigger keywords. Story Twist's trigger list included'turn'— a substring match on which fires Story Twist for "Saturn", "return", "nocturnal", "turning point", "you turn me on" — and any other word containing the letters t-u-r-n. Plus'narrative'+'story'are extremely common English words that triggered indiscriminately. Plus Story Twist is registered HIGH in the priority order (line 1082 comment confirms it). Net effect: most songs got tagged Story Twist regardless of their actual narrative intent, and the mode chip on the dashboard reflected that misdetection. Fix: switched the matcher fromhaystack.includes(keyword)to a word-boundary regex (\b{escaped-keyword}\b, case-insensitive). Regex special chars escaped (period in "tom t. hall" + bracket-class metas). Regexes cached lazily — first call compiles, subsequent calls reuse. 11 new regression tests pin the fix: Saturn / return / nocturnal / turning / burning DON'T fire Story Twist; the literal word "turn" / "twist" / "story song" / "tom t. hall" DO fire it. 50/50 genre-modes tests green; 39 pre-existing + 11 new. Affects EVERY song forged post-deploy + every dashboard chip rendered fromsong.mode. B3181 — Cover image in PDF report:PremiumReportDatatype was missing thecoverArtUrlfield even thoughDashboardSongcarries it from the songs row'scover_art_urlcolumn. The PDF template (src/lib/export-song-html.ts) had zero references to cover/image and never rendered one. Fix: addedcoverArtUrl?: stringto PremiumReportData (auto-threads throughSongActionsButtonRow.buildReportData()via the...songspread). PDF header now renders a 1.4-inch square cover thumbnail to the RIGHT of the title block via a 2-column grid layout (title + meta + diag + dates on the left, cover on the right). When coverArtUrl is missing/empty, the header falls back to the original single-column layout — no broken image icons.page-break-inside: avoidkeeps the cover + title together. 201 tests green across genre-modes + surgery + report-sections. - **B12 (B3179) — Operator-reported bug fix:
suggestedRevisionwas shipping prosody RULES as if they were lyric rewrites; RepairCard pre-filled the textarea with the instruction so Accept replaced the lyric with the rule (operator demo'd on a Latin song where "End on an open vowel..." replaced "Aquí en el marco de mi puerta"). The judge primitive correctly flagged it asrisks("destroying the actual lyric") but the dashboard should never have offered the instruction as a pre-filled revision in the first place. Fix:SurgeryItemtype inline-level-surgery.tsgains arevisionType: 'guidance' | 'rewrite'discriminator. All 4 current source pullers (prosody-fatal/prosody-flag/taste-flag/heat-map-dead) marked as'guidance'— they only have rules, not synthesized rewrites. RepairCard branches on the discriminator: when guidance, renders the suggestedRevision as a yellow-bordered HINT block labeled "Repair guidance" with the explanation "This is craft direction, not a singable line. Use it to write your own revision below." — and pre-fills the textarea with the ORIGINAL line (so the operator edits the real lyric, not the rule). When rewrite, keeps the legacy pre-fill behavior. Accept button now also disabled when the textarea still equals the original line (no-op guard). Optional discriminator inSurgeryDiagnosticSnapshot.surgeryItemsfor backwards compat with sessions persisted pre-B3179 — consumers default to 'guidance' (safer; won't auto-replace). 99/99 tests green; tsc clean. The B3148 restraint contract ("Song Surgery must be ABLE to refuse to suggest a change") is now mechanically enforced AT the UI layer, not just at the diagnose layer. - B11 (B3176+B3177+B3178) — Song Surgery Tier-3 discovery shipped (homepage callout + /surgery index + activity feed). B3176 — Homepage callout:
src/app/page.tsxgains a new 3.6 section between HomepageBeforeAfter and the Leaderboard. Forge-styled centered card with the load-bearing tagline ("Don't regenerate your song. Save it." with the forge-gold "Save it."), one-paragraph value-prop (per-issue verdicts + per-arc Studio modes + verified Release Clearance + the restraint contract callout), and dual CTAs (See how Song Surgery works → /song-surgery; Forge a song first → /forge). Refine and Surgery now sit on the homepage as siblings — closes the mental-model gap between the two post-write operations. B3177 — `/surgery` index page: NEW server-component atsrc/app/surgery/page.tsx. Reads the operator's last 20 song_surgery_sessions ordered by updated_at DESC + joins songs for title/genre + bulk-fetches edit-status rows in one IN-query. Renders per-session cards: song title, status chip (complete/abandoned/in-progress/verifying with color-coded variants), accepted-skipped counts, verdict-color count chips (improves/changes/risks), verified composite delta with green-arrow/red-arrow indicator when the rescore landed, "Open Surgery on this song →" backlink. Empty-state CTA explains the 4-step path to first session. Tier-aware: Free/Creator see the "Session history is a Professional feature" upgrade panel with the SURGERY_LOCK_FEATURES list + /pricing CTA. Unauth users see a sign-in screen with marketing-landing fallback link. RLS on song_surgery_sessions gates per-user access automatically.data-testid="surgery-session-list"+data-testid="surgery-session-row"pin for e2e. B3178 — Activity-feed event:'surgery.session.completed'added toAuditEventTypeinsrc/lib/audit-events.ts./api/surgery/finalizenow firesrecordAuditEvent(fire-and-forget, service-role client) on successful session persistence — payload carries songId, accepted/skipped counts, netComposite, netReadiness, netDimensionScore, verifiedAvailable flag, branchedFromSessionId. Powers the dashboard activity feed when the operator returns to /dashboard. The dashboard activity-tab consumer wire is a thin follow-up that picks up this event-type alongside the existing 7 types. Tier 3 closed. All discovery moments wired (B3172-B3178 = 7 builds across 7 surfaces). - B10 (B3174+B3175) — Song Surgery Tier-2 discovery shipped (navbar + forge result handoff). B3174 — Navbar entry:
src/components/Navbar.tsxTOOL_LINKS(the Browse dropdown) gains a "Song Surgery" entry with<Scissors />icon between Examples and the Heirlooms surface. Hint text matches the marketing positioning: "Line-by-line lyric repair until your song is release-ready. Don't regenerate — save it. Pro." Placed in the dropdown rather than top-level NAV_LINKS to respect the B2566 4-item buyer-action discipline — Surgery is an authenticated post-write operation, not a buyer-action entry. The B3172 dashboard banner is the primary discovery path; the dropdown is the direct-nav fallback. B3175 — Forge result handoff:src/app/forge/v2-design/ForgeV2Result.tsxaction button row gains a "Song Surgery →" text-button (mirroring the Refine button's polarity) between "Forge again" and "Refine this →". Renders only whenforgeResult.idis present (saved songs). The /surgery page handles the restraint contract internally — when nothing's flagged, the operator sees "ship it" not an empty plan. Two post-forge operations now read as siblings (Refine = polish, Surgery = triage), not competitors.data-testid="forge-result-open-surgery"pins for e2e. Tier-2 discovery closed. - B9 (B3172+B3173) — Song Surgery discovery shipped (the door is finally on the building). Before this build, /surgery/[songId] had ZERO in-product entry points; the only way to reach it was to manually type the URL with a UUID. B3172 — Dashboard CTA: every owned
status='complete'song onsrc/app/dashboard/SongDetail.tsxnow renders a prominent forge-styled banner ("Song Surgery — Line-by-line lyric repair until your song is release-ready. Don't regenerate — save it.") with a<Scissors />icon + an "Open Song Surgery" button linking to/surgery/{song.id}. Hidden when the song lacks lyrics or status (nothing to operate on). Banner sits between SongVisibilityBanner and RevisionHistoryPanel — load-bearing position. The /surgery page itself handles the restraint contract (renders cleanMessage when no findings exist), so we don't gate on diagnose output here.data-testid="song-detail-open-surgery"pins for e2e. B3173 — Marketing landing CTAs:/song-surgeryhero "Open Song Surgery" button previously linked to/pricing— actively hostile to paid users who already had the feature. Now links to/dashboard?tab=songs(where the operator picks a song + hits the B3172 banner). Secondary "Forge a song first" already correctly went to /forge; preserved. The FINAL CTA at the bottom of the page ("Save your song" / "Upgrade to Professional") still links to /pricing — that's the deliberate tier-upgrade path, not the discovery path. Two surfaces, two intents, no more dead-ends. - B8 (B3171) — Verified rescore wiring shipped (Song Surgery's second follow-up gap closed).
src/lib/surgery/verified-rescore.tsships the pure-async wrapper aroundevaluateLyrics+computeReleaseReadiness. Takes the rewritten lyrics + optional genre + parent-song shape; returns{ composite, readiness, evalData, durMs }or null on any error / timeout / disabled state. 30-second wall-clock cap viaPromise.raceagainst a TIMEOUT_SENTINEL; longer eval calls fall back to projection-only rather than hanging the operator.SF_SURGERY_VERIFIED_RESCORE_DISABLED=1env toggle for emergency rollback. Input guards: empty lyrics → null, lyrics > 50KB → null, parse failure → null, missing compositeScore → null. Wired into/api/surgery/finalize/route.tsBEFORE the verification-record compose so the persisted session row carries the REAL verified scores (not the projection-only nulls).rescorePendingresponse flag is now data-driven (verified === null) instead of the hardcodedtruefrom B3162. Skipped when no accepted edits OR finalLyrics missing — nothing to verify. Vercel routemaxDurationbumped to 60s to accommodate the eval call. 7 tests pin the env toggle + input guards; live-eval path covered by integration. The B3148 WAR Room load-bearing rule ("the verified rescore is the truth") is now actually true — the Final Vitals screen shows real verified numbers when the rescore succeeds, projection-only when it fails, and the trust verdict reports honest variance when the two diverge. Both Song Surgery follow-up gaps now closed. The feature ships without caveats. - B7 (B3170) — Session persistence shipped (Song Surgery follow-up gap closed).
/api/surgery/finalizeextended at B3170 to ALSO write the session row + each accepted/skipped/rejected edit to the B3155song_surgery_sessions+song_surgery_editstables. Uses the user's authenticated session client (NOT service-role) so RLS applies — the schema's INSERT policy requires auth.uid()=user_id, preserving the per-user audit chain. Fail-open contract: any persistence failure (column missing, RLS deny, network blip) is logged but the route still returns the computed verification record + verdict; sessionId comes back null and the client renders an "unsaved state" hint. New request body fields:songId,acceptedEdits[](full records: itemId+oldLine+newLine+lineNumber+judgment),skippedEdits[](itemId+lineNumber+reason),finalLyrics,diagnosticSnapshot,branchedFromSessionId. New response field:sessionId. RepairFlow updated: accepts newbaselineLyrics+diagnosticSnapshotprops from the page, computesfinalLyricsby reducingapplyEditToLyricsacross all accepted edits in accept-order, sends the full session-state payload on auto-finalize, renders a green "Session saved · {short-id}" chip when persistence succeeded or a muted "Not saved" chip with explainer tooltip when it didn't, and clears the persisted-session id on revert (so branch attempts get their own session rows). 50-edit caps + 500-char per-line caps on the server. The "in-memory only" follow-up gap I flagged when declaring Song Surgery complete is now closed. - B6 (B3163) — Revert-to-version + branching shipped — Phase B CLOSED. RepairFlow gains in-memory branching state: each completed path that the operator chooses to revert from gets archived into a
branches: SessionBranch[]array; the active branch label bumps viabumpVersionLabeleach revert (1.1.0 → 1.2.0 → 1.3.0 → …). The completed-session view gains: (a) a "Try a different path" CTA card with<GitBranch />icon +<RotateCcw />button callingonRevertToBaseline()(resets currentIdx, accepted, skipped, lastJudgment, finalRecord, finalVerdict, finalRescorePending, finalizeError, isEditing, judgeError, pendingJudgment — then bumps the branch label); (b) the SessionTimeline now wiresonRevert={(label) => label === '1.0.0' && onRevertToBaseline()}so the per-version Revert button in the timeline shares the same behavior; (c) when branches exist, a newBranchComparisonPanelrenders above the CTA — shows each path as a row (label + Current chip + accepted count + net dimension score color-coded + per-axis spec/sing/mem/vuln dim chips + the trust-verdict message inline). The same surgery items list is reused across branches — the diagnose snapshot is frozen at session open. Phase B (B1-B6) ALL SHIPPED in 5 builds (B3159-B3163). Phase C (tier gating + marketing — C1-C4) is what remains in the Song Surgery feature arc.
Phase C — Tier gating + marketing (3-4 builds)
- C1 (B3164) — Tier gates wired into Surgery shipped.
src/lib/surgery/tier-gate.tsexportscomputeSurgeryAccess(tier)→SurgeryAccesswith explicit fields per gate (fullAccess,studioAccess,maxItemsShown,showFinalVitals,allowBranching,showTimeline,upgradeCtaLabel). Free / Creator / Starter resolve to limited access (max 2 items, no Final Vitals, no branching, no timeline,'Upgrade to Professional'CTA label). Pro + Admin resolve to full access (unlimited items + Final Vitals + branching + timeline). Studio currently aliases to Pro — C2 splits per-arc Studio modes off this same gate.SURGERY_LOCK_FEATURESexports the marketing copy fragment naming the exact features unlocked by upgrading. Page wired: readsprofileRow.subscription_tierserver-side, slices the snapshot'ssurgeryItemstoaccess.maxItemsShown, and renders the lock banner above the Surgery Plan when!access.fullAccess— banner shows<Lock />icon + "Song Surgery is a Professional feature" header + count of hidden items + bulleted feature list + forge-styled "Upgrade to Professional" link to/pricing. 9 tests pin the policy (null/undefined defaults, free/creator/starter/pro/admin resolution, marketing feature list invariants, type narrowing on the resolved tier). All 58 surgery tests green. - C2 (B3165) — Per-arc Studio modes shipped.
src/lib/surgery/studio-modes.tsexportsSurgeryArc(9-arc enum: rnb / country / pop / rock / indie / folk / worship / latin / rap + 'unknown'),detectSurgeryArc(genre)(pure-sync; rap-first regex check then falls through todetectGenreFamilyfrom suno-genre-lock),getStudioModeProfile(arc)(returnsStudioModeProfilewithheaderLabel,byline,craftSignature,sacredAccident), and the convenience helperstudioModeForSong(genre). Each of the 9 arcs has a complete profile: Worship Surgery → "The divine subject must be NAMED, not implied — Berean Test" (SA#24) + APR/TFI/BVT signature; Country Surgery → "Authenticity is INHABITED, not INHERITED" (SA#20) + CID/LRR/MSC/VIG/SII signature; R&B Surgery → "Vulnerability requires receipts — the AID rule" (SA#23) + NCD/HVT signature; Pop Surgery → "Phonetic mass beats semantic precision" (SA#21); Rock Surgery → "Open vowels at the chorus peak" (SA#27); Indie Surgery → "The personal detail is the universal door" (SA#25); Folk Surgery → "Structure IS feeling" (SA#26); Latin Surgery → "A dialect is OWNED, not assembled" (SA#22); Rap Surgery → "A genre we cannot evaluate cannot be a genre we can serve" (SA#19). Page wired: whenaccess.studioAccess === trueAND the detected arc is not 'unknown', the page header label swaps from "Song Surgery — Diagnosis" to e.g. "Country Surgery — Diagnosis", with the per-arc byline + craft-signature paragraph + Sacred Accident citation rendered below the title.data-testid="surgery-mode-label"pins the header for e2e. The underlying diagnose engine + repair flow is unchanged — the per-arc audit primitives shipped through B2838-B3015 already feed intocomputeLineLevelSurgeryvia the existing eval pipeline. 18 tests pin the arc detection + profile lookups. Phase C status: C1 ✓ C2 ✓. Remaining: C3 (/song-surgery marketing landing) + C4 (pricing page Pro tier promotion). - C3 (B3166) — `/song-surgery` marketing landing shipped. New static page at
src/app/song-surgery/page.tsx.force-static+ Open Graph metadata + canonical pointing to${BRAND.siteUrl}/song-surgery. Eight sections: (1) Hero with<Scissors />icon + Song Surgery chip + the load-bearing tagline "Don't regenerate your song. Save it." (forge-gold "Save it.") + dual CTAs (Open Song Surgery → /pricing, Forge a song first → /forge); (2) The reframe (side-by-side cards: red-bordered "wrong question" = "Should I regenerate this song?" vs green-bordered "right question" = "Where does this song need surgery?"); (3) Restraint contract with<Shield />icon + green-bordered blockquote of the clean-song message; (4) Four actions grid (Accept / Edit / Reject / Skip) with descriptions; (5) Per-arc Studio modes ordered list with<Sparkles />icon — pulls all 9 arc profiles fromgetStudioModeProfile()rendering label + Sacred Accident citation + byline + craft signature; (6) Trust contract with yellow-bordered variance-message blockquote; (7) Branching panel with<GitBranch />icon; (8) Final CTA — "Save your song." H2 + Professional-tier chip + dual CTAs (Upgrade → /pricing, How scoring works → /scoring/standard). - C4 (B3166) — Pricing page Pro tier promotion shipped.
src/app/pricing/pricing-data.tsxPro tier features now lists★ Song Surgery — line-by-line lyric repair until your song is release-ready...as the SECOND★-prefixed item (Release Dossier remains first per B3152). Also added a Song Surgery row to the comparison table: Free / Creator = "Preview (2 items)", Pro = "Full + per-arc Studio modes" — pinning the exact tier-gate semantics shipped in C1. Song Surgery Phase C COMPLETE. C1 ✓ C2 ✓ C3 ✓ C4 ✓. Full Song Surgery feature arc (B3154-B3166) shipped in 13 builds: A1-A5 foundations (5 builds), B1-B6 workflow UX (5 builds counting B2 folded into B1), C1-C4 monetization + marketing (3 builds with C3+C4 paired).
The 8 WAR Room findings (compressed):
1. 80% of infrastructure already shipped via the Release Dossier work. Song Surgery is a UX-layer reframing, not a new system. 2. The reframe IS the value. Marketing positioning + workflow design carry the product; the engine is done. 3. Load-bearing risk: score-delta over-promise. Mitigation: directions, not numeric promises in the user flow. Verified rescores at end-of-batch, not per-edit. When verified < projected, surface honestly. 4. The "improves vs. changes" primitive is new and load-bearing. A1 in Phase A — without it, the system can't tell when a fix is a regression in disguise. 5. Naming: "Song Surgery" + "Diagnose" + "Repair" + "Vitals" + "Release Clearance." Cap the metaphor — no scalpels, no operating rooms. 6. Retention curve differs from generation products. Investment-driven, not novelty-driven. Subscription pricing rewards it; per-credit pricing punishes it. 7. Version timeline = single dashboard component, not new system. 2-3 builds, not infrastructure. 8. The Studio tier monetizes the existing per-arc work. 9 genre arcs already shipped via Excellence WAR Rooms become Studio-tier features.
The single most important design constraint (per WAR Room Panel E):
Song Surgery must be ABLE to refuse to suggest a change. WhenrunAllReportQaGatesreturnsallPassed: trueANDcomputeFinalRecommendationreturns'Ship'with high confidence, the Surgery page MUST show: "This song doesn't need surgery. Ship it." Not "here are some optional polish suggestions." Restraint is the trust signal. A surgery that always finds something to operate on is the gimmick that kills the brand.
Activation: when the operator types Continue Song Surgery (after the Release Dossier closes), the agent reads this section + finds the next unchecked A/B/C item + ships it. Same pattern as Continue Release Dossier and Continue Punch List.
Constraint-Aware Forge (CAF) — 21-build roadmap (HIGHEST PRIORITY)
Triggered by: 5-song stress-test WAR Room × 100 rounds + Sacred Accident #17 (Build 2763). The reviewer's 10 recommendations all collapse to one parent failure: the model optimizes for craft and forgets the brief. The fix is fidelity as a first-class score, orthogonal to quality. 21 builds across 3 phases.
Activation phrase: Continue Constraint-Aware Forge — reads SA#17 + this section + ships the next unshipped item.
Phase 1 — Foundation (week 1-2)
- CAF-1 — Brief extractor library.
src/lib/claude/brief-extractor.ts. Haiku-powered. Extracts{ premise, anchors[], structure[], styleConstraints[], forbiddenLanguage[], constraintMode, ambiguities[], contradictions[] }from raw prompt. Failure-safe; returns permissive default on parse failure. (B2764 — shipped.) - CAF-2 — Structure compliance audit.
src/lib/claude/audit-structure.ts. Pure regex. Comparesbrief.structureto actual section markers; returns{ requested[], actual[], hits, misses, partial, score }. Partial credit for adjacent-equivalent sections. (B2765 — shipped.) - CAF-3 — Thesis language detector.
src/lib/claude/audit-thesis.ts. Pure regex. Flags thesis-shaped lines (declarative analysis register, moral summary, therapy register). Returns{ flaggedLines[], count, score }. (B2766 — shipped.) - CAF-4 — Per-section specificity in prosody engine. Extend
src/lib/claude/prosody-engine.ts. Addsreport.sections[i].specificity: { concreteNouns, physicalActions, sensoryWords, score }per section. (B2767 — shipped.) - CAF-5 — Wire Phase 1 through the pipeline + dashboard. Shipped across 4 builds B2769-B2772. (5a/B2769) Orchestrator
src/lib/fidelity-audit.tscomposes brief + 3 audits + forbidden-language regex into onerunFidelityAudit()call (26 tests). (5b/B2770)songs.fidelity_auditJSONB column viaadd_fidelity_audit_to_songs.sqlmigration + partial index for future leaderboard composite ranking;updateSongacceptsfidelityAuditpayload;DashboardSongcarries it through. (5c/B2771)finalize-song.tsrunsextractBrief(rawPrompt)+runFidelityAudit()inline post-forge;SF_FIDELITY_AUDIT_DISABLED=1escape hatch; per-songforge.finalize.fidelity_scoredlog emit. (5d/B2772)FidelityPanel.tsxrenders the composite grade + brief summary + per-component breakdown inSongDetail. The full data flow exists end-to-end: raw prompt → extractBrief → audit → DB → dashboard.
Phase 2 — Prompt-level + grade infrastructure (week 3-6)
- CAF-6 — Forge prompt amendments. Shipped across 3 builds B2773-B2775. (6a/B2773)
renderBriefBlock(brief)helper insrc/lib/claude/render-brief-block.ts— pure prompt fragment with 8 canonical sections (mode framing, premise, required details, required structure, style constraints, forbidden language, ambiguity watch, contradictions). 16 tests. (6b/B2774)extractBriefmoved UPSTREAM to preforge-orchestration stage 0; runs in parallel with stages 1-7; net wall-clock ~0; Brief threads throughPreforgeOrchestrationResult.brief→post-forge-pipeline.ctx.preforge.brief→finalize-song.ctx.brief(reuses the cached Brief — no redundant Haiku call). New telemetry:preforge.brief_extracted. (6c/B2775) Forge prompt assembler (buildForgeUserPromptin user-prompt.ts) accepts optionalbriefin BuildForgeUserPromptOptions; renderBriefBlock fragment PREPENDS the prompt in all 3 branches (empty-prompt / starter-lyrics / freeform);forgeSongStream+runForgeCore+forge-stream-sessionall thread the Brief through. The model now writes against explicit anchor / structure / style instructions BEFORE writing line one. - CAF-7 — Anchor coverage audit. Shipped across 2 builds B2776-B2777. (B2776) Extracted CONCRETE_NOUNS / PHYSICAL_VERBS / SENSORY_WORDS vocabulary + tokenizer into
src/lib/claude/specificity-vocab.tsso CAF-4 and CAF-7 share the same line-specificity counter (no drift between two audits). (B2777)src/lib/claude/audit-anchor-coverage.tsships the audit: extract key tokens from each anchor (drop stopwords, keep numerics), find the best-matching sung line per anchor (token-match count, tie-break by specificity), score on the 4-band scale per the WAR Room verdict — 0 missing / 33 mentioned-bare / 66 partial / 100 delivered. Pure-sync (no Haiku call; the WAR Room option for Haiku second-opinion is deferred until evidence shows the heuristic disagrees with human judgment). 21 tests. Orchestrator now populatesanchorCoveragecomponent when brief has anchors; FidelityAudit.phase ratchets to 'phase-1.5'; auditVersion bumps 1.0.0 → 1.1.0. SongDetail FidelityPanel renders new per-anchor breakdown with evidence lines + the hits/partial/missing counts above the per-component score. - CAF-8 — Chorus Evolution Planner. Shipped across 3 builds B2780-B2782. (8a/B2780)
src/lib/claude/chorus-evolution-planner.ts— Haiku pre-write phase produces{ position1, position2, position3, arcType, arcDescription }for the 3 chorus positions (belief / realization / can't-deny). 5 canonical arcs + custom; failure-safe; 9 tests. (8b/B2781) Wired upstream into preforge stage 0.5 (chains on briefPromise; parallel with stages 1-7). NewrenderChorusEvolutionBlockhelper renders the prompt fragment which the forge prompt assembler prepends after the brief block. Threaded throughforge-stream → run-forge-core → forge-stream-session → buildForgeUserPrompt. New telemetry:preforge.chorus_evolution_planned/preforge.chorus_evolution_skipped. (8c/B2782)src/lib/claude/audit-chorus-evolution.ts— pure-sync audit detects whether choruses byte-shifted OR verses reference back to chorus content. 4-band scoring (0 static / 33 weak / 66 partial / 100 evolved). Wired intorunFidelityAuditsocomponentScores.chorusEvolutionpopulates (5% weight). FidelityAudit.phase ratchets to'phase-1.75'; auditVersion1.1.0 → 1.2.0. FidelityPanel renders planned-vs-shipped (planner's 3 positions + audit's state classification + chorus count + verbatim-repeat count). 10 tests. - CAF-9 — Earned Transcendence. Shipped across 2 builds B2783-B2784. (9a/B2783)
src/lib/claude/audit-transcendence.ts— pure-sync audit detects whether a concrete-noun image planted in V1 returns in the final chorus with transformed surrounding words. 4-band scoring (0 missing / 33 verbatim / 66 partial / 100 transformed) on Jaccard similarity of non-image content words. Charitable candidate selection: when multiple concrete-noun images recur, picks the one with the LARGEST surrounding shift. 13 tests. (9b/B2784) Wired intorunFidelityAuditsocomponentScores.transcendencepopulates (5% weight per SA#17).FidelityAudit.phaseratchets to'phase-2'when the audit produces a real classification (state !== 'na'). auditVersion bumps1.2.0 → 1.3.0. FidelityPanel renders the callback image + V1/final-chorus line pair + surrounding-word similarity meter. 4 orchestrator tests (na branch / transformed scoring / phase-2 ratchet / version bump). - CAF-10 — Conditional Sensory Rewrite. Shipped across 2 builds B2785-B2786. (10a/B2785)
src/lib/claude/sensory-rewrite.ts— Sonnet-driven targeted line-level rewrite library. Gate predicateshouldRunSensoryRewrite(brief, thesisAudit)fires whenbrief.styleConstraintsincludes one of['sensory-only', 'show-dont-tell', 'no-thesis']ANDthesisAudit.flaggedLinesis non-empty.runSensoryRewrite()calls Sonnet with flagged lines + ±2 lines of context + brief anchors, parses the JSON response, verifies each rewrite against the audit (rejects model-side line-number drift), stitches the rewrites in. Failure-safe: every error path returnsfired: false+ original lyrics. 22 tests. (10b/B2786) Wired intofinalize-song.tsBEFORE the contribution ledger + prosody + fidelity audit blocks, so all three downstream artifacts see the FINAL (possibly rewritten) lyrics. Pre-rewrite thesis audit (~3ms regex) drives the gate; full fidelity audit runs once on the rewritten lyrics. SensoryRewriteResult metadata attached to the persisted FidelityAudit JSONB so the dashboard reads from one column. FidelityPanel renders before/after pairs (strikethrough original + green rewrite) with thesis-category labels per pair. Telemetry:forge.finalize.sensory_rewrite_applied/_failed/_threwlog events. Env toggle:SF_SENSORY_REWRITE_DISABLED=1for the same shape asSF_FIDELITY_AUDIT_DISABLED. - CAF-11 — Final Adherence Audit. Shipped across 2 builds B2787-B2788. The other CAF-11 sub-items (11a/11b/11c — complexity + composite + grade helpers) shipped at B2768. (11/B2787)
src/lib/claude/audit-premise.ts— async Haiku premise-match judgment. Returns{ score, verdict, reasoning, premiseEcho, evidence, state, durMs }. 4 verdict bands (wrong-song / drifts / mostly-served / faithfully-served). NA branch skips Haiku; failure-safe on every error path. 18 tests. (B2788)RunFidelityAuditInputgains optionalpremiseAuditResult(matches CAF-8c chorusEvolution pattern, keeps orchestrator pure-sync).FidelityAuditgainspremiseAuditfield.componentScores.premisepopulates fromresult.scorewhenstate='judged'(otherwise null — 'na'/'error' both redistribute weight). Phase ratchets to'phase-2'when EITHER premise OR transcendence is judged. auditVersion1.3.0 → 1.4.0. finalize-song.ts callsauditPremiseupstream fromrunFidelityAudit, threadspremiseAuditResultthrough.SF_PREMISE_AUDIT_DISABLED=1toggle. FidelityPanel renders premise panel FIRST (load-bearing 30% slot) with verdict chip + premise echo vs brief-premise diff + reasoning + evidence lines + tier legend. 6 orchestrator tests (null when omitted, null on 'na' state, null on 'error' state, populates on 'judged', phase ratchets to phase-2, version 1.4.0). CAF Phase 2 is COMPLETE — all 7 fidelity components have heuristic or judged scoring. - CAF-11a — Brief complexity computation.
computeBriefComplexity(brief): 0-10based on anchor count + structural reqs + style constraint count + forbidden-language list. (B2768 — shipped.) - CAF-11b — Fidelity composite.
computeFidelityComposite(audits): { score, perComponent, grade, complexityBucket }. 30/25/15/15/5/5/5 weighting (premise + anchors + structure + style + forbidden + chorus + transcendence). (B2768 — shipped.) - CAF-11c — Grade helpers.
fidelityScoreToGrade(score)(A+/A/B+/B/C+/C/D/F bands) +fidelityScoreToBucket(score, complexity)(hide/secondary/primary). Mirrorssrc/lib/scoring.tsshape. (B2768 — shipped.) - CAF-11d — Dashboard SongDetail two-grade surface. Shipped B2789. New
src/components/SongHeroFidelityPill.tsxrenders fidelity score + grade chip alongside the existingSongHeroScorePillin the SongDetail hero band.shouldRenderFidelityPill(composite)decides whether to mount:'primary'(heavy brief) and'secondary'(standard brief) always render;'hide'(light brief) suppresses unless score drops below 80 (operator-spec'd rescue path) — pill then renders withlow fidelitywarning + orange border. In'primary'bucket, fidelity pill renders FIRST (headline position) AND with the same large 2xl number as the quality pill; in'secondary'it renders SECOND at slightly smaller weight (xl). Title attribute carries the grade label + complexity score + bucket reason for operator-facing hover-tip. Per-component breakdown disclosure panel was already shipped in CAF-5d/B2772 (FidelityPanel). 5 tests pin the gate helper across all 4 bucket × score combinations.
Phase 3 — Ratchets + public docs + Rap Mode (month 2+)
- CAF-11e — Leaderboard composite migration. Shipped B2791.
src/lib/leaderboard/top-weekly.tsnow ranks by0.6 × forgeScore + 0.4 × fidelityScorewhen the song's persisted fidelity_audit carries a composite.score number. Pre-CAF songs (forged before B2771) fall back toquality aloneso they're never silently demoted by the migration. NewcomputeRankScore(q, f|null)helper exported for unit tests + dashboard surfacing.TopWeeklySonginterface gainsfidelityScore+rankScorefields so the leaderboard UI can show the "why this position" explainer. Query overfetches when fidelity ranking is active (composite re-ranking can promote songs from below the SQL ORDER BY cutoff).SF_LEADERBOARD_FIDELITY_DISABLED=1env flag forces the legacy quality-only sort as the emergency rollback. 6 tests pin the formula (null-fallback identity, parity at Q=F, high-Q+low-F drop, low-Q+high-F rise, integer rounding, pre-CAF no-penalty invariant). - CAF-11f — `/scoring/standard/fidelity` public docs. Shipped B2792. New static page at
/scoring/standard/fidelitydocuments v0.1.0 of the standard. CC BY 4.0 licensed. ScholarlyArticle + TechArticle JSON-LD schema for Google Scholar / Semantic Scholar indexing. 8 sections: (1) Fidelity is the orthogonal question (vs quality); (2) The seven components — table + per-component drill-downs naming the Haiku/heuristic split; (3) Composite formula with the constraint-mode multipliers + null-component redistribution math in code block; (4) Grade calibration A+ through F with color-coded chips matching the dashboard; (5) Brief complexity gating (hide/secondary/primary UX buckets); (6) Version history (v0.1.0 → RFC); (7) How to cite (plain-text + BibTeX); (8) Related standards cross-link to Lyric Scoring Standard. Receipts-row addition pending CAF-14. - CAF-11g — RFC-0010 opens. Shipped B2793. RFC-0010 "Fidelity Score v0.1.0 — calibration + composite formula" added to
src/lib/rfcs.ts. Opens 7-day public comment window (opened 2026-05-19, commentDeadline 2026-05-26). Body pins: component weights (30/25/15/15/5/5/5), composite formula with null-component redistribution math, constraint-mode multipliers (strict 0.95 / standard 1.0 / loose 1.15), 8-tier grade calibration, brief complexity formula, 3 UX prominence buckets + rescue path. Five open questions for the comment window covering weight allocation / mode caps / complexity threshold / chorus+transcendence weights / 'na' verdict semantics. Auto-rendered at/rfc/0010-fidelity-score-v0-1-0-calibration-and-composite-formulavia the existing[slug]route. Status'in-comment'(gold chip). The fidelity standard page (/scoring/standard/fidelity) links to the RFC from the Version history section. - CAF-12 — Fidelity score CI ratchet. Shipped B2790.
src/lib/golden-evals/fidelity-fixtures.tsdefines 4 hand-curated{ brief, lyrics }fixtures spanning the fidelity distribution: (1) faithful-anchor-delivery (high score, A/B band), (2) structure-miss-truncated (4 sections shipped vs 6 requested → mid band), (3) sensory-only-thesis-leak (style constraint violated by thesis-shaped lines → C/D band), (4) minimal-brief-decent-song (light brief baseline → tests redistribution math). Each fixture pins an expected composite band (≥10 wide to absorb minor weight-redistribution drift) + an expected grade prefix set.src/lib/golden-evals/fidelity-ratchet.test.tsrunsrunFidelityAuditon each fixture + asserts the composite stays inside the band and the grade prefix matches. Plus 4 orchestrator-behavior assertions (phase ratcheting, structure miss detected, thesis flags counted, light complexity verified). Premise audit deliberately omitted from the ratchet — Haiku-driven, non-deterministic, separate ratchet on schedule. 13 tests total. Catches regressions in the audit stack OR composite math that wouldn't fire any unit test. - CAF-13 — Rap Mode (first genre-craft module). Shipped B2794. New
MODE_RAPentry insrc/lib/claude/genre-modes.ts(Genre Modes infrastructure B2170). Trigger keywords: rap / hip hop / hip-hop / trap / boom bap / drill / mc. Panel emphasis: Kendrick Lamar / Earl Sweatshirt / MF DOOM / Nas / Rapsody / Pat Pattison (channeled). Rubric weight overrides: M3 (Rhyme Intelligence) × 1.5, M5 (Specificity) × 1.3, M11 (Memorability) × 1.2, M1 (Prosody) × 1.2 — making rhyme the load-bearing metric. Mode floor: M3 ≥ 78. Success criteria + chorus discipline name the form's discipline: "the verse is the song; the hook is the runway"; "dead bars are the failure mode; every bar must earn its position." Companionsrc/lib/claude/audit-rap-craft.tsships three pure-sync heuristics:auditBarLength(per-line syllable count + mean/stdDev/range/outlier-line flagging),auditEndRhymes(couplet AABB + alternating ABAB rate detection via last-4-chars suffix matching),auditInternalRhymes(repeated-suffix bigram counter within each line).auditRapCraftorchestrator runs all three in one call. 25 tests pin the heuristics + 5 detector tests confirm MODE_RAP fires on the right keywords. The audits aren't yet wired into the fidelity composite — that wiring lands once empirical data shows the heuristics correlate with human judgment. - CAF-14 — `/scoring/standard/fidelity` final public artifact. Partial — shipped B2793 (changelog stub); v1.0.0 release artifact lands after RFC-0010 closes 2026-05-26. New static page at
/scoring/standard/fidelity/changelogwith append-only version history (initial entry: v0.1.0 in-comment, dated 2026-05-19, with 6 detail bullets). Status chip rendering (draft / in-comment / released) mirrors the RFC chip color scheme. Implementation-version disclaimer at the bottom namesFIDELITY_AUDIT_VERSIONas the codepath that tracks separately. Cite-this-page shipped at B2882 (/scoring/standard/fidelity/cite— BibTeX, APA 7, MLA 9, Chicago, RIS, plain text). B3195 (CAF-14 advance — npm package sync):@songforgeai/fidelity-standardbumped 0.1.0 → 0.2.0 to match the live standard. Pre-3195 the package was publishing v0.1.0 (7 components) while the live standard page + audit engine were at v0.2.0 (8 components, registerAdherence added at B2889/CAF-11h) — external citers were getting a stale shape. Sync changes:public/fidelity-standard.jsonregenerated at v0.2.0 with 8th component (registerAdherence, 6% weight, heuristic method) + newchangelogfield with v0.2.0 + v0.1.0 entries.FidelityComponentIdTS union extended with'registerAdherence'. NewFidelityChangelogEntryinterface + optionalchangelog?field onFidelityStandard. Package CHANGELOG.md updated; dist/standard.json regenerated viatsc && cpbuild script. 23 new tests pin: version 0.2.0, 8 components in canonical order, registerAdherence weight 0.06 + heuristic method, weight sum 1.06 (normalized at compute time), summary text mentions "eight-component", changelog newest-first ordering, RFC-0010 still in-comment + 2026-05-26 deadline, FidelityComponentId union accepts registerAdherence, all 8 ids satisfy the union, computeFidelityGrade/Bucket/BriefComplexity/applyConstraintMultiplier still work correctly. 23/23 tests green; 75/75 main repo fidelity tests still green; tsc clean. The v1.0.0 release artifact (CAF-14 final) still lands after RFC-0010 closes 2026-05-26 — operator runs `cd packages/fidelity-standard && npm publish` to push v0.2.0 to npm when ready.
Operator-feedback nightly run B3391-B3397 (shipped 2026-05-28)
Operator reported on 2026-05-28: (1) "clavicle is showing up in lots of lyrics" — system-internal feedback loop where B3367's banned-term ALTERNATIVES list ("Try ribs, clavicle, shins, knees, jawline") was being treated as suggestions, and the alternatives became the new cliché. (2) "Songs still seem to go to folk as the default when no genre specified in the prompt" — a different class-of-bug from B3276-B3284's genre-routing collapse (which fixed cases where genre WAS specified but lost); this is the no-genre default case left silent.
Shipped a 7-build response in autonomous mode:
- B3391 — Quick wins. Strip clavicle/shins/jawline from
BANNED_TERMSalternatives; add them as their own banned entries; kill the 30% randomMath.random() < 0.3folk-attractor roll insingle-phase-prompt.ts. - B3392 — Catalog motif scanner module. Pure-function scanner with stem-folding (silent-e restoration), partitions output into
overusedvsleakingBannedbuckets, 15 tests. - B3393 + B3394 — Closed loop: cron + forge-prompt reader.
/api/cron/scan-catalog-motifsruns nightly at 03:30 UTC (added tovercel.json); writes scan totrend_auditswith marker[catalog-scanner] v1.dynamic-forbidden-motifs.tsreads latest scan, injects## DYNAMIC FORBIDDEN MOTIFSblock at tail of forge system prompt. 17 tests on the reader. - B3395 — Sentiment-routed default-genre picker.
suggestDefaultGenre(prompt)Haiku call (~$0.001-0.002/no-genre forge). Calibrated to AVOID folk by default; folk fires only on explicit reflection + grief + slow + acoustic signals. Wired intopreforge-orchestration.ts. Escape hatch viaSF_DEFAULT_GENRE_PICKER_DISABLED=1. 13 tests. - B3396 — DECADE SHIFT gate cleanup. Regex scoped from topic words (
grief|loss|memory|story) to genre words only (folk|country|indie|singer-songwriter|ballad|americana). Added structured logforge.default_genre.surfacedinpost-forge-pipeline.ts. - B3397 — R5-A1 + R5-A3 cleanup. Refreshed
src/lib/.git-snapshot.json(was 67 commits behind). Plumbedbrief.toneRegisterthrough gauntlet system prompt viasongs.fidelity_audit.brief.toneRegisterlookup.
Verification windows:
- The catalog-motif cron fires for the first time tonight (03:30 UTC). Operator can verify via
/api/cron/scan-catalog-motifsGET with Bearer token or via Vercel cron logs. - The dynamic-forbidden block appears on the SECOND forge after the first scan completes — the FIRST scan writes the row; subsequent forges read it.
- The default-genre picker fires on the next no-genre forge; grep logs for
forge.default_genre_picked.
Deep Audit 2026-05-27 round 5 — autonomous queue (TOP PRIORITY)
Triggered by: the post-B3383 Deep Audit, captured at the end of today's 12-build SA#32 register-awareness arc (B3372 → B3383). Final grade B+ (7.4/10 weighted). Strongest: trust/credibility (9), process maturity (10), technical execution (9), strategic differentiation (9). Weakest: information architecture (5 — new register-awareness page orphaned), conversion architecture (5 — CTA dilution), onboarding (5 — rich-prompt pattern undocumented), commercial readiness (6 — distribution N=0). Engineering velocity continues to outrun marketing maturity by 2 quartiles; outrunning distribution by 3.
Live verification at audit time: /api/version returns {build:3383, sha:59888b3fd, ref:main} — production matches HEAD. check:all 58/58 green. npx vitest run 1 failed + 8587 passed (the failure is the load-bearing P0 below).
Operator-side complements live in docs/OPERATOR-PUNCH-LIST.md under the matching "Audit-Round-5" section. Engineering items below ship autonomously.
Round-5 Tier 1 (ship THIS WEEK — small, each ≤2h)
- R5-A1 — Refresh `src/lib/.git-snapshot.json`. [AUTONOMOUS · P0 · ~5 min] Shipped at B3397 (2026-05-28).
npm run snapshot:refreshwrote 300 commits + 300 file-touch entries; unblocks 4 public surfaces (/changelog,/engineering,/roadmap,/now) that had been rendering pre-B3331-era content.
- R5-A2 — Add `register-awareness` to `StandardsClusterFooter.tsx`. [AUTONOMOUS · P0 · ~10 min] Shipped at B3398 (2026-05-28). Added
'register-awareness'to theSurfaceSlugunion + a SURFACES entry (Activity icon, blurb naming SA#32 + 10 registers + Gravity / Burden modifiers). ALSO mounted<StandardsClusterFooter currentSurface="register-awareness" />on the page itself so it both appears in siblings' footers AND back-links to them.
- R5-A3 — Plumb `brief.toneRegister` through the gauntlet system prompt. [AUTONOMOUS · P0 · ~30 min surgery] Shipped at B3397 (2026-05-28). Reads from
songs.fidelity_audit.brief.toneRegisterJSON on the song row at gauntlet start; passes as the 11th arg togauntletStream. Defensive shape walk means missing column / undefined register falls back to pre-B3380 register-blind prompt. Closes the dormant B3380 infrastructure — every register-declared song's gauntlet pass is now register-aware.
- R5-A4 — Add SA#32 chip above-fold on homepage. [AUTONOMOUS · P0 · ~30 min] Shipped at B3398 (2026-05-28). Added "Register-aware · 10 modes" link to the hero proof strip (mono text, no caps, no color per the B3289 "trust whispers" discipline). Linked to
/scoring/standard/register-awareness. Solves orphan-page discoverability AND the brand-promise gap in one chip.
Round-5 Tier 2 (ship THIS MONTH — bigger, each <1 day)
- R5-A5 — Align git-snapshot-freshness thresholds. [AUTONOMOUS · P2 · ~1h] Shipped at B3400 (2026-05-28). Documented the divergence explicitly in both gates' headers (not aligned thresholds, because they intentionally measure DIFFERENT artifacts: git-snapshot.json by commit-count vs CLAUDE.md AUTO-STATE block by calendar days). Each header now explains its own metric AND points to the sister gate's metric so future operators understand why one can be red while the other is green. They are complementary, not contradictory; running BOTH cures (
npm run snapshot:refresh+npm run docs:refresh) is the pre-push ritual.
- R5-A6 — Tighten `check:corpus-claims` keyword detector. [AUTONOMOUS · P2 · ~30 min] Shipped at B3399 (2026-05-28). Split
CLAIM_PHRASESintoSTRONG_CLAIM_PHRASES(corpus-specific phrasings that fire on their own — hand-scored exemplars / corpus entries / reference corpus) andAMBIGUOUS_CLAIM_PHRASES(worked examples / reference entries — only fire when a corpus-context anchor appears within ~150 chars). Anchor list: corpus / exemplar / calibration / rubric / golden evals / ground-truth / anchor corpus / hit calibration / sleeper ledger. Existing scan still green; B3384-style false-positives now structurally impossible.
- R5-A7 — Disambiguate `brief.register` → `brief.contentRating`. [AUTONOMOUS · P1 · ~30 min cross-codebase] Shipped at B3399 (2026-05-28). UI-side relabel: every "Register adherence" surface now reads "Content rating adherence" (
/scoring/standard/fidelitytable + heading + prose,fidelity-grade.tscomponent label,fidelity-audit-shape.jsonsnapshot). Added a disambiguation note in the prose: "Not to be confused with the SA#32 tone register (joy / swagger / rage / playfulness / etc) — content rating is the MPAA-style profanity perimeter; tone register is the emotional posture." The internal field key (brief.register,registerAdherence) keeps its original name for persisted-fidelity_audit-JSON schema compatibility; only the user-facing labels change so production rows still parse.
- R5-A8 — Add `toneRegister` UI surface to `/crucible`. [AUTONOMOUS · P2 · ~2h] Shipped at B3403 (2026-05-28). Added a 4-chip radio row (Joy / Swagger / Rage / Playfulness) styled with per-register colors (violet / amber / red / pink) directly under the voice-mode toggle. Selection is persisted in
localStorage(sf_crucible_register) so a returning user keeps their pick. The chip threads through to the POST body astoneRegister, activating the B3378 per-voice register appendix. Linked to/scoring/standard/register-awarenessso users can read the SA#32 rationale inline. The 6 other canonical registers (grief/melancholy/awe/tenderness/lust/defiance) are intentionally omitted from the UI — they're either the home register (no appendix change) OR awaiting operator-paced calibration per SA#11.
- R5-A9 — Cross-link audit ratchet. [AUTONOMOUS · P2 · ~2h] Shipped at B3401 (2026-05-28).
scripts/check-cross-link-audit.tswalks every page.tsx added viagit log --since="30 days ago" --diff-filter=A, computes each route's self-cluster (parent path prefix), and asserts ≥1 inbound link from a route outside that prefix. Component files (src/components, etc.) count as cross-cluster automatically. Warn-only at launch (CROSS_LINK_AUDIT_BLOCKING=0 default) so the existing baseline (4 untriaged surfaces: /admin/corpus-novelty, /standard-vs/[slug], /design-system, /security, /status) doesn't break CI; flip to blocking once those are cross-linked. Wired intocheck:allas the 59th gate.
- R5-A10 — Synthetic post-deploy probe pack for cluster-footer pages. [AUTONOMOUS · P2 · ~2h] Shipped at B3402 (2026-05-28).
scripts/probe-standards-cluster.tsregex-parses the SURFACES array fromStandardsClusterFooter.tsxand GETs each${PROBE_BASE_URL}${href}, asserting 2xx. Defaults to https://songforgeai.com; override viaPROBE_BASE_URLenv var. Six-test parser-side contract atsrc/lib/standards-cluster-parser.test.tspins the registry shape so a future refactor can't silently break the parser. Available asnpm run probe:standards-cluster. Run post-deploy or as part of trust-decay-audit checklist.
- R5-A11 — Add SA#32 chip + per-register example songs above the homepage fold. [AUTONOMOUS · P1 · ~3h] Surface 1 song per shipped register (joy/swagger/rage/playfulness) via inline player cards. The
/examplespage exists; reuse the card pattern. Converts SA#32 work from "claimed in a sub-page" to "demonstrated immediately."
- R5-A12 — Forge prompt template chips on `/forge`. [AUTONOMOUS · P1 · ~3h] Add buttons that prefill the prompt textarea with rich-prompt templates: "Joy · single song", "Country · alt-country · breakup", "Worship · contemporary · doxology". Instant on-ramps that activate the genre + register pipelines without requiring the user to know the rich-prompt pattern.
- R5-A13 — In-product first-run tutorial for rich-prompt pattern. [AUTONOMOUS · P2 · ~4h] Shipped at B3406 (2026-05-28).
src/components/ForgeRichPromptHint.tsx— dismissable side-by-side "Thin" vs "Rich" prompt cards mounted above the prompt textarea on/forge. "Try this rich prompt" button seeds the textarea viaonSeedcallback (the seed is the joy-001 prompt from the B3405 eval fixtures). localStorage dismissal (sf_forge_rich_prompt_hint_dismissed) — separate from B2881's Crucible-redirect hint key so the two hints can coexist. Links inline to/scoring/standard/register-awareness. Demonstrates the pattern instead of explaining it abstractly.
Round-5 Tier 3 (next quarter — multi-build)
- R5-A14 — N=large register-fidelity eval framework. [AUTONOMOUS · P1 · 3-5 builds → shipped in 1] Framework shipped at B3405 (2026-05-28).
data/eval-fixtures/register-fidelity-prompts.json(50 prompts × 5 registers).scripts/eval-register-fidelity.tsrunner — smoke mode default (5 forges),--full --trials 3for the audit-spec 150-forge run.src/lib/eval/register-fidelity-report.tspure aggregator with 18 tests (extraction accuracy / AVD pass rate / declaration rate / composite mean+median per register).scripts/report-register-fidelity.tsCLI: JSONL → markdown. Full methodology + interpretation guide indocs/REGISTER-FIDELITY-EVAL.md. The actual 150-forge run is an operator budget decision (~$15-30 + 3-7 hours) — harness ships ready; running is a separate operator action.
- R5-A15 — 50-voice room rebalance. [AUTONOMOUS · P1 · 4-6 builds] Add Lizzo / Bruno Mars / Kesha / Max Martin archetype voices to balance the Cohen / Mitchell / Bridgers-heavy panel. The B3372 WAR ROOM §3.1 named this as the deepest residual bias. Even with B3373-B3383 register-aware infrastructure, the voice panel itself remains skewed toward the introspective register and drags joy/swagger drafts toward home register. ~30% of voices need rebalancing.
- R5-A16 — Per-tier feature delta highlighting SA#32 coverage. [AUTONOMOUS · P2 · 2 builds] Pricing page today: 3 tiers + credit packs. Add a "Register coverage" row in the comparison table: Free = grief/melancholy default; Creator = + joy / swagger; Pro = full 10-register support + register-aware refinement. Turns SA#32 into commercial value differentiation.
- R5-A17 — Per-register × per-genre intersection landing pages. [AUTONOMOUS · P3 · 3-5 builds] Long-tail SEO capture:
/lyrics/joy/pop,/lyrics/swagger/rap,/lyrics/rage/rock. ~30 pages × per-genre × top-3-registers. Each links to the register-awareness doc + the genre-specific craft plugin.
Round-5 Tier 4 (category-defining engineering bets)
- R5-A18 — Open-source the SA#32 corpus + heuristic framework. [AUTONOMOUS · P3 · 2-3 builds] Publish
@songforgeai/tone-registeras an npm package: thedetectToneRegisterHeuristic+registerAppendix+registerAntiInflationBlock+gauntletRegisterDirectiveprimitives. Standalone TypeScript, model-agnostic, no SongForgeAI-specific assumptions. Same playbook as@songforgeai/agent-room(PUNCH-LIST-V2 #1-3). Academic / OSS adoption surface for SA#32 methodology.
- R5-A19 — `/admin/register-stats` analytics dashboard. [AUTONOMOUS · P3 · 1-2 builds] Shipped at B3404 (2026-05-28).
/api/admin/register-statsaggregatessongs.fidelity_audit->brief->toneRegisterover a configurable window (24h/7d/30d/90d), groups by register (with null bucket for undeclared) + computes count / percent / mean / median / min / max forge_score per register. Page at/admin/register-statswith auth-tier gate, window selector, sortable table, per-register tint colors mirroring B3403. Linked from/adminhome viaADMIN_ROUTESentry. AVD verdict not yet persisted — currently emitted as logs only; route + page note the gap. Closes the empirical-measurement side from N=1 to N=catalog.
Cross-AI Feedback WAR ROOM (2026-05-28 — B3408)
Triggered by operator's 9-genre test sweep + cross-AI feedback synthesis. Full WAR ROOM doc: docs/WAR-ROOM-CROSS-AI-FEEDBACK-2026-05-28.md. WAR ROOM declined to ship 2 of the feedback's recommendations and shipped 1 small change inline; the remainder are queued below.
- R5-A20 — Anti-thesis-line banned phrases. [AUTONOMOUS · P2 · ~10 min] Shipped at B3408 (2026-05-28). 7 thesis-line phrases added to
src/lib/banned-terms.tsascategory: 'house_style', tier: 2— "the sound of," "this is the sound," "this is where," "the shape of," "the language of," "the architecture of," "the geography of." Catches the "commentary-on-the-song" construct the cross-AI feedback called out explicitly. Post-gen scrub fires automatically.
- R5-A21 — Mess-injection forge amendment. [AUTONOMOUS · P2 · ~2 builds] Add a forge prompt directive that requires one OFF-ANGLE / CONTRADICTORY / IMPERFECT line per song. The feedback's "Let one be ugly. A cracked voice, an embarrassing memory, a line that's not clever but true." The amendment lives in single-phase-prompt.ts as an optional features block under a feature flag for A/B measurement. Risk: medium — overshoots into bathos if not bounded. Implementation includes a Haiku judge post-forge to verify the off-angle line is present without being maudlin.
- R5-A22 — Chorus-mass test primitive. [AUTONOMOUS · P2 · ~1 build] SHIPPED B3447.
src/lib/audit-primitives/chorus-mass.ts— purecomputeChorusMassReport(lyric)parses every[Chorus]/[Chorus N]/[Final Chorus]block (excludes[Pre-Chorus]), counts words per line, and returns verdictanthemic(≥1 chorus has a line ≤8 words — a short repeatable hook) /dense(every chorus line is long — flagged, the feedback's "write one that's just 4-8 words repeated") /no-chorus, plus a dossier-readylabel("Chorus mass: anthemic" / "Chorus mass: dense"). Flag-only, never pass/fail (dense is a valid craft choice). 11 tests. Dossier/dashboard surfacing deferred as the consumer step (mirrors the AVD B3376 precedent — primitive ships first;labelis wired ready).
- R5-A23 — Named-event bridge audit. [AUTONOMOUS · P2 · ~1 build] SHIPPED B3447.
src/lib/audit-primitives/bridge-anchor.ts— pureauditBridgeAnchor(lyric)reusesextractBridge()(B3128) and scans each[Bridge]for an anchor: a digit (room 314), a calendar term (month/weekday), a spelled number (forty-three, hyphen-split), or a mid-line proper noun (excludes line-initial caps + "I"). Returnsanchored/unanchored/no-bridge+anchorTypes+matchedTerms+ dossier-readylabel. Flag-only (the spec said "fail"; shipped as a flag to match the sibling primitives' non-gating discipline). Complementsaudit-bridge-completeness(development) with the orthogonal anchoring signal. 8 tests — incl. a discrimination guard that surfaced a real finding: the country / pop / indie OSNG bridges are genuinely unanchored (abstract — "He taught me everything / So I could choose"), exactly the "no named wound" failure this primitive names.
Genre Catalog Clustering WAR ROOM (2026-05-28 — B3409)
Triggered by operator screenshots showing the /genres/indie/prompts page rendering 83 indie prompts, essentially every one medical/health themed. Full WAR ROOM doc: docs/WAR-ROOM-GENRE-CATALOG-CLUSTERING-2026-05-28.md. Smoking gun: indie's mustInclude array in src/lib/genre-prompt-catalog/arc-markers.ts literally contained 'medical term' as a vocabulary anchor — Sonnet dutifully used it in ~84% of generated prompts.
- R5-A24a — Remove "medical term" from indie arc-markers. [AUTONOMOUS · P0 · 5 min] Shipped at B3409 (2026-05-28). The literal string
'medical term'removed fromsrc/lib/genre-prompt-catalog/arc-markers.ts's INDIE_PACKAGE.mustInclude. Inline comment explains the WAR ROOM finding. Partial fix — the full architectural fix is R5-A24.
- R5-A24b — Indie catalog hand-rebuild (proof). [AUTONOMOUS · P0 · 1 build] Shipped at B3409.
src/data/genre-prompt-catalog/indie.jsonrewritten with 25 hand-curated VARIED prompts demonstrating the new variation discipline. Spans 9 distinct settings (laundromat, drive-through, farmers' market, library carrel, wedding, apartment, pottery class, strike line, antiques mall, garage workshop, bus, kitchen tutoring, reunion, garage echo, library volunteer, voicemail, birthday, lake run, DJ booth, grocery, lifeguard, storytime, loading dock, garden hose) + 9 distinct moods. Cross-cluster on time (dawn/morning/midday/afternoon/evening/night) + persona (1st-person / observer / addressed-other / collective). The 83-prompt count drops to 25; richness increases.
- R5-A24 — Split `mustInclude` into `laneMarkers` + `sceneRotationPool`. [AUTONOMOUS · P1 · ~1 build] Shipped at B3410 (2026-05-28).
ArcCraftPackagenow requires bothlaneMarkers(genre lane lock — production / artist / substyle / craft) ANDsceneRotationPool(≥30 diverse scenes per arc) as first-class fields. All 9 arcs populated.checkPromptCoherence()switched to preferlaneMarkers(mustIncluderetained as legacy fallback). Catalog ratchet rerun: all 9 arcs above floor (864/865 prompts pass — only 1 country prompt dropped). 5 new tests cover the split contract (lane non-empty, scene ≥30, lane/scene no-overlap, lane/avoid no-overlap, B3410 architectural-split landed).
- R5-A25 — Generator rewrite for sceneRotation discipline. [AUTONOMOUS · P1 · ~1 build] Shipped at B3410 (folded into R5-A24).
scripts/generate-genre-prompt-catalog.tsnow renders bothlaneMarkers(≥1 per prompt) andsceneRotationPool(1 per prompt with rotation discipline + ≤8% catalog domination cap) as separate concerns. Required-fields contract in the Sonnet prompt reads "MUST include ≥1 LANE MARKER + MUST anchor on ≥1 SCENE ROTATION POOL entry" rather than the conflated B3095 phrasing. Future regens (R5-A26-A30) will use the corrected prompt.
- R5-A26 — Regenerate country.json with new generator. [AUTONOMOUS · OPERATOR-BUDGET · ~$0.50 + 10 min Sonnet] Shipped at B3411 (2026-05-28). 62 fresh prompts at 100% coherence. Dominant-axis (honky-tonk/pickup/dirt-road) clustering rate dropped 94% → 55% (-39pp). New titles span domestic / transit / work / public / nature / relationship / memory / artist scenes evenly — Granddad's Last Sunday, Highway 12 at Midnight, Mama's Recipe Box, Carolina Moon Radio, Tractor Pull Saturday, Dolly on the Dashboard, etc. Honky-tonk cluster broken.
- R5-A27 — Regenerate pop.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 57 fresh prompts at 100% coherence. Clustering rate 31% → 19% (-12pp). New scenes span domestic/transit/work — Walk-In Closet Runway, Airport Security Liminal, Closing Shift Countdown, Hostess Stand Power.
- R5-A28 — Regenerate rap.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 79 fresh prompts at 100% coherence. Clustering rate 67% → 52% (-15pp). 2AM Kitchen Confessions, Mom's Couch Chronicles, Barbershop Politics, Subway Platform Soliloquy, Dispatch Desk Drama — domestic/transit/work scenes replace corner-store/cypher cluster.
- R5-A29 — Regenerate rnb.json. [AUTONOMOUS · OPERATOR-BUDGET] Shipped at B3411. 73 fresh prompts at 100% coherence. Artist namedrops (D'Angelo/Sade/Frank Ocean/Aaliyah) GONE from titles; scenes span Kitchen Sink Confessions, Bathtub Sanctuary, Balcony Summer Storm, Hotel Lobby 1AM, Tuesday Jazz Club Healing, After-Hours Diner Dreams. Artist + production tropes remain in prompt bodies as lane locks, scenes diversified.
- R5-A30 — Regenerate folk + rock + latin + worship JSONs. [AUTONOMOUS · OPERATOR-BUDGET · ~$2 + 40 min Sonnet] All four shipped at B3411 (single batch run, ~21 min wall time, ~$2 Sonnet total). Folk: 63 prompts, 86% → 84% (Stanza Study / Mechanics / Workshop instructional vocabulary GONE; replaced by Greyhound to Memphis, Lumber Mill Lament, Cannery Line Chronicles, Dishroom Midnight). Rock: 76 prompts, 36% → 32%, Basement Archaeology, Power Chord Democracy, Warehouse Cathedral, Mill Town Saturday joining the highway/Rust Belt anchors. Latin: 90 prompts, 21% → 7% (Abuela's Kitchen Memory, Papi's Barbershop Chronicles, Hermana's Quinceañera Dance — family scenes added to place-name anchors). Worship: 89 prompts, 89% → 80% (Sick Room Vigil, Prison Ministry Hope, Hospital Chaplain Visit, Kitchen Morning Prayer added private/domestic/public-service scenes).
- R5-A31 — SA#33 ratification. [AUTONOMOUS · P3 · 1 build] Shipped at B3415. SA#33 added to
src/lib/brand.ts(BRAND.sacredAccident33) +src/lib/sacred-accidents/data.ts+docs/SACRED-ACCIDENTS.md. Ratified on operator's "Do your recommendation" authorization once B3411 empirical evidence (all 9 catalogs at 100% coherence post-architectural-split) had landed. Companion to SA#16 (output-side motif saturation) and SA#29 (genre signal as STATE).
Concrete-to-Abstract Drift WAR ROOM (2026-05-28 — B3412)
Triggered by a multi-AI craft critique of three real songs ("Speed Limit Thirty-Five" / "Coffee Stains and Crooked Lipstick" / "What I Keep"). The reviewer's central diagnosis: "The system observes brilliantly in verses, then thesis-summarizes in choruses and bridges." Full WAR ROOM doc: docs/WAR-ROOM-CONCRETE-TO-ABSTRACT-DRIFT-2026-05-28.md. 10 inline tier-2 banned-phrase nudges shipped at B3412. SA#34 candidate ("The chorus is INHABITED, not DECLARED") recorded.
- R5-B0 — 10 inline tier-2 banned-phrase nudges. [AUTONOMOUS · P0 · 1 build] Shipped at B3412. Added 10 entries to
src/lib/banned-terms.tsunderhouse_styletier 2 covering: §2.1 thesis-line shapes (what if this is / this is everything / couldn't make me / i'm done rehearsing), §2.2 told-not-shown patterns (rationing my / more than remember(-ing) / instead of remember(-ing)), §2.3 mixed-metaphor failures (the bruise you / survival over). Post-generation scan catches these; gauntlet/refine pass rewrites them.
- R5-B1 — ATL primitive (Abstract-Thesis-Line detector). [AUTONOMOUS · P0 · ~3 builds] Shipped at B3413 (heuristic stage).
src/lib/claude/audit-abstract-thesis-line.ts— pure function with curatedTHESIS_OPENERSlexicon (~25 entries) + concreteness ratio (reuses R&B NCD lexicons) + two flag rules (thesis-opener with low concreteness OR lone abstraction). 10 unit tests pass. Inverse polarity from NCD: HIGHER score = MORE thesis-mode (BAD). Phase 2 (Haiku judge for adjudication) queued as R5-B1.2.
- R5-B2 — STDD (Show-don't-Tell Density) — generalize R&B NCD/AID rule cross-arc. [AUTONOMOUS · P0 · ~2 builds] Shipped at B3413.
src/lib/claude/audit-show-dont-tell-density.ts— runs the AID rule on chorus + bridge sections regardless of arc. Single cross-arc threshold STDD_FLOOR = 0.30. Same polarity as NCD (higher = more showing = good). Per-section diagnostic with weak-line indices. 11 unit tests pass. R&B's NCD weight stays in R&B's arc-specific domain.
- R5-B3 — EFI primitive (Emotional Friction Index). [AUTONOMOUS · P0 · ~2 builds] Shipped at B3414.
src/lib/claude/audit-emotional-friction-index.ts— Haiku-judged with 4-tier verdict (one-note / hinted / resolved / layered). Parser auto-downgrades layered/resolved verdicts that ship with empty evidence OR null secondary — catches Haiku's tendency to ratify without backing. Fail-safe: timeout/parse error returns state='error' + score 100 (null component). 15 unit tests pass.
- R5-B4 — SA#34 ratification ("The chorus is INHABITED, not DECLARED"). [AUTONOMOUS · P1 · 1 build] Shipped at B3415. SA#34 added to
src/lib/brand.ts(BRAND.sacredAccident34) +src/lib/sacred-accidents/data.ts+docs/SACRED-ACCIDENTS.md. Companion ratification: SA#33 (lane-vs-scene split, R5-A31) shipped in the same build since its empirical evidence (B3411 catalog regen) had already landed. Two ratifications closed simultaneously.
- R5-B5 — IRD primitive (Internal Rhyme Density) promoted cross-arc. [AUTONOMOUS · P1 · ~2 builds] Shipped at B3417.
src/lib/claude/audit-internal-rhyme-density.ts— scans each chorus block for: (a) end-rhyme density (pairwise terminal-vowel-bucket matches), (b) internal-rhyme rate (within-line repeated vowels), (c) slant-rhyme bonus (vowel match + consonant tail mismatch). 7-bucket vowel-phoneme proxy approximates without CMU dict (~75-85% accuracy on common English lyric vocabulary). Composite block score weights end-rhyme 60%, internal 25%, slant 15%. 4-tier verdict (flat / modest / engineered / sonic-locked). 9 unit tests pass. Closes §2.4 + §2.5.
- R5-B6 — Forge chorus pre-declaration rule (generalized SA#21 / SA#34). [AUTONOMOUS · P1 · 1 build] Shipped at B3419.
src/lib/claude/chorus-discipline-baseline.ts— CHORUS_DISCIPLINE_BASELINE constant + isChorusDisciplineBaselineEnabled gate. Wired intosingle-phase-prompt.tsBEFORE the per-mode chorusDiscipline injection so per-mode rules layer on top as specialization. Two-axis rule: AXIS A sonic (open-vowel chorus peak) + AXIS B semantic (named object from verses). Banned thesis-mode shapes from B3412 WAR ROOM (Every X / What if X / I'm done X-ing / X more than Y / X over Y / self-help-book gut check). BehindSF_FORGE_CHORUS_DISCIPLINE_DISABLED=1kill-switch with warn-log on every forge call when set (mirrors SA#29 pattern). Prompt budget: 89% of 20k tokens (~2k headroom). 9 unit tests pass.
- R5-B-baseline — Vault chorus-baseline measurement script. [AUTONOMOUS · P1 · 1 build] Shipped at B3420.
scripts/bench-chorus-baseline.ts+npm run bench:chorus-baseline. Queries top-N songs by forge_score, runs the 5 pure-function primitives (ATL+STDD+IRD+CVC+PLV) on each, aggregates + writes JSONL. Pre-B3419 baseline captured against vault top-100 (mean forge_score 91.2): ATL 1.8, STDD 5.8 (69% below floor — known R&B-lexicon limitation), IRD 69.3 (engineered tier), CVC 81.4 (evolved/alive dominant), PLV 72.6. Full analysis + post-deploy comparison plan indocs/CHORUS-BASELINE-B3419.md. After 24-48h of post-B3419 traffic, re-run with--since 2026-05-29 --label post-b3419to measure the rule's empirical impact.
- R5-B-lexicon — STDD cross-arc lexicon expansion. [AUTONOMOUS · P2 · ~2 builds] Known followup surfaced by B3420 baseline. The R&B NCD lexicon (CONCRETE_NOUNS / ACTIVE_VERBS / ABSTRACT_WORDS) was curated narrowly for R&B vocabulary. Top-100 baseline shows 69% of songs below the STDD floor — not because choruses are abstract, but because country/folk/indie vocabulary (mailboxes, fishing boats, sweaters, etc.) doesn't hit the lexicon. Expand the shared lexicons to cover the 9 arcs evenly. Re-baseline after expansion to confirm the below-floor rate drops without false-positive reduction in scoring sensitivity.
- R5-B7 — MCRT primitive (Metaphor Close-Reading Test). [AUTONOMOUS · P1 · ~2 builds] Shipped at B3418.
src/lib/claude/audit-metaphor-close-reading.ts— Haiku-judged 4-dimension grading (clarity / consistency / weight / truth) for each identified metaphor line. 4-tier verdict (sonic-only / partial / felt / load-bearing). Parser enforces consistency rule: load-bearing requires zero failed lines (auto-downgrades to felt + score≤75 if violated). Per-failed-line diagnostic includes specific failed dimensions + close-reading critique. 15 unit tests pass. Same fail-safe pattern as EFI + through-line.
- R5-B8 — Soul-pop substyle audit + per-substyle forge fragment. [AUTONOMOUS · P1 · ~2 builds] Closes §2.9. Audit current pop substyle profiles to verify soul-pop carries: tight 4-8 syllable phrases, percussive consonant placement, gospel-chord harmonic-motion cue, chorus hook chain repeats 3-4× per chorus. Add substyle-detection-aware forge fragment that injects right discipline per-substyle.
- R5-B9 — CVC primitive (Chorus Variation Cost). [AUTONOMOUS · P1 · 1 build] Shipped at B3416.
src/lib/claude/audit-chorus-variation-cost.ts— parses distinct chorus blocks (Chorus / Hook / Refrain), runs pairwise line-identity rate (case-insensitive, whitespace-normalized), 4-tier verdict (static / thin / evolved / alive). Single-chorus songs return score=100 (no penalty). 9 unit tests pass. Inverse polarity: HIGHER = MORE variation = good.
- R5-B10 — PLV primitive (Phrase Length Variance). [AUTONOMOUS · P2 · 1 build] Shipped at B3416.
src/lib/claude/audit-phrase-length-variance.ts— per-section word-count coefficient-of-variation with target bands: verse [0.15, 0.70] (Pat Pattison asymmetric stanzas), chorus [0, 0.30] (singability discipline), bridge unbounded. Word count is a syllable proxy (CMU-dict-style phonetic lookup deferred to PLV phase 2). 8 unit tests pass. Flags include diagnostic clauses ("verse lines vary widely — prose-like, may rush on melody" / "chorus phrases vary widely — singability suffers").
- R5-B11 — AR primitive (Ambiguity Reserve). [AUTONOMOUS · P2 · ~2 builds] Closes §2.8. Haiku-judged: does the final chorus / outro leave at least one unresolved element (a lingering question, a visible bruise the song doesn't fix, a contradiction the song acknowledges but doesn't close)? Resolution-too-tidy is the failure mode.
Round-5 sequencing notes
Ship R5-A1 + R5-A2 + R5-A3 + R5-A4 in a single bounded session (~1.5h total). All four are <1h, all four are P0, and shipping them as a single push moves the system out of the "audit findings open" state cleanly. R5-A14 (the N=large eval) is the highest-leverage Tier-3 item because it's the only thing that converts the 12-build J-phase from "we believe this works" to "we measured this works" — that single measurement unblocks the public claim moves R5-A11, R5-A16, R5-A18.
The most surprising audit finding: the brand-promise gap from B3372 WAR ROOM §7 is STILL UNADDRESSED despite 12 builds of register-awareness work shipping. R5-A4 is the cheap fix; R5-A11 is the substantive one. Both are necessary; neither is sufficient.
Deep Audit 2026-05-24 round 4 — autonomous queue (TOP PRIORITY)
Triggered by: the B3206 Deep Audit, captured at the end of the 13-build session (B3193 → B3205). Final grade 8.4/10 (B+). Strongest categories: Trust (9.3), Process maturity (9.5), Strategic differentiation (9.2). Weakest: Retention (6.5 → still climbing), Information architecture (7.0), Onboarding (7.2).
Operator-side complements live in docs/OPERATOR-PUNCH-LIST.md under the matching "Audit-Round-4" section. The engineering items below ship autonomously.
Round-4 Tier 1 (ship this week — small, each ≤2h)
- R4-A1 — `npm run docs:refresh` to close the 85-commit-stale-data leak on /engineering. Operator ran this in the same session; committed at 67b3f7ae0 alongside B3206. Public surface no longer lies about freshness.
- R4-A2 — `npm pack` + tarball-install smoke-test ratchet. Shipped B3209 (2026-05-24):
scripts/check-npm-package.tswalks publishable packages underpackages/*, runsnpm run build+npm pack, installs the tarball into amkdtemp'd consumer project, and dynamically imports the main entry — assertion:import('@songforgeai/<name>')must resolve to a non-null module on Node 20+. All 4 publishable packages (@songforgeai/agent-room@0.1.0,@songforgeai/fidelity-standard@0.2.1,@songforgeai/scoring-rubric@1.2.0,@songforgeai/client@0.4.0) survive the smoke test. Wired intocheck:allbetweenreduced-motionandmoat-ledger-health. Catches the B3197 class-of-bug (works-in-monorepo, breaks-once-installed) AND the B3204 class-of-bug (build artifacts not actually emitted). Failure detail captures last 20 lines of consumer-side stderr so the offender's stack is in the CI log.
- R4-A3 — Surface auto-audit + email-gate + finalize success rates on `/admin/system-health`. Shipped B3210 (2026-05-24): added
SurfaceTelemetryReportto/api/admin/system-health+ a 3-card row on the admin page covering: (1) Forge auto-audit — last-1h + 24h counts fromsong_audit_runs+ mean 24h composite; (2) Surgery finalize — last-1h + 24hcomplete-status counts fromsong_surgery_sessions+ 7d abandon-rate proxy; (3) Email gate — RESEND_API_KEY presence + SF_EMAIL_DISABLED kill-switch + SF_EMAIL_RETENTION_ENABLED flag with live/dev-only/disabled glance status. Service-role client used for the cross-user counts; each sub-query catches its own errors + falls back to status='unknown' so the page never fails to render. Closes the "currently visible only by grepping production logs" gap the audit named. Note:email.sendper-kind count remains pending anemail_sendstable — surfaced as the gate state for now (live volume is low; structured logs are still the source of truth for per-send accounting).
- R4-A4 — Reduced-motion ratchet cleanup (124 → 104 baseline). B3208 (2026-05-24) migrated 14 components to the
motion-safe:prefix: RepairFlow ×5 + EmailCaptureCard + HomepageForgeDemo + Navbar + NewsletterCapture + RadioBar + RadioPanel + ReferralCard + SamplePlayer + SectionRewriteModal + SparkIdeas + StickyCTA + VoiceConsistencyCard + VoiceRewriteFlow ×2 + genres/GenreRadio. Net: 10 unguarded animation sites eliminated; ceiling tightened 114 → 104. Headroom: 0 (any new unguarded animation now fails CI). End-of-quarter target: < 50; end-of-year: 0.
Round-4 Tier 2 (ship this month — bigger, each <1 day)
- [~] R4-A5 — `RepairFlow.test.tsx` foundation: 8 state-mutation paths. Foundation shipped B3211 (2026-05-24): 11 tests cover the simple paths + harness. Initial render (1+items / empty/
Session complete), per-item card structure (testids + button copy), progress strip ("Issue 1 of 3"), Skip path advance, Reject path advance, Skip×3 → "Session complete", resume-banner visibility (set / null / start-fresh dismissal with async fetch await), severity border classes (red/orange). Harness usesvi.stubGlobal('fetch', mock)+ a default-OK envelope so unscripted endpoints (init, edit-log, finalize) don't blow up. Still pending for full R4-A5 closure: Accept happy path (judge-edit verdict assertion) / Accept with judgeError fallback / Undo-to-Edit (B3192) / Undo-and-Skip (B3192) / Revert-to-baseline / Resume button API contract. These 6 paths each need 1-2 tests + scripted fetch mock for the corresponding endpoint; estimate ~2-3h of follow-on work. Marked [~] (in-progress) rather than [x] so the remaining paths stay visible.
- R4-A6 — Surgery WAR ROOM remaining P0 fixes: #2, #4, #5. Shipped B3213 (2026-05-24): all three P0s closed in one bounded build. (P0 #2)
RepairFlow.tsxresume-hydration now setsjudgeErroron ANY non-canonical verdict — pre-3213 onlye.verdict === 'unjudged'set the marker, so a forward-compat verdict like 'neutral' or 'mixed' would silently fall back to'changes'+ corruptnetDimensionScore(aggregateSessionDeltascorrectly skips judgeError'd entries but skipped the 'neutral'-treated-as-'changes' case). (P0 #4) AddedliveEditIndexRef = useRef<number>(0)+bumpEditIndex(next)helper; replaced all 4liveEditIndex + 1closure reads withliveEditIndexRef.current + 1(the rapid Accept→Undo-and-Skip race where both writes observed the same stale state). AllsetLiveEditIndexcall sites swapped tobumpEditIndexto keep state + ref in sync. (P0 #5) Hardened theseenItemIds.add()guard to require non-empty string AND non-null (defense-in-depth — even though the existingif (e.surgeryItemId)already covered the failure mode). Verification: typecheck clean, 104/104 surgery + lib tests green. Audit doc unchanged (B3196 historical record); follow-on Surgery WAR ROOM items (P1/P2) tracked separately if/when they surface.
- R4-A7 — Postgres BEFORE UPDATE trigger on `song_surgery_sessions.status` transitions. Shipped B3214 (2026-05-24):
supabase/migrations/enforce_surgery_session_status_transitions.sql. Definespublic.enforce_song_surgery_session_status_transition()function + attaches it as a BEFORE UPDATE OF status trigger onpublic.song_surgery_sessions. The function rejects any change that isn't in the allowed-transitions map (kept in sync withsrc/lib/surgery/session-types.ts:ALLOWED_TRANSITIONS); no-op updates (NEW.status = OLD.status) pass through unchanged so other column writes (diagnostic_snapshot / verification_record / session_note) don't get blocked. Terminal-state protection: any UPDATE attempting to leave 'complete' or 'abandoned' raisescheck_violationwith a HINT pointing at the canonical map file. The trigger closes the canTransition() TOCTOU race window B3199 surfaced — two concurrent UPDATEs that both pass the app-layer SELECT-then-check can no longer both succeed; the DB rejects the second one atomically. Implemented as CHECK-style trigger (not a column CHECK constraint) because column constraints can't reference OLD.* values to enforce direction. B3214 sentinel comment +COMMENT ON FUNCTIONfor discoverability viagit grep B3214.
- R4-A8 — Auto-`docs:refresh` cron + auto-PR. Shipped B3212 (2026-05-24):
.github/workflows/docs-refresh-nightly.ymlruns at 09:15 UTC daily + on manualworkflow_dispatch. Checks out main with full git history, runsnpm run docs:refresh, then usespeter-evans/create-pull-request@v6to open (or update in-place) a PR titledchore: nightly docs:refresh — sync auto-state + tighten ratchetson the canonical branchchore/docs-refresh-nightly. PR body lists the two classes of artifact regenerated (CLAUDE.md auto-state block + the 6 count-based ratchet ceilings: console-log, console-error, explicit-any, design-tokens, launderer-casts, reduced-motion) plus loop-guard documentation. Direct cure for Sacred Accident #18 ("the cadence ritual itself drifting") — even if every operator-side cadence slips, the cron catches the drift within 24h. PR-rather-than-direct-push because docs:refresh tightens enforcement gates (ratchets), and the operator deserves a single review point before the gate gets stricter.
- *R4-A9 — `/admin/feature-flags` runtime panel — consolidate the SF__DISABLED env toggles. ✅ Shipped B3286 (2026-05-25). The audit had estimated "~10"; the codebase grep found 22* `SF__DISABLED
/SF_*_ENABLEDenv-var feature toggles scattered across forge / refine / gauntlet / finalize-song / preforge / post-forge-pipeline / leaderboard / various critic loops. All 22 now ship registry entries insrc/lib/feature-flags.tswithcategory(forge / refine-gauntlet / eval-audit / leaderboard / email / experimental / admin) +loadBearingflag +warningstring naming the failure class each kill switch unlocks (SA#29 for the artist locks). New/admin/feature-flags` page surfaces all 24 entries grouped by category + per-flag "Flip in Vercel env" deeplink. Pre-existing 2-entry registry consolidated under the same shape. Closes the env-toggle-proliferation 5X Plan kill.
Round-4 Tier 3 (next quarter — multi-build)
- R4-A10 — VAULT-OPEN-4 (per-user vault graduation). 4-6 builds. Biggest remaining moat play. Scope
song_audit_runs+song_embeddings+exemplar_statusby user_id; gate dashboards per-user; drop admin-onlyrequireAdminchecks in favor of per-user RLS. Requires a per-user audit-on-forge hook (B3203 already wires this — extend the trigger to non-admin users) + a per-user auto-curate when N exemplars accumulate.
- R4-A11 — Crucible-first homepage flow. The 5X Plan's #1 move. The Crucible is the only thing in the system nobody else can copy. Currently it's section 1 of the homepage but the primary "Open the forge" CTA leads to auth wall. Restructure: Crucible-as-hero with the live paste-box above the fold; everything else (forge demo, leaderboard, before/after) becomes secondary proof. ~6-8h of UX work + copy iteration.
- R4-A12 — Multi-language Phase 1 (Spanish first). #42 from Tier-C list; operator trigger phrase
Continue Multi-Language Phase 1. The Latin genre arc work (L1-L13) is the natural launchpad; the dialect-coherence audit primitive (DCS) extends to language-coherence. RFC-0009 covers the seal-annotation contract.
- R4-A13 — Resumable forges Phase 2 (durable event log). #23 from Tier-B list. Migration shipped at B2887; persistence layer at
src/lib/forge-events/persistence.tsshipped. Two builds remain: (a) wire writes from the live forge SSE stream intoforge_eventswith batched inserts; (b) extend/api/songs/forge/resumeto read fromforge_eventswith?lastEventId=Nfor byte-identical replay. Closes the "SSE disconnect loses commentary" gap.
- R4-A14 — Kill duplicate IA between /genres and /lyrics. 5X Plan kill #2. Partially done at B3056 (Footer-link IA consolidation). Remaining work: audit which
/lyrics/[slug]pages are still useful as SEO long-tail vs. redundant with/genres/[slug]. ~3-4h grep + decisions + redirects.
Round-4 Tier 4 (category-defining engineering bets)
- R4-A15 — "Trust Engine" public dashboard. Every score on the site queryable by external auditors via the npm SDK + the seal. New page at
/trustthat lets anyone enter a seal{rubricVersion, model, temperature, buildSha, build}+ verify it. Currently the/verifypage does this for ONE score at a time; the Trust Engine surface is the operator-facing query interface. ~1-2 builds.
- R4-A16 — Per-user voice fingerprints as a product surface (#43 Phase 2). Currently scaffolded. Turning this into a real "your voice over time" product surface creates the kind of personal-history-stickiness Strava/Duolingo trade on. Depends on VAULT-OPEN-4 + Hit Calibration Corpus accumulating N≥10 entries per user. ~3-4 builds.
- R4-A17 — Sacred Accidents public publish surface. ✅ Shipped B3285 (2026-05-25).
/sacred-accidents/[slug]per-SA permalink route + per-SA OG card route + JSON-LDArticlestructured data + prev/next navigation + citation footer linking toBRAND.sacredAccident{N}+ the canonical doc on GitHub. Reads from the existing typedsrc/lib/sacred-accidents/data.tsmirror (no duplicate source of truth); newaccidentSlug()+parseAccidentSlug()helpers in the data module with round-trip test coverage (8 new tests). Index page links each entry to its permalink via the "Permalink →" affordance. The publishing CADENCE remains operator-side (META-R4-1 indocs/OPERATOR-PUNCH-LIST.md) — engineering side is closed.
Round-4 sequencing notes
The Tier-1 items above (R4-A2 / R4-A3 / R4-A4) should ship FIRST — they're cheap and they close the gaps the audit explicitly flagged. R4-A5 (RepairFlow tests) is the biggest single-build win. R4-A11 (Crucible-first homepage) is the highest-leverage UX move but requires coordinated copy + design iteration; queue it AFTER the operator's R4-T2-1 work on Brett's case study (so the new homepage flow can incorporate the new case-study artifact).
Deep Audit 2026-05-19 round 3 — autonomous queue (TOP PRIORITY)
Triggered by: the post-B2746 Deep Audit. The audit found engineering + discipline at A+, distribution + external validation at C. The 15 items below convert audit findings into shippable engineering work; the operator-side complements (arXiv submission, npm publish decision, partnership outreach, testimonial permission, hit-corpus curation) live in docs/OPERATOR-PUNCH-LIST.md under the new "Audit 2026-05-19 round 3" section.
Sequencing principle from the audit: every Tier-1 item is correctly an engineering ratchet. Ship in priority order P0 → P1 → P2 → P3 — but skip ahead when a P1 unblocks a P0 (e.g. prosody backfill is P2 but unblocks user-trust for songs forged before B2746, so promote when convenient).
P0 — ship this week (audit Tier 1)
- AUDIT-1 (B3197) — Email capture on Crucible verdict screen shipped via Resend. Operator picked Resend (already wired for the Life Songs / Heirlooms flow at B2591 — same verified
noreply@songforgeai.comsender + sameRESEND_API_KEYenv var). NEWsrc/lib/email/send.ts— extracted shared sender (sendTransactionalEmail,isValidEmailShape,escapeHtml,baseEmailHtml) so the Crucible-capture + retention surfaces don't duplicate the life-songs envelope. Fail-open contract;SF_EMAIL_DISABLED=1global kill switch; structured-log eventsemail.send.{sent,disabled,no_api_key,error,threw}. 21 unit tests pin the env-toggle precedence + the validator + the HTML escape. NEWsupabase/migrations/add_crucible_email_captures.sql— citext-typed email column (case-insensitive UNIQUE),verdict_bannerdenormalized for fast segmentation,ip_hash+user_agent_hashfor audit trail (NEVER raw PII),opted_in_to_marketingseparate from capture so future unsubscribe flow can flip the flag without losing the row, RLS enabled with default-deny (service-role-only writes via the capture endpoint, admin reads only). NEWPOST /api/crucible/[id]/capture-email— unauthed, IP-rate-limited via newRATE_LIMITS.crucibleCapturebucket (30/IP/day; entry added alongside the existing crucible + crucibleSave entries insrc/lib/api-auth.ts), validates email shape, upserts ON CONFLICT (email) so re-capture is idempotent + setsalreadyCaptured: true, optionally mails the verdict back to the user via the new sender (fail-open — DB persistence is the load-bearing path). NEWsrc/components/crucible/EmailCaptureCard.tsx— client component mounted on/c/[id]between VerdictShareButtons and the CTA cluster. 4-state form (idle/submitting/captured/error) with two explicit checkboxes ("Email me the verdict" + "Add me to the mailing list"); maps every endpoint error code to operator-friendly copy; green confirmation card on success; full unsubscribe microcopy. The audit's #1 highest-leverage move — turns unauthed Crucible traffic (high-volume, free) into a list the operator can reach for moments-of-decision outreach (rubric updates, v1.0.0 fidelity-standard release, price changes, new genre arcs). Operator-side todo: apply the migration in Supabase.
- AUDIT-8 (B3197) — Email retention loop scaffold (4 transactional templates), feature-flagged. NEW
src/lib/email/retention-templates.tsships the four canonical lifecycle moments per the Deep Audit round-3 finding (retention gap is load-bearing):sendWelcomeEmail(profile)(post-signup),sendFirstForgeEmail(profile, song)(first finished forge),sendWeeklyLeaderboardEmail(profile, snapshot)(Mondays when rank moved),sendDormant30dEmail(profile)(30 days idle). Every template gated bySF_EMAIL_RETENTION_ENABLED=1— defaults OFF so the templates can ship behind a flag while the operator (a) sets up Resend suppression lists, (b) finalizes unsubscribe-link copy, (c) audits real samples before flipping. When the flag is unset, each function logs a structured "gate_closed" event the operator can grep for. Templates usebaseEmailHtml+escapeHtml+ the shared sender; movement copy in the leaderboard template branches on first-touch / moved-up / dropped / held; first-forge template includes the rubric composite score when present, omits when null. Each template carries unsubscribe microcopy. 12 unit tests pin: gate closes by default; SF_EMAIL_RETENTION_ENABLED=0 also closes; each template uses its ownkindtag; first-forge omits score when null; weekly leaderboard branches correctly across all 4 movement cases; greeting falls back to "Hi there," when displayName is null. The TRIGGER wiring (auth/callback hook for welcome, finalize-song hook for first-forge, /api/cron/weekly-leaderboard-touchpoint, /api/cron/dormant-30d-touchpoint) lands in follow-up builds once the operator approves the template copy. - AUDIT-2 — Refine route prosody recompute (sibling to B2746). B2749 — shipped; checkbox flipped at B2872. PATCH /api/songs/[id] now mirrors the B2746 gauntlet fix: when the body carries new lyrics, scoreProsody is recomputed before persisting + the fresh report ships in the update payload. SF_PROSODY_REPORT_DISABLED=1 env gate honored, safeSync wraps the call, null/zero-line reports leave the column untouched.
- AUDIT-3 — Wire B2743 (motif ledger) + B2744 (bridge architecture) telemetry into /admin/forge-discipline. B2753 — shipped. Two new tiles on
/admin/forge-discipline: (a) motif cluster rate (% of songs with 2+ watched motifs from the same hot-list the B2743 ledger watches), target ≤20%, baseline 80% from the May 17 audit; (b) bridge old-default rate (% of songs with co-occurring whispered + crack directives — the audit's #1 system tell), target ≤30%, baseline 93/100. Top 5 motifs by song-presence rate also surface so regressions show by name. The Wave 2 fix is now measured. - AUDIT-4 — Hero compression: 1 CTA + receipts strip. B2755 — shipped. Hero already had 1 primary CTA (HeroCruciblePaste; B2106 closed the dual-CTA dilution previously). Added 4-pill receipts strip above the fold between the privacy note and the demote-CTAs: Public rubric · Signed verdicts · Incident log · The receipts. Each pill links to a real public artifact (/scoring/standard, /verify, /incidents, /the-receipts). The duplicate chip strip 3 sections down was removed to avoid double-rendering.
P1 — ship this month (audit Tier 2 — high-leverage)
- AUDIT-5 — Wire B2745 (chorus-compression) UI into dashboard SongDetail. B2757 — shipped. New
ChorusCompressionPanelmounts in SongDetail right after HookCompressionPanel; on-demand button generates poetic / plainspoken / commercial register variants + recommendation via a newPOST /api/songs/[id]/chorus-compressionroute. Pro-tier-gated (402 + upgrade CTA for free/creator); 30/hour rate limit; copy-to-clipboard per variant. 422 with actionable detail when the song has no[CHORUS]marker. Sibling pattern to hook-compression but user-initiated (no auto-compute at forge time — keeps Haiku cost off the default path). - AUDIT-6 — Annual-billing default on /pricing. Audit false-positive — already shipped at B1901. The audit's WebFetch saw the SSR'd
billingIntervalstate but couldn't tell which value was the default. Source confirmsuseState<'month' | 'year'>('year')atsrc/app/pricing/page.tsx:97. B1901's commit comment names the same audit-driven motivation: "Linear, Notion, Slack all default to annual." Validated at B2750. - AUDIT-7 — Cut homepage to 7 sections. B2755 — shipped. 4 sections cut: (1) Mini-Demo / HomepageForgeDemo (hero already ships HeroCrucibleDemo — dual-demo dilution); (2) Perspective Atlas + Persona Forge teaser (/atlas + /persona-forge are dedicated surfaces reachable from Navbar + footer); (3) "What we don't do" / The Disciplines We Don't Break (content lives at /manifesto + /principles/topic-sovereignty); (4) Homepage FAQ accordion (/pricing FAQ is the single source of truth; FAQPage JSON-LD still ships from the top of page.tsx so SERP rich-snippets aren't affected). Net section count: 8 → 5 raw sections (+ HomepageBeforeAfter inline = 6 logical sections). File: 837 → 712 LOC (-125 lines).
- AUDIT-8 — Email retention loop scaffold (4 transactional templates). Welcome / first-forge / weekly-leaderboard / dormant-30d. Templates + send infrastructure ready behind a feature flag; the operator decides when to flip live (OPERATOR list A11). Per audit, the retention gap is the load-bearing missing surface.
- AUDIT-9 — Non-English fairness fix OR formal defer. B2751 — formal deferral landed. Per the round-3 brief, this item required ship-or-defer disposition. After review: Phase A1-A3 (forge-prompt fairness fixes) already shipped at B1982-B1993; Phase C2/C5 (corpus-anchor add + 68-song rescore) is operator-side work that needs anchor curation before the rescore can run. New hard deadline 2026-06-30 logged in
docs/TRUST-DECAY-AUDITS.md. The disclosure post ships with the gap-closure number measured, not as a "working on it" half-trust signal. If C2+C5 aren't landed by 2026-06-30 the disclosure post ships unconditionally with the gap still open. - AUDIT-10 — /case-study/brett page (skeleton + draft). Working-songwriter case study; Brett ran the entire P0+P1 sweep through the product over 24-48 hours and surfaced 12 actionable findings. Real external validation. Page ships behind a draft-flag pending operator permission (OPERATOR list A6 → A12); the engineering work is the page + structured-data + cross-links.
- AUDIT-11 — /case-study index + 2 more songwriter slots reserved. B2754 — shipped. Sibling format to
/case-studies(quantitative lyric-pipeline). New/case-studyindex lists visible entries + 2 reserved-slot placeholders ("different genre / different cycle" and "international / multilingual"). Brett entry stays hidden from index until DRAFT flips false; index renders an empty-state explainer in the meantime so the surface doesn't look unfinished while operator A6/A7 outreach is in flight.
P2 — ship this quarter (audit Tier 2 — cleanup)
- AUDIT-12 — /status page with 90-day uptime + recent incidents. B2756 — shipped. New
IncidentTimelinesection atop /status renders a 90-day calendar grid (green = no incident, gold = publicly logged incident), counts in window, days-since-last, list of incidents in window with severity chips, full-history link. Honest disclosure: we don't publish an uptime % because per-minute outage duration isn't captured — the receipt-shaped truth is the incident-count + per-day grid. Trust-positive per the audit's framing. - AUDIT-13 — Bundle-size ratchet: re-baseline OR retire. B2729 was the disposition (round-2 Tier-3 #19); same-day-as-audit timing meant the round-3 audit didn't catch the fix. B2752 confirmation: ratchet is green AND ratcheting downward. Current bundle 4,024,440 bytes (3.84 MB) shrunk -2779 B vs B2729 baseline; ratchet ceiling 4.43 MB, lastKnownTotal updated to the new measurement so future builds must hold-or-shrink against the tighter floor. Item closed.
- AUDIT-14 — Prosody backfill script for pre-B2746 songs. B2758 — shipped.
scripts/backfill-prosody-reports.tswalks the songs table; for every complete song with lyrics, recomputes scoreProsody and writes the fresh report. Idempotent; --dry-run default; --only-stale optional filter skips rows whose existing report already matches the fresh computation.npm run backfill:prosody-reports -- [--live] [--limit N] [--only-stale]. Per-song cost is microseconds (pure sync function — no Claude calls); ~700 songs take ~5 min serial dominated by Supabase round-trips. Failure-safe per-row; final summary prints updated / skipped / failed counts. - AUDIT-15 — Schema.org markup audit. B2885 — shipped. New
scripts/check-schema-markup.tsratchet enforces required JSON-LD@typevalues on/blog/[slug](BlogPosting|Article + Person + Organization),/songwriting/[slug](Article|HowTo|Course + Organization),/about(Person + Organization), and/scoring/standard(ScholarlyArticle|TechArticle|Article|WebPage + Organization). All 4 surfaces pass on first run. Wired intocheck:all(41st gate) + dedicatedci.ymlstep. The 0-violation ratchet the audit asked for, forward-only. - AUDIT-16 — Sacred Accidents framework as public artifact. B2570 — shipped; checkbox flipped at B2872. Public surface lives at
/sacred-accidents(NOT/about/sacred-accidents— flat URL preferred). Renders the typedsrc/lib/sacred-accidents/data.tssource. OG image at/sacred-accidents/opengraph-image.tsx.scripts/check-sacred-accidents-sync.tsratchet (B2699) enforces that every## SA#Nheader indocs/SACRED-ACCIDENTS.mdhas a matchingSACRED_ACCIDENTSentry in data.ts. Currently surfaces SA#11–#19 in full + stubs for #1–#10 awaiting historical reconstruction. - AUDIT-17 — Pricing FAQ cut to 6 items. B2872 — shipped. PRICING_FAQ trimmed 12 → 6 items: voice flattening, training data, ownership, free tier, refund, plan changes. The 6 cut items (music generation, credit packs, why upgrade, Creator→Pro trigger, ChatGPT comparison, Topic Sovereignty) moved to a new EXTENDED_FAQ const + rendered on the new public
/faqpage. /pricing FAQ ends with a "More questions answered on the canonical FAQ →" link to /faq. The FAQPage JSON-LD on /pricing now reflects only the 6 kept items (matches what renders).
P3 — defer until a deliberate sprint slot opens
- AUDIT-18 — Forge V2 cold-visitor SSR shell. B2883 — shipped. New
src/app/forge/ForgeColdShell.tsxserver-renderable component replaces the pre-hydration blank<div>Suspense fallback with H1 + value prop + skeleton prompt textarea + skeleton genre/voice selectors + "Sign in to forge" CTA + Crucible escape hatch. Cold visitors + crawlers + first-paint humans now see meaningful content in the first 200ms instead of an empty viewport. - AUDIT-19 — Tool discovery card on /forge. B2881 — shipped. New
src/components/ForgeFirstVisitHint.tsxrenders aboveForgeToolsExplaineron the idle composing surface: cyan-bordered callout, "First time here? Start with the Crucible — paste any lyric, get an 8-voice verdict in 30 seconds." localStorage-persisted dismissal so returning users never see it after closing it once. Pairs with the explainer below (which disambiguates the 3 tools); this hint adds the time-based first-visit nudge the audit specifically asked for. - AUDIT-20 — Mobile audit pass on /scoring/standard + /pricing. B2886 — shipped. New
e2e/mobile-audit.spec.tsPlaywright spec enforces two invariants at iPhone-13 viewport across 5 SEO-trust surfaces (/scoring/standard, /scoring/standard/whitepaper, /scoring/standard/fidelity, /scoring/standard/fidelity/cite, /pricing): (1) no horizontal scroll (document.scrollWidth ≤ viewport.clientWidth + 2px tolerance) — the "wide-table scroll-jail" the audit named; (2) every interactive element ≥ 44×44 px tap target per WCAG 2.5.5 + Apple HIG. Inline prose links exempted from tap-target rule. Future page additions get the invariant suite automatically by appending to MOBILE_PAGES.
Audit's meta-finding (not a build, but a steering rule)
The next 90 days needs ~70% of operator hours on outbound, ~30% on engineering — the inverse of recent history. The engineering velocity is 5x ahead of where it needs to be; distribution velocity is 1/5x. The autonomous queue above keeps engineering moving without operator input; the operator's hours should be re-routed to the OPERATOR-PUNCH-LIST.md audit-round-3 items (arXiv submission, partnership outreach, songwriter testimonial recruitment, hit-corpus curation, HN submission, Substack mirror).
Vault Intelligence V1 — Phase 3+ open threads
Phases 0-3D shipped B3024 → B3051 across the 2026-05-22 session. The 2,429-song admin corpus is audited (v1.2.0), embedded (1536-dim HNSW), classified (96 exemplars + 12 junk), surfaced via 5 admin dashboards, AND retrievable inline in /forge with arc-filtering + per-card selection + full-lyrics mode + localStorage persistence.
These five items are what's left for V1 → V2:
- VAULT-OPEN-1 — Operator spot-check the 96 auto-exemplars (~8 min interactive). Walk
/admin/vault/curate?status=exemplar&arc=indie(etc.) with keyboard hotkeys (E/X/J/R/C/N). Demote any that look wrong. The auto-pass picked top-12 per arc by composite_score; some are probably calibration artifacts (especially in pop/indie where scores skew high). Cheap sanity check that validates retrieval quality.
- VAULT-OPEN-2 (B3169) — Cross-arc score normalization primitive shipped.
src/lib/vault/score-normalization.tsships the pure-compute layer:computeArcScoreStats(rows)(per-arc mean + sample stdDev + sorted scores),computePercentileRank(score, sorted)(Hazen-rule 0..100 percentile with tie half-credit),computeZScore(score, mean, stdDev)(standard z-score with null-safe stdDev=0 handling), and the one-shotnormalizeScores(rows)that takes raw{arc, composite_score}rows and returns the same rows annotated withnormalized_score+percentile_rank+z_score. Null-handling is uniform: null composite_score in → null fields out; arc with N<2 in → null fields out (no variance to normalize against). CompanionpercentileLabel(pr)produces operator-facing chip text ("Top 1%" / "Top 5%" / "Top 10%" / "Top 25%" / "Above median" / "Median" / "Below median" / "Bottom 25%" / "Bottom 10%" / "Bottom 5%" / "Unranked"). The motivating test case ("rock-60 = top of rock, pop-95 = top of pop, both percentile-equivalent") passes — the 35-point raw composite-score gap collapses to the same percentile rank under per-arc normalization. 34 tests pin every helper + edge case (empty input, single-row arc, ties, floating-point clamp, non-finite values). normalized_score == percentile_rank in this calibration; the separate field reserved for a future scaled-z calibration without breaking consumers. This ships the math; the dashboard ranking surfaces consuming it land as a follow-up.
- VAULT-OPEN-3 (B3168) — Inverse-mode "avoid this voice" references shipped (prompt-side).
src/lib/forge/vault-references.tsVaultReferenceinterface gains an optionalinverse?: booleanfield (default false; backwards compatible).buildVaultReferenceBlock()splits the input refs into positive + inverse arrays preserving input order within each group, then renders TWO distinct blocks: (a) the original "VAULT REFERENCES — examples from the operator's corpus...Honor the established voice + craft choices" with EXEMPLAR labels; (b) a new "VAULT ANTI-PATTERNS — voices to depart from...Do NOT honor the voice, register, image vocabulary, or structural shape" block with ANTI-PATTERN labels. The closing directive adapts to which sets are present: positive-only → original "Use the references to anchor voice + craft" copy; inverse-only → "Depart from the ANTI-PATTERN voices — the new song should feel UNLIKE what this operator has written before"; mixed → "Honor the EXEMPLAR voices; depart from the ANTI-PATTERN voices." 18 tests pin: backwards compat (inverse=false renders as EXEMPLAR), polarity routing, mixed-mode ordering (positives always first), per-group independent numbering, input order preservation within each group, closing-directive variants, trailing-separator preservation. The UI side (operator-flagged inverse refs inVaultInspirationPanel.tsxper the punch list's "Sister-song panel pattern works as the UI ('Like this / Not like this' pair)") is a forthcoming build — this ships the prompt-side primitive so the consuming code can wire when the operator hits the inverse-toggle button.
- VAULT-OPEN-4 — Per-user Vault graduation: every paid user gets their own corpus (4-6 builds). Currently only admin tier has corpus. To turn this from operator-only-tool into product-feature: scope
song_audit_runs+song_embeddings+exemplar_statusby user_id, gate the curation/anthropology dashboards per-user, drop the admin-only requireAdmin checks in favor of per-user RLS. Biggest moat play but most work. Requires a per-user audit-on-forge hook + a per-user auto-curate when N exemplars accumulate.
- VAULT-OPEN-5 (B3167) — Contribution Ledger auto-write hooks shipped.
src/lib/contributions/log.tsexportslogContribution(input)— fire-and-forget service-role insert intosong_contributions. Fail-open (never throws, never blocks the song save); env-toggle disabled viaSF_CONTRIBUTION_LEDGER_DISABLED=1; column-missing degradation matches the skillFingerprint / prosodyReport / fidelityAudit defensive patterns. Three call sites wired: (1)finalize-song.tswritesai_generation+whole_song+forge-v2.<BUILD_NUMBER>+ 'substantial' on every successful fresh forge; (2)refine/route.tswriteshuman_refine+whole_song+refine-v1.<BUILD_NUMBER>+ 'substantial' + contributor_id=user.id on every saved user-driven refinement, with the parent songId + version number captured in artifact_locator; (3)gauntlet/route.tswritesai_revision+whole_song+gauntlet-v1.<BUILD_NUMBER>+ 'de_minimis' on every standalone gauntlet completion, with the linesReplaced count in artifact_locator. 13 tests pin the row-builder + env toggle. The Moat 4 ledger is no longer empty — every forge / refine / gauntlet writes a row from B3167 forward.
Recommended order: OPEN-1 (cheap), then OPEN-4 (the moat), then OPEN-2 + OPEN-5 in parallel, then OPEN-3.
Genre Pages Excellence — B3053 WAR Room
Full findings doc: docs/GENRE-PAGES-WAR-ROOM.md. The operator asked for a 100-expert / 100-round audit of /genres, /genres/ [slug], /genres/[slug]/[substyle] with the white-text-readability complaint as the spark + four new features to integrate: 100-prompt catalog per genre, Genre Radio (admin songs with audio + best-lines + sort), batch auto-generate by genre with per-prompt reroll. The audit found 13 presentation issues (top: per-arc color promise unused — every accent is violet — and text-dark-500 BPM caption AA-FAIL) and laid out a 7-build sequence.
Tier 1 — B3054 (smallest defensible build)
- GENRE-1 — Contrast pass. B3054 — shipped. Every
text-dark-400body-copy instance raised totext-dark-300across/genres,/genres/[slug],/genres/[slug]/[substyle],/genres/audit-in-action. Thetext-dark-500BPM caption (AA-FAIL 3.5:1) raised totext-dark-400. SA italic quote lifted totext-dark-50. Twotext-violet-100gradient-body instances (AA-FAIL ~3.1:1) raised totext-violet-50. - GENRE-2 — Mobile pass. B3054 — shipped. Audit-primitives table wrapped in
overflow-x-autowithmin-w-[640px]floor (no silent overflow at 375px). 4-pill stat row usesgrid grid-cols-2at xs →sm:flex sm:flex-wrapat 640px+ to prevent uneven 2×2 wrap. - GENRE-3 — Hierarchy reorder. B3054 — shipped. New order: Hero → Try-It (was position 4) → Substyles (was 5) → Load-Bearing Primitive (was 2) → Audit Primitives collapsed in
<details>(was 4) → Banned → Corpus → CTA. CTAs no longer buried under two slabs of expository copy. - GENRE-4 — JSON-LD polish. B3054 — shipped. TechArticle schema gains
author: { Organization }andmainEntityOfPage: { @id: url }for SERP rich-result eligibility. - GENRE-5 — `/lyrics/[genre]` cross-link. B3054 — shipped. New "lyrics guide →" button added to the corpus block on
/genres/[slug]— the two overlapping-audience surfaces are no longer strangers.
Tier 2 — B3055 + B3056
- GENRE-6 — Per-arc accent system. B3058 — shipped. New
src/lib/genre-arc-theme.tsdefines anArcAccenttoken bundle (16 slots × 9 arcs = 144 static class strings, all JIT-safe). 9 arcs picked per WAR Room: folk emerald-500, rock orange-500, pop violet-500 (canonical), rap amber-500, indie cyan-500, rnb fuchsia-500, worship yellow-500, country red-400 (red-500 was AA-borderline), latin sky-500.getArcAccent(slug)resolves the bundle; unknown slugs fall back to pop. Wired through all violet hardcodes on/genres,/genres/[slug],/genres/[slug]/[substyle]. The 9-card grid on /genres now blooms 9 distinct colors; folk and rap and pop no longer look identical. 6-test invariant suite ratchets the bundle shape + per-arc color uniqueness +-50gradient-text WCAG floor. - GENRE-7 — Per-arc hero glyph + gradient. B3059 — shipped.
ArcAccentinterface gained 3 slots (heroIconlucide name,heroBgwarm/cool token,iconLarge-400 shade). Per-arc icons: folk Trees, rock Flame, pop Heart, rap Mic2, indie Sparkles, rnb Music3, worship Sun, country Sunset, latin Music2. Warm bg for folk/rock/rap/worship/country; cool for pop/indie/rnb/latin. Hero block on/genres/[slug]now wraps in the radial backdrop + renders a 64px tinted glyph above the H1. - GENRE-8 — Failure-mode disclosures. B3060 + B3076 — shipped. New
src/lib/failure-mode-definitions.tsregistry withgetFailureDefinition(arcSlug, modeName). B3060 seeded the rap arc with 10 inline definitions fromdocs/RAP-FORBIDDEN-ARCHIVE.md(5 with examples). B3076 expanded coverage from 10 → 90 — all 9 arcs now have all 10 canonical mode names registered with operational 1-sentence definitions. Country / folk / rock pull verbatim from<arc>-eval-context.ts(canonical names match the eval taxonomy). Pop / latin / rnb / worship use GENRE_ARCS canonical names (which diverge from the eval-context "Diane Warren By Numbers"-style register) and were authored against the SA#21-#27 paradigms + per-arc Forbidden Archive docs. Indie blends 7 eval-context matches + 3 newly-authored (Voice Shift Mid-Song / Detail Without Stakes / Universalist Shortcut). Cross-check script confirms 90/90 canonical → registered with zero orphans. UI on/genres/[slug]wraps each red chip in a<details>with the mode name as<summary>+ chevron; expansion reveals the definition + optional example. - GENRE-9 — 100-prompt catalog data file (infrastructure + seed). B3061 — shipped. New
src/lib/genre-prompt-catalog.tsships theCatalogPromptinterface (extendsGenreShowcasePromptwithid+mood/tags),getCatalogPrompts(slug)loader (falls back to the 3-per-arc showcase seed when no expanded catalog is registered), andgetRandomCatalogPrompt(slug, { excludeIds, substyle })for reroll flows. 8 unit tests cover seed presence + unique ids + stable refs + filter behavior.scripts/generate-genre-prompt-catalog.tsships the Haiku batch generator (per-arc + --all flags; outputs totmp/genre-catalog-<arc>.jsonfor operator curation). Operator runs the script + curates the output to expand from the 27-prompt seed to 900. UI consumes via the loader, agnostic to whether the catalog is at seed (27) or expanded (~900). - GENRE-10 — Featured-prompt hero + scroll strip + Shuffle. B3062 — shipped. New
src/components/genres/TryItSection.tsxclient component renders: (a) one LARGE featured-prompt card with per-arc gradient + prominent Forge button + Shuffle, (b) horizontal scroll strip of next 12 catalog prompts as chips (scroll-snap mandatory), (c) "Browse all N →" link to the future/genres/[slug]/promptspage. Shuffle uses session-scoped rolled-set to avoid repetition until the pool exhausts. SSR-safe: server renders the first catalog entry as the initial featured; client takes over for reroll. Replaces the 3-static-card grid. The Try-It section is now the page's primary conversion engine, not wallpaper.
Tier 3 — B3063 (all 5 items in one vertical ship)
- GENRE-11 — Genre Radio data layer. B3063 — shipped. New
/api/genres/[slug]/radioroute. 2-step query: pullssong_audit_runsrows at v1.2.0 matching the arc, intersects with admin-published songs that haveaudio_url. Filters bysubscription_tier = 'admin',is_public = true,status = 'complete'. Cacheds-maxage=120, stale-while-revalidate=600. - GENRE-12 — Audio player integration. B3063 — shipped. Inline native
<audio controls autoPlay>in each row. Clicking "Play" promotes the row to autoplay; another row's Play replaces. Pause sync via onPause/onEnded. - GENRE-13 — Best-line extraction. B3063 — shipped. Server-side join into
evaluations.transcendent_lines; surfaces the first line per song under the title with a Sparkles icon + italic quote treatment. - GENRE-14 — Sort controls. B3063 — shipped. 3-button pill row:
top score(default) /newest/most played. Toggle pauses any playing audio (list reorders, target row may scroll out). - GENRE-15 — Loading skeleton. B3063 — shipped. 5-row pulse skeleton with staggered animationDelay for initial fetch + sort changes. Errors degrade silently to "Radio temporarily unavailable" copy.
Tier 4 — B3064 (5 items shipped, full /forge batch-state integration deferred)
- GENRE-16 — Batch Genre-Catalog mode. B3064 — shipped as standalone batch picker. New
/genres/[slug]/promptspage +<PromptBatchPicker>client island. Operator picks N (1-10) prompts via per-card toggle; selected prompts pool into a sticky-bottom batch tray. The handoff to /forge is per-prompt deep-link (each "Open #N" button opens the forge in a fresh tab with?prompt=...&genre=...). Full integration into the existinguse-batch-state.tsstate machine deferred to a follow-up — would require batch-state to accept a multi-prompt URL contract; mechanical but tangled, not worth tying up this round. - GENRE-17 — Preview state with reroll/remove. B3064 — shipped. Each batch-tray row has Reroll + Remove icons. Reroll swaps the slot in-place using
getRandomCatalogPrompt(slug, { excludeIds })so the new pick respects the rest of the selection. Visual flash (ring-amber-400) on the rerolled card for 800ms. - GENRE-18 — Anti-dupe weighting. B3064 — shipped via excludeIds. The reroll passes the current selection's ids as the
excludeIdsset — the picker can never duplicate within a batch unless the operator manually adds a prompt twice (which the toggle blocks anyway). - GENRE-19 — `/genres/[slug]/prompts` browse page. B3064 — shipped. SSR'd
/genres/[slug]/promptsroute. Static-generated per arc. Renders the full catalog with substyle-filter pills (count per substyle), per-card Forge + Add-to-batch toggle, and the picker tray below. Sparse when catalog is at seed (3/arc); fills out after Haiku expansion. - GENRE-20 — Telemetry. B3064 — shipped. Fire-and-forget
posthog.capture('genre_prompt_clicked', { arc, prompt_id, substyle, exercise_axis, source })on Forge-link clicks.sourcedistinguishes card-click vs batch-tray click. Drops silently when PostHog isn't loaded yet — best-effort.
Commit footer for B3054–B3060: Moat: Distribution (presentation + funnel improvements aimed at converting browse traffic).
Genre Pages Audit Round 2 — B3065 (operator-reported "doesn't flow with the rest of the site")
Full findings doc: docs/GENRE-PAGES-AUDIT-B3065.md. After shipping the 11-build B3054–B3064 sequence, the operator photographed the country page rendering with cream/pink Forbidden Archive chips + cream/pink Load-Bearing Primitive box. The WAR Room audit confirmed: the B3058 per-arc accent system was authored with `dark:` Tailwind prefixes that DO NOT FIRE because `tailwind.config.ts` has no `darkMode:` setting (defaults to media, which gates on OS preference). Light-OS visitors see the LIGHT-mode classes win. The 144 light/dark class pairs in src/lib/genre-arc-theme.ts plus 9 inline sites across the genre pages are all affected. Smoking gun confirmed.
Tier P0 — B3070 (must — the dark-mode bug blocks correct rendering for ~half of visitors)
- GENRE-22 — Drop light-mode classes from all 144 accent slots. B3070 — shipped.
src/lib/genre-arc-theme.tsrewritten: 9 bundles × ~16 slots, every pairedbg-X-50 dark:bg-X-950/40collapsed tobg-X-950/40. 6 inline dark:-prefix sites onsrc/app/genres/[slug]/page.tsx(breadcrumb link + Haiku pill + Forbidden Archive disclosure chips) collapsed to dark-tone classes only. 8 inline sites onsrc/app/genres/page.tsx(hero pre-header + 4 stat numbers + 3 "Why this matters" section icons) collapsed via bulk replace. Newscripts/check-no-dark-prefix-in-genre-pages.tsratchet walkssrc/lib/genre-arc-theme.ts+src/app/genres/**+src/components/genres/**and fails the build on anydark:prefix reappearing. Currently green: 0 violations across 9 gated files.
Tier P1 — B3067 (brand parity)
- GENRE-23 — Replace solid-fill gradient blocks with site-canonical pattern. B3077 — shipped.
ArcAccentgained actaBorder: 'border-{color}-500/30'slot (all 9 arcs). 5 surfaces converted:/genres/[slug]CTA,/genres/[slug]/[substyle]CTA,/genres/audit-in-actionCTA (now per-arc instead of hardcoded violet),TryItSectionfeatured card,PromptBatchPickerselected-prompt pills. New pattern:rounded-2xl border ${accent.ctaBorder} bg-dark-900/60, heading usesaccent.boxLabelfor tint, body usestext-dark-200. CTA buttons route through<Button intent="forge">. Per-arc identity now lives on tinted border + heading; surface is site-canonical dark elevated card. 7-test invariant suite ratchets the new slot's shape (/^border-[a-z]+-500\/30$/). - GENRE-24 — Define `font-display` alias. B3077 — shipped.
tailwind.config.tsgaineddisplay: ['Sora', 'Inter', 'system-ui', 'sans-serif']matching the canonicalheadingfamily. Pre-3077, 58font-displayheadings across 19 files (admin/vault dashboards, /genres/* surfaces, components, homepage genre grid) silently inherited Inter because the alias was undefined. One-line config fix; affects rendering site-wide. - GENRE-25 — Standardize breadcrumb pattern across 3 genre routes. B3074 — shipped. All 3 sibling routes now use one idiom: segments separated by
/, current page as a non-link intext-dark-300, parent segments as accent-colored Link components,aria-label="Breadcrumb"on the nav element./genres/[slug]=All genres / {Genre}(2-segment);/genres/[slug]/[substyle]=All genres / {Genre} / {Substyle}(already 3-segment, unchanged);/genres/[slug]/prompts=All genres / {Genre} / Prompts(was a single back-link with arrow icon).
Tier P2 — B3068 (flow + density)
- GENRE-26 — Compress `/genres/[slug]` 8 sections → 5. B3075 — shipped (partial). Load-Bearing Primitive section folded into the hero as an inline accent-bordered chip; the standalone rounded-xl section is gone. GenreRadio moved from position 4 (after Substyles) to position 6 (just before Corpus) — proof cluster now adjacent. Corpus rounded-xl card flattened into a
border-t pt-6block with shortened button labels. Final order: Hero (with LB chip) → Try-It → Substyles → Audit Primitives (collapsed) → Forbidden Archive → Genre Radio → Corpus → CTA. The LB section absorbed entirely; the Corpus section's vertical real estate roughly halved. "Other substyles" deletion from the substyle deep page deferred to a follow-up. - *GENRE-27 — Focus-visible rings on custom buttons in `src/components/genres/
.** *B3074 — shipped.* Canonicalfocus-visible:outline-none focus-visible:ring-2 focus-visible:ring-cyan-400/70 focus-visible:ring-offset-2 focus-visible:ring-offset-dark-950` added to: TryItSection Shuffle button; GenreRadio 3 sort pills (top-score / newest / most-played); PromptBatchPicker substyle filter pills (× ~5-17 per arc) + per-card batch toggle + Reroll + Remove icon buttons. Keyboard users tabbing through the genre surfaces now see the canonical cyan ring that matches the rest of the site. - GENRE-28 — `motion-reduce:` modifiers. B3074 — shipped. Every animated/transition surface in
src/components/genres/*now honorsprefers-reduced-motion: TryItSection Shuffle button, GenreRadio sort pills, all PromptBatchPicker buttons getmotion-reduce:transition-none. PromptBatchPicker reroll flash (the 800msring-2 ring-amber-400/60) downgrades toring-1under reduced-motion — still a visible cue, lower intensity.
Tier P3 — Backlog
- GENRE-29 — Per-arc accent propagation (Option A: propagate). B3080 — shipped. Operator chose A: extend arc accents beyond
/genres/*. Three surfaces threaded:
1. `/lyrics/[genre]` — hub badge, "Why the profile is tuned" heading + check icons, leaderboard heading + score numbers all pick up the resolved arc accent. Hub-only slugs (hip-hop) translate to arc slugs (rap) via the new resolveArcSlug() helper. 2. Navbar — when the route matches /genres/X or /lyrics/X, a 2px absolute-positioned bar tinted in the arc's accent renders at the navbar's lower edge. Uses usePathname only (no useSearchParams) so global static-rendering stays intact. New navTopBar ArcAccent slot (bg-{color}-500/40) all 9 arcs. 3. `/forge?genre=X` — "Tuned for {Label}" chip with per-arc tint surfaces above the composing pad when the deep-link arrives with a recognized arc. Mirrors the persona-locked chip idiom; clicking "what this means →" deep-links to /genres/{arc}. Free-text genres still pass through unaccented (pre-3080 byte-identical behavior).
New library exports: resolveArcSlug(input) handles arc slugs, hub slugs (hip-hop → rap), labels (Hip-Hop, R&B), and punctuation-stripped fallbacks; getArcAccentForGenreInput(input) returns the bundle or null (NOT the violet default — lets the Navbar avoid misleading tint on non-genre routes). 7 new tests cover the resolver edge cases.
- GENRE-30 — Custom audio player for GenreRadio rows. B3077 — shipped. New inline
RadioAudioPlayercomponent replaces the native<audio controls autoPlay>UA chrome. Layout: per-arc-tinted play/pause toggle button (usesaccent.pillBg+accent.pillText), monospace tabular-nums elapsed/total time, native<input type=range>seek bar styled withaccent-current, surface wrapped inborder ${accent.ctaBorder} bg-dark-900/60. Honors canonical focus-visible ring + motion-reduce; autoplay-on-mount with graceful Safari-block fallback (button visible, click → play). SameonEndedcontract as the old native player. ~75 LOC inline (kept local to GenreRadio since no other surface needs it). - GENRE-31 — Reroll flash arc-accent threading. B3078 — shipped.
ArcAccentgained two new slots:flashRing: 'ring-{color}-400/60'+flashBorder: 'border-{color}-500/40'for all 9 arcs. PromptBatchPicker's 2 hardcodedring-amber-400/60usages (the prompt-card flash + the batch-tray-row flash) now resolve throughaccent.flashRing+accent.flashBorder. Country / worship / indie / etc. each flash in their own arc color on reroll instead of all 9 flashing with rap's amber. Test suite gained matching invariants (/^ring-[a-z]+-400\/60$/+/^border-[a-z]+-500\/40$/). - GENRE-32 — Sticky tray + Footer collision. B3078 — shipped.
/genres/[slug]/promptsmain element gainedpb-32(split from the symmetricpy-12 sm:py-16intopt-12 sm:pt-16 pb-32). The sticky batch tray inside PromptBatchPicker (sticky bottom-3 z-10) was colliding with the global Footer on short pages + at scroll bottom. ~128px of bottom-padding gives the tray clear room above the Footer. - GENRE-33 — Mobile filter-pill compression. B3078 — shipped. On a 375px viewport, 17-substyle Latin was filling 4-5 rows of pills because every pill ran at
px-3 py-1 text-xswith a· Ncount suffix. Post-3078: tighter mobile sizing (px-2 py-0.5 text-[10px]mobile,sm:px-3 sm:py-1 sm:text-xsdesktop) + count suffix hidden on mobile via<span className="hidden sm:inline">. Same 17 pills now fit in 2 rows on mobile + 1 row on desktop. Accessibility preserved viaaria-label="Filter to {s} ({n} prompts)"so screen readers still read the count.
Tier 1.5 — Operator-spotted IA mismatch (small standalone)
- GENRE-21 — Footer `/lyrics/` hubs don't match Navbar `/genres`. B3056 — shipped (Option B). The 8-link "Lyrics by genre" row was retired from the Product column. A single "Genres · 9 deep" link sits alongside Forge / Refine / Crucible at the top of the Product column, pointing at
/genres. The /lyrics/[slug] SEO long-tail surfaces remain reachable from /genres/[slug] corpus blocks (B3054 GENRE-5) + from the /lyrics index. Two competing IAs (one with "hip-hop" + missing latin, one with all 9) collapsed into one canonical surface.
Today's top 5 (Tier A first wave)
These are the foundational discipline items. Each is one focused day's work. All five = one work week. Each one ratchets a CI gate so regression is impossible.
- #1 — Kill the `explicit-any` allowlist down to ZERO. Every
: any, everyas any, every<any>becomes a precise type or typedunknown+ guard. Ratchet ceiling at 0; CI now fails on a single new occurrence. (B1178 partial 194 → 110, B1182 closeout 110 → 0. Total: 194 anys eliminated across the codebase. Manhattan + data-intelligence supabase params typed via SupabaseClient bulk pass; gauntlet/route.ts gained a ParsedGauntletOutput interface; episodic-trace gained EpisodicTraceDbRow; deep-consolidation, dreaming, plateau-detection all typed against EpisodicTrace. extractLearnings widened to `unknown` to remove launderer casts at call sites.) - #2 — Wire jsdom + write the FIRST real `renderHook` test. Add
jsdomdevDep, configure vitest with a*.behavior.test.tsglob in jsdom env. ConvertuseSongDetailEdittest from structural-source-text to real renderHook + act() + state assertions. Path opens for converting all extracted hooks. (B1174) - #3 — Stand up `src/lib/design-tokens.ts`. Semantic color names (
tokens.score.high/mid/low,tokens.status.error/warning/success). Migrate/s/[slug]score-color logic + ScoreBadge as proof. Lint rule: forbidden literal hex/rgb in JSX. (B1175 stand-up; B2098 milestone — ratchet hit ZERO. 414 → 478 → 0 across the B1175→B2098 arc. Every JSX hex/rgb/rgba in `src/app` + `src/components` now lives in `design-tokens.ts`. CI ceiling fixed at 0; any new literal fails immediately.) - #4 — Bundle-size CI gate flips from snapshot to blocking.
bundle-size-ceiling.jsonexists; not enforced. Writescripts/check-bundle-size.ts, gate CI. Shrink-or-stay rule. (B1176) - #5 — Golden-eval regression gate ships.
src/lib/golden-evals/exists but invisible. Pick 3 prompts (country/pop/hip-hop), forge once, snapshot lyrics + scoring output, commit. CI re-forges with same model+prompt+temp and asserts byte-identical output. Locks the engine against silent prompt drift. (B1177 — shipped as 4 deterministic system-prompt snapshots: evaluator, refine@50, gauntlet@high, gauntlet@low. The forge prompt is non-deterministic by design (Math.random feature gates) so the byte-identical gate runs against the eval/refine/gauntlet prompts where prompt-drift actually matters.)
Net by end of week 1: +2 CI gates, 1 new test pattern, 1 new design primitive, 194 fewer anys.
Tonight\u2019s run (B1571\u2014B1584, 2026-04-27)
A 14-build sprint triggered by one operator question ("How are we doing to be able to make an Italian opera piece or Gregorian Chant?"). What started as a capability lift exposed the deeper "capability without surface" bug class, triggered a Deep Audit, and produced:
- 2 new CI ratchets installed (design-canon at 662, capability-
surface at 0 gaps \u2014 both compounding-down).
- 8 new one-click Tradition presets on V2 forge (Operatic,
Sacred Chant, Delta Blues, Murder Ballad, Sea Shanty, Spiritual, Hymn, Spoken Word). Each pairs a dedicated ghost voice + DNA entry + pure-mode prompt path + unit-test contract.
- 6 new ghost voices in GHOST_PROFILES (verdi, hildegard,
crossroads, reaper, mariner, witness, hymnwright, speaker).
- 8 new genre DNA entries in GENRE_DNA_DATABASE
(opera, chant, murderballad, shanty, spiritual, hymn, spokenword + 4 surfaced from existing DNA: gospel/punk/ reggaeton/dancehall).
- Multilingual UI: language picker surfaced on V2 (was hidden
since B1420), enabling Latin chant + Italian opera one-click.
- Conversion CTA: songs-remaining chip on V2 forge converts
the dashboard-only counter into a forge-page upgrade prompt.
- First-time-user mode: zero-prior-songs detection reduces
decision overhead with a "Show all options" escape.
- Console-log ceiling: 19 \u2192 10 (\~96% cumulative reduction
since B1037 baseline of 251).
- Unchecked-index allowlist: 10 \u2192 8 (-20% this build).
- 2 new evergreen guides at /songwriting (chant + opera craft).
- 1 new blog post at /blog (the build-narrative).
The full list of builds + their punch-list mappings:
- B1571: opera + chant capability lift (Verdi + Hildegard ghosts)
- B1572: chant capability fixed (DNA + pure-mode prompt + V1 tiles)
- B1573: design-canon CI ratchet
- B1574: language picker on V2 (capability-surface gap closed)
- B1575: capability-surface ratchet (4 genre gaps closed)
- B1576: V2 Tradition presets (Sacred Chant + Operatic) + unit tests
- B1577: blog post on the 6-build chant/opera narrative
- B1578: songs-remaining chip on V2
- B1579: console-log ceiling 19 \u2192 10 (9 calls migrated)
- B1580: chant + opera evergreen guides at /songwriting
- B1581: first-time-user simplified mode for V2
- B1582: Tier-1 traditions batch 1 (Delta Blues + Murder Ballad + Sea Shanty)
- B1583: Tier-1 traditions batch 2 (Spiritual + Hymn + Spoken Word)
- B1584: unchecked-index allowlist 10 \u2192 8
The pattern worth keeping: every operator critique becomes either a fix OR a CI ratchet that prevents the bug class. Capability/ surface was the second ratchet installed this run; the first (design-canon) had been Quality Council #1 forward-looking item all week.
Tier A — Foundation (Linear parity)
- #1 — Kill explicit-any allowlist (B1178 partial 194→110, B1182 closeout 110→0. Ratchet at 0; one new any fails CI.)
- #2 — jsdom + first real renderHook test (B1174)
- #3 — Design tokens module (B1175)
- #4 — Bundle-size CI gate (B1176)
- #5 — Golden-evals CI gate (B1177)
- #6 — Kill the
unchecked-indexallowlist (B1597 CLOSEOUT: cumulative 166 → 0 (100%). B1595 cleaned components/CoverArt.tsx (15 \u2014 hexToRgb signature widened to absorb 12 + RGBA pixel-stride asserts for 3); B1596 cleaned lib/genre-profile.ts (15 \u2014 `!` on every GENRE_PROFILES static-key lookup since each key is guaranteed in the record); B1597 cleaned admin/lineage/page.tsx (30 \u2014 narrow-once locals in the Fruchterman-Reingold force-layout loops). Allowlist now empty. Tier-A item closed.) - #7 — Kill the
strict-tscallowlist (B1600 CLOSEOUT: cumulative 37 → 0 (100%). B1598 cleaned dashboard/page (4 dead-code orphans deleted: EvalDeleteButton + handleDeleteEval + upgrades pipeline). B1599 cleaned dashboard/SongDetail (10 violations: unused destructured hook returns + handleSunoExport orphan). B1600 cleaned forge/page (56 violations \u2014 ~40 dead imports from EXTRACT-1 factory pattern + 9 dead destructured hook fields + 2 orphan closures + cascading dead-setter cleanup). Allowlist now empty. Tier-A item closed.) - #8 — Kill the
console-logallowlist (B1586 CLOSEOUT: 10 → 0. Cumulative since B1037: 251 → 0. Tier-A target hit. Final 8 calls migrated to logger.info: forge/criticalSend (2 in create-stream-plumbing), mycelial-pathways, data-intelligence/orchestrator, manhattan/creative-debt, manhattan/skill-frontier, side-effects/cover-art, side-effects/focus-group. Two legitimate exclusions documented in the script: observability/logger.ts (the canonical sink) and developer/page.tsx (a JS code snippet inside a multi-line backtick template). The exclusions are commented inline in scripts/check-console-log.ts.) - #9 — Visual regression with Playwright
toHaveScreenshot()(every page, every breakpoint) (B1185: e2e/visual.spec.ts snapshots 8 public surfaces with dynamic-content masks (SongCounter, ShippingThisWeek, leaderboard, hero canvas). Per-platform baseline path (Linux is the CI source of truth). Tightened threshold 0.20 → 0.15. CI auto-uploads new baselines + diff PNGs as artifact for PR review. Single-breakpoint (1280x800) for now; mobile + tablet breakpoints land as follow-up when the desktop baseline is stable.) - #10 — A11y audit with
axe-corein CI, block on serious violations (B1179: e2e/a11y.spec.ts scans 10 public surfaces; block on serious+critical, advisory warnings on moderate/minor) - #11 — Storybook for every primitive (Button, Card, Badge, Input, ScoreBadge, etc.) (B1602 deferred; B2358 stand-up + B2372 closeout. B2358 installed Storybook 8.6 + the first 3 stories (Button 12, Chip 10, ScoreBadge 9). B2372 shipped the remaining 7 (Disclosure, Surface, MetadataRow, NextActionPair, TrustBlock, FooterLadder, CrucibleVoiceIcons) in a single Audit 2026-05-14 A16 batch — every design-system primitive now has a story file with realistic args + variant coverage. Item flipped to closed at B2421 (the punch list lagged the actual shipment by ~50 builds; B2421 caught the drift during a punch-list audit).)
- #12 — Real component library (replace ad-hoc Tailwind with
<Button intent="forge" size="md">) (B1602 kickoff: Button primitive shipped \u2014 typed `intent="forge|ai|outline"` + `size="sm|md|lg"`, polymorphic <button> / <Link> via href prop, loading state with spinner, leadingIcon/trailingIcon slots. First migration: dashboard BillingTab upgrade CTA (next/link \u2192 Button href=...). 11 contract tests assert intent/size class composition + canon-compliance (no off-palette utilities, no rounded-3xl, no inline gradients). Closes when 30+ btn- call sites are migrated to Button + the design-canon gate adds a forbid-rule for raw <button className="btn-">.) - #13 — Type-generated Supabase schema (
database.types.tsregenerated on migration; no hand-typed table shapes) (B1186 prep: workflow + npm script + placeholder + freshness check shipped. CLOSEOUT B1761: four-build chain landed the criteria — B1756 typed factory `createTypedServiceRoleClient()` scaffold + 6 contract tests; B1757 operator ran `npm run db:gen-types` + types regenerated against live schema (32 tables, 4 RPCs, schema 14.4); B1758 wave 1 migration (admin/regression-check + admin/costs, launderer ceiling 13 → 11); B1761 wave 2 migration (admin/forge-metrics + admin/cost-per-song, ceiling 11 → 8, including FULL/FALLBACK dynamic-select dead-code removal). Five admin routes now use the typed client; row types narrow against the live schema; the placeholder is gone; the freshness check would fail loudly if a migration drifted the types. Item closed.) - #14 — Zod runtime validation on every SSE event (no trusted casts) (B1180: PipelineEventSchema + CrucibleEventSchema; runtime validation in withSSEStream, forge create-stream-plumbing, and crucible route. 34 schema tests covering valid variants + drift catalog. Catches typo'd type discriminators, missing required fields, mixed envelope drift (text vs message). Critical events that fail validation get replaced with a generic error envelope so the client never sees a corrupt payload.)
- #15 — Discriminated-union state for ForgeResult (
Draft | Evaluated | Gauntleted, each with required fields) (B1181: ForgeSongState union + deriveSongState() pure derive function in src/app/forge/song-state.ts. useSongState() hook combines useForgeSessionState + useGauntletState through deriveSongState. 20 unit tests cover every transition path including the partial-Gauntleted guard. Source-of-truth state stays in the existing hooks; consumers migrate to the union as files are touched.)
Tier B — "Actually beat Linear"
- #16 — xstate forge state machine (every transition explicit, every state typed) (B1045 hand-rolled state-machine reducer + B1532-B1536 V2 cutover delivered the same outcome xstate would have: explicit typed states, impossible-state representations rejected at compile time, forbidden transitions throw at the call site. Hand-rolled saved a 15-40KB bundle hit + the second mental model. The B1405 scaffold doc remains as the migration plan IF a future requirement makes the runtime visualizer worth the dependency. As of B1664 the forge state surface is tested at src/app/forge/forge-state-machine.test.ts + use-forge-state-machine.test.ts. Goal achieved without xstate.)
- #17 — Forge
page.tsxunder 500 LOC (logic behind hooks/machines, page = pure composition) (B1626 V1 deletion brought page.tsx from ~1,909 LOC to 63 LOC. The page is now a Suspense + searchParams-keyed mount of <ForgeV2 />. All forge logic lives behind extracted hooks + the state machine. Item closed.) - #18 — Real-time presence for batch forging (multiple windows / collaborators see live progress)
- #19 — Optimistic UI everywhere (public toggle, delete, rename — all instant + reconcile) (public-toggle B894, delete B1222, share-link B1254, collection-assign B1268, B1278 closeout: rename + collection-bulk-move + cover-art-regen. handleUpdateSong is now optimistic across the board (every PATCH lands locally before the round-trip; per-song snapshot restores on failure). handleBulkMoveToCollection clears selection + applies move IMMEDIATELY then reverts only the rows whose PATCH failed. regenerateCoverArt closes the prompt dialog + clears the input on submit so the loading spinner takes the thumbnail unobstructed; on failure the dialog reopens with the prompt re-primed and the error inline.)
- #20 — Soft-delete with 30-day undo (no destructive action without recovery) (B1264 migration draft + B1275 API/UI wiring. DELETE /api/songs/[id] now soft-deletes (sets deleted_at + deleted_reason); falls back to hard-delete when the migration column is missing so the button never breaks. New POST /api/songs/[id]/restore + GET /api/songs/trash + /dashboard/trash page with one-click recover. Migration applies via supabase/migrations/add_soft_delete_to_songs.sql; trigger auto-fills the 30-day window. Reaper cron for past-window hard-delete is the follow-on.)
- #21 — Per-user audit log (B1220 surface, B1228 SQL draft, B1269 read endpoint with migration-pending fallback, B1328 wired writes from finalize-song / restore / share / DELETE/PATCH on /api/songs/[id]. CLOSEOUT: operator applied the migration; B1727 swapped /dashboard/activity from songs-only derivation to user_audit_events as the authoritative source, with songs-derived events filling in for entries the audit table doesn't cover (older songs, pre-B1313 forges).)
- #22 — Inngest async job queue for forge → score → gauntlet (currently inline) (B1407 scaffold + B1429 Phase 1 (cover-art) + B1550 Phase 2A (focus group) + B1726 Phase 2B CLOSEOUT (audio generation lift) — `forge/audio.requested` Inngest event + `generateAudioJob` handler in src/lib/queue/jobs.ts; src/app/api/songs/forge/post-forge-side-effects.ts dispatches the event before falling through to the inline path. 3-retry exponential backoff replaces the legacy fire-and-forget IIFE. Phase 2C (eval) + Phase 2D (gauntlet/superstyle) are UX-deferred indefinitely — they emit user-visible output during the SSE stream and breaking that apart for queue purity isn't worth the cost. Item flipped to closed at B2421 — the punch list lagged the actual shipment by ~700 builds; B2421 caught the drift during a punch-list audit. Status detail + Phase 2B work plan in docs/INNGEST-ASYNC-QUEUE-SCAFFOLD.md.)
- #23 — Resumable forges (SSE disconnect → next request picks up where it left off) (B2357 Phase 1 SHIPPED: `GET /api/songs/forge/resume?id=<songId>` returns an SSE stream that polls the song row every 1.5s and emits phase/result/error events matching the fresh-forge envelope. Auth + ownership checks; hard cap at 285s; resumed events carry `resumed: true`. Client helper at `src/app/forge/resume-forge.ts` (5 tests). Phase 2 — durable event log + true event-replay — open; would require per-forge event persistence so a resumed stream can replay from a last-seen event id. Phase 1 gives the user a working resume path; Phase 2 makes resumed sessions byte-identical to the original.)
- #24 — Public SDK on npm —
@songforgeai/client, used by external tools (v0.2.0 published B1214; v0.3.0 published B1859 with verifySeal + voiceFingerprint. B3194 closeout: v0.4.0 adds `crucible()` — the 8-voice adversarial critique via POST `/api/v1/crucible`. SDK now covers every public `/api/v1/endpoint surface. CrucibleVoiceResult + CrucibleVerdict + CrucibleRequest + CrucibleResponse interfaces; loose-string typing onvoiceId/verdict/banner/panelAgreementso future server-side additions don't break consumers. 4 new tests pin POST shape + invalid_lyrics error mapping + rate_limit_exceeded handling + traditionScope preservation. 18/18 SDK tests green; tsc clean. Operator runscd packages/sdk && npm publish` to push v0.4.0 to the registry when ready.)* - #25 — OpenAPI spec for
/api/v1/score+ auto-generated docs at/developer/api-reference(B988 — pre-existing.) - #26 — Distributed tracing with spans per pipeline phase (Honeycomb or DataDog) (B1226 Phase 1 — in-memory ring + tracing.ts; B1403 Phase 2 — forge SSE phases auto-instrumented via phase-logger bridge + /admin/trace UI with Gantt-style flame graphs. Vendor integration (Phase 3) deferred — in-process data informs the choice before commit.)
- #27 — Sentry integration (every uncaught error categorized, source maps) (B1301-B1326 + B1379 shipped: Sentry runtime FULLY live \u2014 SDK init in 3 config files, captureException at 10+ surfaces, structured tags (build/sha/env/runtime/userId-hash) auto-attached. B1601 attempted to add withSentryConfig wrapper for build-time source-map upload but used the v8 API on a v10 SDK \u2014 the third options arg was silently dropped, removing the graceful-degradation guards, which broke SSE streaming + auth flow in production. B1617 reverted. B1725 CLOSEOUT: re-enabled withSentryConfig with the v10-correct 2-arg signature + sourcemaps.disable defaulting to !SENTRY_AUTH_TOKEN. The wrapper degrades to a no-op for every build environment that doesn't have the operator-supplied SENTRY_AUTH_TOKEN, so dev + contributors + token-less CI runs are unaffected. Activates source-map upload immediately when the operator sets SENTRY_AUTH_TOKEN + SENTRY_ORG + SENTRY_PROJECT in Vercel env. The B1618 OpenTelemetry-hang root cause was fixed via force-dynamic on the three affected pages, so the wrapper no longer triggers a static-gen Supabase deadlock.)
- #28 — Per-route p99 latency dashboard at
/admin/perfwith regression alerts (B1204 — in-memory ring + /api/admin/perf snapshot + /admin/perf live table. /api/v1/score is the first wrapped route; expand by copying the recordLatency pattern.) - #29 — Feature flags via GrowthBook (every UX change ships behind a flag, A/B tested) (B1227 cohort-flag primitive + B1270 /admin/flags integration. Ready for vendor swap when GrowthBook decision is made; surface stays identical so consumer code doesn't change.)
- #30 — Lighthouse score ≥ 95 on every public route, gated in CI (B1237 ratchet ceiling at .github/lighthouse-ceiling.json + scripts/check-lighthouse.ts gate; B1259 operator-run snapshot script; B1304 flipped the GitHub Actions workflow from soft-warn to BLOCKING. Current per-category floors enforced in CI: performance ≥ 80, accessibility ≥ 95, best-practices ≥ 90, SEO ≥ 95. Three of four categories meet the literal "≥ 95" target; performance pegged at 80 because real-world cold-start performance on a 9-route Next.js app with images + Vercel SSR is the realistic ceiling — bumping to 95 would either fail every push or require routes to ship without any non-trivial assets. Closeout B1414. Audit annotation if the literal-95-everywhere interpretation matters: the perf floor lives in `.github/workflows/lighthouse.yml` line 123 and `.github/lighthouse-ceiling.json`; raise both in lockstep with measured-improvement evidence, never speculatively.)
Tier C — "Past Linear" (our unique edge)
- #31 — Public scoring SDK becomes the standard (third parties cite our npm package)
- #32 — Per-metric API endpoints —
/api/v1/metric/specificity(test individual signals) (B1210 — POST + GET on /api/v1/metric/[slug]. Same eval pipeline + auth + rate limits + reproducibility seal as /api/v1/score; projects out the requested metric.) - #33 — Scoring corpus opens — 1000+ human-scored lyrics, public, used to version the rubric
- #34 — Public model card at
/scoring/standard/model-card(prompts, temperatures, training rationale) (B1197 — pipeline tables + temperature rationale + reproducibility-fields table + 5 known limitations. Cross-linked from /scoring/standard.) - #35 — Quarterly rubric versioning (
v1.1,v1.2published with diff + migration notes) (B1211 partial: scaffold proven via v1.0.1 PATCH bump. Cadence policy formalized in RFC-0001. Closes when the first MINOR/MAJOR bump ships through the published process.) - #36 — Reproducibility seal (every score includes rubric version + model + temp; results reproducible from public docs) (B1199 — `seal: { rubricVersion, model, temperature, buildSha, build }` attached to every shapeScoreResponse output. Wires through both sync /api/v1/score and async/jobs routes.)
- #37 — Internal CLI —
sfai forge "prompt"from terminal,sfai score --file lyrics.txt(B1203 — bin/sfai.ts first cut. Verbs shipped: help / status / score / counts. `npm run sfai -- <verb>`.) - #38 — Crucible API (third parties hit the 8-voice critique programmatically) (B1248 — /api/v1/crucible synchronous JSON endpoint with Bearer auth, 60/hr per key, reproducibility seal.)
- #39 — OG card per metric page (
/scoring/metrics/voice/opengraph-imagerendered programmatically) (B1202 — generateImageMetadata fans out 12 PNGs; tier-colored accent, giant metric number, gradient name, definition + URL.) - #40 — Embedded score widget (third-party blogs/sites embed a "scored by SongForgeAI" card with verifiable hash) (B1209 — /api/v1/badge SVG endpoint + /developer/embed snippet generator. Verifiable hash deferred to a future RFC; v1 trusts URL.)
- #41 — Webhook outbound —
song.forged,song.scored,gauntlet.completedevents to user-supplied URLs (B1238 song.scored from /api/v1/score; B1274 song.forged from forge finalize-song + gauntlet.completed from PATCH /api/songs/[id]. Shared dispatch-helpers.loadSubscribersForUser plus per-event subscription filter on `subscribedTo`. All three cohort-gated on `webhook_dispatch_live` (default 0%); fire-and-forget delivery; failures captured to Sentry. Env-var subscriber list for v1; per-user table swap is a future build.) - #42 — Multi-language scoring (Spanish, French, Japanese rubric pages + scoring support) (B1404 prep: RFC-0009 opened in public comment through 2026-05-03. Pins seal-annotation contract + 3-phase methodology BEFORE implementation lands so the first non-English score carries honest disclosure. Phase 1 implementation (language param + UI dropdown) deferred to "Continue Multi-Language Phase 1" trigger phrase.)
- #43 — Voice fingerprint analysis (given a corpus of one writer's lyrics, score Voice consistency over time) (B1413 scaffold doc + B1422 Phase-1-part-1 pure function: computeVoiceConsistencyIndex shipped in src/lib/voice-fingerprint.ts. B1685 Phase-1 endpoint shipped at /api/v1/writer/me/voice-consistency. B1687 Phase-1 dashboard surface shipped as VoiceConsistencyCard mounted in /dashboard?tab=stats — loading / insufficient-samples / error / ok all rendered, behavior tests cover all four states. Phase 1 (a) dashboard surface = DONE. Item stays open until Phase 2: multi-axis VCI with per-song fingerprint vectors persisted (RFC-0009 deferred until calibration corpus reaches threshold).)
- #44 — Versioned songs UI (every revision on a timeline; restore any version) (B1255 gauntlet revision card + B1276 full timeline closeout. SongDetail.tsx now shows the complete version history (newest first) ungated from versionNumber > 1, so the feature is discoverable on day-one songs. Restore button snapshots the CURRENT state as a new version BEFORE overwriting, so any revert is itself reversible — round-trip safe. Empty state explains the contract; "Most recent" badge marks the freshest snapshot. Existing /api/songs/version POST + GET endpoints unchanged.)
- #45 — Public weekly engineering report at
/engineering(auto-generated from git log + changelog + telemetry) (B1201 — built from `git log` at request time. Highlights / ratchets / area heatmap / all commits. Force-dynamic; no telemetry plumbed in so the report is reproducible from outside.) - #51 — Hit-Song Calibration Corpus (Tier 3 of Brett review). Verified-hit lyrics scored against the rubric, used to detect calibration misses (rubric marks #1 hits as "Collapsed") and trigger Gravity-Rule recalibration. Closes when (a) corpus has ≥15 scored entries per genre, (b) at least one genre clears ≥90% pass-rate against its floor, and (c)
/scoring/standard/hitsships as a public artifact with the reproducibility seal. Trigger: Brett (Nashville pro) found the rubric scored a verified hit "Collapsed" — credibility blocker for any commercial writer.
- Infrastructure shipped (B2078–B2083): src/lib/calibration/hit-corpus.ts schema + helpers (findCalibrationMisses, summarizeCorpusHealth, CALIBRATION_FLOOR_BY_GENRE), 11 vitest tests, scripts/calibration-init.ts + scripts/calibration-report.ts + scripts/calibration-set-lyric.ts, private-corpus/ gitignored, docs/HIT-CORPUS-CURATION.md (policy) + docs/HIT-CORPUS-WALKTHROUGH.md (literal step-by-step). - Operator pause point (2026-05-05): operator started populating with "Last Night" by Morgan Wallen as the first country entry. Got blocked on JSON escaping, fixed via B2083 set-lyric helper. Recovery recipe is in chat history; the workflow is: save lyric to lyric.txt → npm run calibration:set-lyric -- <id> ./lyric.txt → score on songforgeai.com → record scores block manually → npm run calibration:report. - To resume: trigger phrase Continue Hit-Song Calibration should re-load docs/HIT-CORPUS-WALKTHROUGH.md and pick up at "score the first entry" step. Or operator follows the walkthrough independently and pings when ready to ship /scoring/standard/hits. - Definition of done: country pass-rate ≥90% AND /scoring/standard/hits page live AND reproducibility seal showing build number + scored-at timestamp on every entry. At that point the rubric publicly stands on chart-agreement, not curator opinion. Brett gets a magic link to that page for the week-10 re-review. - Why this is in Tier C (and not Tier D or a separate tier): it's our deepest standards moat. Linear has nothing equivalent because they don't ship a craft rubric. The closest peer move is the Lyric Scoring Standard whitepaper (B1093). Hit-Corpus is what makes the Standard credible to commercial writers, not just literary ones.
Tier D — "Why-would-Linear-do-this"
- #46 — Open the operating principles as a doc (how decisions get made, what gets prioritized, what gets cut) (B1200 — /about/principles. 10 principles, each paired with the inversion we reject. 4 decision gates. Linked from /about + footer Company column.)
- #47 — Public RFC process (every major change ships after a 7-day public RFC at
/rfc/<slug>) (B1208 — /rfc index + /rfc/[slug] detail + RFC-0001 (rubric versioning policy) opened in-comment through 2026-05-02.) - #48 — Reproducible deploys (
git revin footer matches what's running, every time, signed) (B1198 — src/lib/build-info.ts reads VERCEL_GIT_COMMIT_; footer renders "build N · shortSha" linked to GitHub commit; /api/version JSON endpoint for uptime checks.)* - #49 — Deterministic CI (same commit produces same artifact bytes, every time) (B1272 scaffold doc + B2099 Phase 1 + B2356 Phase 2 + Phase 4 infra shipped. `scripts/check-deterministic-build.ts` walks two `.next/` directories, hashes every artifact, categorizes diffs. `.github/workflows/deterministic-build-check.yml` runs weekly + on workflow_dispatch. B2356 added: SOURCE_DATE_EPOCH + TZ=UTC + LC_ALL=C.UTF-8 pinned in the build env (Phase 2); `--strict` flag + `.github/deterministic-build-allowlist.txt` allowlist loader + workflow_dispatch `strict` input for Phase 4 dry-runs. 10 new unit tests lock the allowlist parser + partitioner. Item stays open until the default-flip: strict-mode default goes from `false` to `true` and `push:` triggers land, gated on a clean Phase 2 measurement run.)
- #50 — Postmortems published (every incident gets a public writeup at
/incidents/<id>) (B1207 — /incidents index + /incidents/[slug] detail + first meta-postmortem entry (channel exists before first real incident).)
How items flip from open to shipped
When a build closes an item:
1. Replace [ ] with [x]. 2. Append a build reference: *(B1234, B1235)*. 3. If the item shipped partially, leave it open and add a sub-bullet describing what landed: - B1234: zod added to forge SSE events; refine + crucible still pending. 4. If the item is GHOST (already shipped before this list existed), mark [x] with *(pre-existing as of B<N>)* and a one-line note.
Every five items shipped: refresh the Tier A summary line at the top so the at-a-glance view stays honest.
Why this list and not the old "Known open items"
The old CLAUDE.md backlog accumulated audit notes from months of context. By Build 1160 it was carrying four shipped-ghost items. The Linear List is prospective discipline, not reactive tracking — every item is a deliberate move toward 100th-percentile, picked because it ratchets a real metric or unlocks a real capability.
The previous ### Known open items block in CLAUDE.md has been deprecated in favor of this file. New session prompts should Read docs/LINEAR-LIST.md first.
Known costs
- Tier A items 1-5 will each take ~1 day of focused work. The allowlist purges (#1, #6-#8) involve grinding through real code that's been allowed to drift.
- Tier B items typically span multiple builds — xstate migration (#16) is probably 5-10 builds done carefully.
- Tier C items have external dependencies (npm package, OpenAPI tooling) and should be planned as small projects.
- Tier D items are cultural commitments more than code commitments. They land when the team is ready to defend them publicly.
The point isn't to ship all 50. The point is never to coast. Every build either advances this list, or has a written reason it doesn't.