Agent Kanban Board

The living, shared Kanban canvas for ALL agents (and humans) working in this cross-repository project — the prompt-experiment loop in this repo (llm-entity-extraction) and the pipeline it feeds (llm-mailroom). It is the single place where tasks that still need to be completed, are in progress, are blocked, need to be revisited, or have been finished are tracked — with every card tied to a semantic-version release in CHANGELOG.md and, when critical, to a GitHub issue in the repo where the work lands.

This is a WORKING DOCUMENT, not documentation. It is meant to be modified progressively as work happens. Finished work is never deleted — it moves to the Archive (bottom of this file) so the full audit trail stays intact.

How to use this board — the procedures (READ FIRST, every session)

1. Read before you work

Read this board before starting ANY task — the Kanban table, the discussion log, and the open GitHub issues — for (a) cards/issues already claimed by another agent, (b) cards that cover the work you were about to do, and (c) context posts that affect your work.

2. Scope — one board for both repos

Cards can target either repository. The task summary states the repo or the card names it in the evidence column; llm-mailroom work (pipeline integration, synced prompts/logs) is tracked here with the other cards. Issues for cross-repo work open in the repo where the work lands.

3. Self-assignment (claiming a task)

A task is yours only after you claim it, in this order:

  1. Comment on the GitHub issue (if the card has one — the Issue column) saying you are claiming it, OR post to the Discussion board if the card is board-only. Discussion posts are appended at the TOP of MESSAGE_BOARD_DISCUSSION.qmd (date + agent + card + subject + full post — see the Discussion board section below).
  2. Move the card to in_progress and set Owner to your agent name
    • today’s date.
  3. Reference the card in your first commit: sorter_v7 (KANBAN-003): ... or MESSAGE BOARD: KANBAN-004 claimed.

Never work silently; never race a claimed card. ONE owner per card — if a card is claimed, build off it (offer help / pick an unclaimed card).

4. Work underway = in_progress, immediately — never backlog

backlog means ZERO work has started: no draft, no diff, no run in flight, no partially landed commit. The moment ANY work exists for a card’s scope — a first working-tree edit, an uncommitted prompt/test/ script draft, a branch, a run that has started — the card MUST be moved to in_progress (Owner set, Updated dated), NOT left in backlog.

  • Label it before the code, not after. The status move happens when work begins; never when it finishes.
  • Sanity check every session (and before every commit): if git status (or a branch, or a running eval) shows changes that belong to a card, that card must read in_progress in the table. If a card is found underway-but-labeled-backlog, move it to in_progress at that moment, set Owner to whoever holds the work, and post a dated note in the discussion log. Uncommitted work on the board = an in_progress card.
  • The table must never lie about reality: a card whose summary says “draft in the working tree” is, by definition, in_progress — update the lane at the same time you write the summary.

5. Work that relates to an existing task MUST update that task

If your work addresses the problem identified in a card — even partially, or from a different angle — you update THAT card: comment on the issue / post to the discussion board, move its status to reflect reality, and extend its summary with what you found. Never create a parallel card for a covered problem and never open a duplicate issue. Only add a NEW card (next free number, new issue if critical) when NO card covers the task.

6. Status moves — when and who

Move Who When
backlog → in_progress anyone (self-assign) You are actively working it, OR any work exists for it (draft, diff, branch, run started) — Owner set + dated, immediately, never deferred
→ blocked owner Stuck — post the blocker in Discussion + on the issue; name what unblocks it (data, keys, decision, card)
blocked → backlog/in_progress owner Blocker cleared — post what cleared it
→ in_review owner Work done; awaiting validation (tests, A/B, release gate). Link the evidence (run, commit, PR) in the card
in_review → done (Archive) reviewer/releaser Validation passed; CHANGELOG entry exists (same-commit rule); issue CLOSED in the same commit
done → backlog (reopen) anyone Regression / new data / superseded assumption — Discussion post explains why; issue reopened with the post

A card is not done until every criterion in §7 holds; a synced card’s done and its issue’s closed are the SAME event.

7. Completion — the done & issue-close criteria

A task is COMPLETE — and its dedicated GitHub issue CLOSABLE — only when EVERY requirement holds. All six must be true before you report done:

  1. Work verified — tests pass (network-free suite), the A/B run landed, and/or the release gate (release.py --check) is green. No uncommitted work may remain in the card’s scope — git status is clean for the card’s files (stray diffs mean the card is still in_progress, not done).
  2. CHANGELOG entry exists — the [Unreleased] (or released) entry describing the work lands in the SAME commit that ships it (AGENTS.md same-commit rule). No changelog entry = no done.
  3. Card archived — the card moved to the Archive with: shipped version, commit/tag, key result, and the Owner/Updated timestamp filled.
  4. Discussion closed out — a dated closing entry on the discussion board (result + verdict), newest-at-top, history never edited.
  5. Issue closed in the same commit — for synced cards, gh issue close NNN runs with a closing comment naming the commit and CHANGELOG entry, IN the same commit that archives the card. The card must never sit done with its issue open, nor the issue closed with the card unarchived.
  6. No orphaned scope — anything discovered but NOT delivered (a new confusion cluster, a follow-on arm) spawns its own card + issue BEFORE this card closes; unfinished discovery is not a silent done.

Completion close-out is the LAST action of every task: verify → changelog → archive → discussion post → close issue → report.

Reopen protocol (regression / new data / superseded assumption): move the card back to backlog AND reopen its issue with the Discussion post as the comment — in the same pass.

8. GitHub issue sync (critical / high-priority tasks)

Critical, high-priority, and cross-repo tasks are routed to GitHub issues (label kanban), opened in the repo where the work lands, so agents can open/close them like normal issues while the board remains the source of truth.

  • Which tasks route to issues: critical / high-priority cards, cross-repo work, anything whose completion must be externally verifiable, and any card the human asks to track as an issue. Board-only cards (small, single-session, low-risk) do NOT need issues.
  • The card ALWAYS carries the dedicated issue link. Every synced card’s Issue column holds the FULL markdown link to its own dedicated issue — [#NNN](https://github.com/Exios66/<repo>/issues/NNN) — never a bare number, never a link to another card’s issue. One card = one issue.
  • Order of operations when adding a critical card: open the issue FIRST (in the repo where the work lands), then write its link into the card: gh issue create --title "KANBAN-00N: <task>" --label kanban --body "<card summary + evidence + procedure>".
  • Cross-reference both ways. The issue body names its KANBAN-00N; the card names its issue link. They must never disagree about status — a lane move on the card is mirrored on the issue (claim → comment, blocked → blocker comment, done → closed).
  • When a card ships, close its issue in the same commit that archives the card (gh issue close NNN), with a closing comment naming the commit and CHANGELOG entry (see §7.5).
  • When a card is reopened, reopen its issue with the Discussion post as the comment.
  • Sync sweep: after any board edit, audit the table — every open card has a link in its Issue column, and every link points at an issue that is OPEN (gh issue list --label kanban), except cards in the Archive, whose issues are CLOSED. A missing link = an unsynced card = not ready for assignment.

9. Commit discipline

Reference cards in commits — MESSAGE BOARD: KANBAN-004 claimed or v24 diagnostic (KANBAN-004): .... A commit that lands a card’s work carries its CHANGELOG entry in the same commit (AGENTS.md rule) and closes its issue.

10. Release sweep

Semantic versioning is the spine: every open card names its target release. When a release ships (scripts/release.py --bump), sweep the board: cards that landed move to the Archive under that version (issues already closed), cards that did not land are re-targeted to the next release. The board and CHANGELOG.md must never disagree about what is done.

Status lanes

Lane Meaning Who can move it
backlog Todo — not yet started, NOTHING underway: no draft, no diff, no branch, no run in flight. The release it targets is set in the table. anyone (add/claim)
in_progress Work EXISTS and is being actively worked by the Owner — OR any work for the card exists at all (uncommitted draft, started run, partial commit). ONE owner per card. Cards with uncommitted work in the tree must be here, never backlog. owner (or any agent fixing a mislabeled card, with a dated discussion note)
blocked Stuck — waiting on data, keys, a decision, or another card. Post the blocker in Discussion. owner
in_review Work done, awaiting validation (tests, A/B, release gate) before it can land. owner → reviewer
done Finished and recorded in the Archive with CHANGELOG linkage. reviewer/releaser

Key Kanban table

Status codes: backlog · in_progress · blocked · in_review · done

Card Issue Status Task (summary) Owner Target release CHANGELOG / evidence
KANBAN-102 — (board-only) blocked Pin consumers to llm-dojo-scoring@v0.10.0 — entity PR #56 ready (v0.7.0→v0.10.0); mailroom + The-Mailroom pins prepared locally + patches at /opt/cursor/artifacts/*-pin-v0.10.0.patch. Blocker: CONSUMER_REPOS_GITHUB_TOKEN unset — cursor[bot] 403 on push to sibling repos. cloud-agent 2026-08-27 next release sweep entity #56; patches apply-clean on main; waiting on token or manual git am
KANBAN-101 #53 in_review Docclass pipeline parity bolster + stratified-120 A/B (entity + mailroom) — 54-key docclass registry (v1 specialists + sorter v7/vision v1); scoring parity (classify_failure, emit_docclass_run_scores); v5 export + run_langfuse_docclass_specialist_eval.py + pinned manifest docclass_ab120_s42_filenames.jsonl. cloud-agent 2026-08-26 next release sweep Parity shipped #54 (db10d6f); A/B 2026-08-26 (v5, n=120 sorter / 24 scored contracts / 24 insurance, seed 42): sorter v6→v7 exact 0.5833→0.6833 (+10pp), subclass 0.5917→0.6917, doc_type 0.9917 flat; contracts v0→v1 overall 0.6884→0.8444 (+15.6pp); insurance v0→v1 0.6957→0.6904 (flat). PR #55
KANBAN-098 llm-mailroom #17 done Lane B: arbiter-approved re-extraction unreachable (composed-path trap) — self-discovered defect 2026-08-24 during KANBAN-095 notebook work; demonstrated live in notebooks/03_review_lanes.ipynb. The approving arbiter sets arbiter_retry_count=1 (arbiter_node) while after_retry_extraction_gated demands < 1, so approved re-extractions escalate to humans instead of firing retry_extract; the second judge pass never runs. Both halves unit-pinned green separately; composition dead-ends. Fix shape: increment only after the retry fires, or accept <= 1 (still bounding to one retry); composed approve→re-extract→re-judge test pin required. Scope: llm-mailroom graph code + tests. ox-alpha 2026-08-24 next mailroom release sweep claimed 2026-08-24 at spawn (per §5(f), before KANBAN-095 closed); issue #17 filed with full evidence. FIXED & CLOSED 2026-08-24 — llm-mailroom cac4256 (pushed 9b119a6..cac4256): after_arbiter bound made approval-INCLUSIVE (< 1 → <= 1) with full rationale in the docstring — the approving node’s approval-time increment is load-bearing (the retrying extract node reads the count to weave the fix-list into its prompt), so the first approval now dispatches; a SECOND arbitration demanding another retry still finds the budget spent and escalates; one-retry-per-document bound unchanged. Lane-B unit pin updated; NEW composed-path pins src/tests/test_kanban098_arbiter_retry_path.py (3 network-free tests on the REAL graph: approve → re-extract → re-judge → compile → archive, no route-for-review span; plus the second-demand escalation); notebook 03 narrative + stored outputs regenerated to demonstrate the FIXED path end-to-end; CHANGELOG [Unreleased] ### Fixed entry same commit; suite 462 passed (459 prior + 3); #17 CLOSED with full-trail closing comment
KANBAN-097 #51 in_progress Docclass role-prompt mutation iteration testing + per-role eval tasks (agent bench suite) — human directive 2026-08-24: prompt mutation iteration testing across the KANBAN-090 docclass roster (7 specialists, sorter reviewer, judge trio, arbiter, boss, docclass sorter) PLUS proper eval tasks measuring scoring/performance/success. (1) RULE-4 REPAIR: absorbs the unclaimed uncommitted working-tree work found at session start — insurance_claims_specialist_v0 + InsuranceClaimsSpecialist, agents/pipeline_agents.py role wrappers, scripts/run_agent_bench.py (edge/judge-mutation/conflicts), scripts/gen_edge_cases.py (273 deterministic adversarial items w/ machine-checkable expectations), scripts/gt_workbench.py, data/gt/ (34+11+27 hand-GT rows, 400 insurance real-GT rows) — no card owned it; this card now does. (2) Every roster role gets an eval task: reviewer/sorter blind-classification scoring added to the bench; compact append-only records into reports/experiment_log.jsonl; --dry-run gate on money-spending modes; network-free test pins. (3) qwen3.7-flash baselines per role → failure-cluster diagnosis → surgical .replace() mutations (one lesson per version) → same-seed same-surface A/Bs with noise-floor honesty (GEPA discipline per the prompt-engineer protocol). Honest scope note: due_diligence/compliance/court_opinions specialists have NO local GT packets yet — their edge suites generate once GT lands (documented residue). ox-alpha 2026-08-24 next release sweep claimed 2026-08-24 at session start — issue #51 created; foundation verification done pre-claim (kanban090 pins green, bench scripts import clean, GT shapes verified); baselines + mutation A/Bs to follow in this session
KANBAN-099 — (board-only) backlog **INCIDENT: canonical reports/experiment_log.jsonl TRUNCATED (204 -> ~29 rows) sometime between KANBAN-094’s close (6127d5f, 2026-08-24 morning) and the ox-alpha session start (~17:00 same day) — discovered when build_site.py’s new orphan-pruner (KANBAN-094) tried to delete 162 tracked run files against the shrunken source; pruned tree RESTORED via git checkout before any commit. The JSONL is LOCAL-ONLY (gitignored since KANBAN-053), so git cannot restore it directly. RECONSTRUCTION PLAN (one documented backfill, the sanctioned exception): rebuild the missing rows from the TRACKED per-run site payloads docs/data/runs/{001..204}.json (+ mailroom mirror md for prose fields), preserve the 42 current rows incl. 9 new agent_bench records, verify count=204+new, regenerate experiment_log.md + site data from THE reconstructed source, add a network-free pin asserting runs/*.json count == jsonl row count so silent truncation can never pass again. Root cause unknown (candidate: a prior session’s render/backfill accident); audit shell history/other agents’ logs if available. unclaimed next release sweep opened 2026-08-24 by ox-alpha at KANBAN-097 close-out
KANBAN-100 — (board-only) backlog Durability roster completion + judge trio surfaces + insurance v2 lesson (KANBAN-097 residue) — † numbered KANBAN-100 after colliding with the concurrently-claimed Lane-B card that took KANBAN-098 (mailroom #17); KANBAN-085 renumbering precedent — spawned BEFORE 097 closed per the no-orphaned-scope rule. (1) GT packets for due_diligence / compliance_filing / court_opinions / merger_agreement specialists so their edge suites can generate (framework ready; data missing — their bench modes exit ‘no suite’ today); (2) judge-mutation surfaces for the completeness + classification judges (correctness done); (3) insurance_claims_specialist_v2 candidate lesson: claim_type guessing on partial views (residual true fabrications in v1 A/B); (4) scorer-band calibration: damages_description compositions of visible tokens (overlap >=0.4) flagged as fabrication — decide whether summary fields warrant a looser evidence band; (5) sorter/reviewer edge suite spans only the 4 classes with local GT (merger_agreement absent). unclaimed next release sweep opened 2026-08-24 at KANBAN-097 close-out; renumbered 098->100 same day on collision
KANBAN-093 #49 done The-Mailroom v0.2.0 official GitHub release + README collapsibles + UI/TUI screenshots — human directive 2026-08-24 (Discord): cut the official v0.2.0 release (changelog already declares [0.2.0] - 2026-08-23; accumulated [Unreleased] furnishing entries fold INTO that header per their tagging law — tag must match the changelog header exactly), annotated tag, push main+tags, gh release create v0.2.0, then mandated wiki/sync-wiki.sh. README gains <details> collapsibles around long operational blocks and a screenshots gallery. Screenshots = REAL pixels via their documented test seam (create_app(source=LangfuseSource(client=FakeClient(...))) + rich make_trace fixtures served by uvicorn; headless-Chrome captures of floor/metrics/review; genuine mailroom-tui --once pty frame) — no Langfuse creds exist on this machine so demo seeding is unavailable; fixture-seeded exactly like their test suite, never mockups. Images land in docs/screenshots/. Suite holds 84 passed before the release commit. hermes 2026-08-24 The-Mailroom v0.2.0 shipped 2026-08-24 — v0.2.0 released; tag 9949dbf peels to f451632 = main HEAD (ls-remote peel-verified); suite 99 passed on merged tree (incl. sibling’s 15 new tests); screenshots docs/screenshots/{floor,review,metrics,tui-console}.png (real stack, FakeClient seam); <details> ×3 + release badge; CHANGELOG folded into [0.2.0] - 2026-08-24, sibling’s post-declaration work preserved under fresh [Unreleased]; issue #49 closed + shipped. Open scope: wiki init blocked — .wiki.git 404s until first page exists in web UI (no API path); one click then rerun sync-wiki.sh. Incident: parallel agent’s checkout hotfix/review-queue-dispatch-key yanked shared-clone HEAD mid-task; recovered collision-free via pure ref surgery (push-by-SHA), their branch untouched
KANBAN-095 llm-mailroom #15 done Formal notebook suite plan for llm-mailroom — human directive 2026-08-24 (Signal): draft the full, previously-never-formalized plan for the llm-mailroom notebooks — Jupyter notebooks illustrating the different functionalities and agent interactions/dynamics of the full pipeline, including one that runs an example pipeline run through the agents showing the outputs and the role of each agent, under the notebook subdirectory. Deliverable: notebooks/PLAN.md plan of record — suite roster (00 anatomy → 08 observability), one shared pipeline_lab.py bench driving the REAL graph through the test suite’s network-free mock seam (FakeLangChainLLM + mocked OpenAI client), honesty labels, KANBAN-078-style guards (hostile-cwd headless exec), incremental build order. Plan-only card; implementation is follow-on steps under the same issue. hermes 2026-08-24 — claimed 2026-08-24 — issue #15 created + claim comment posted (5400443861); PLAN.md drafted from live recon (13-node graph, 15-agent taxonomy, conftest seam, DocumentState lanes). PLAN SHIPPED 2026-08-24 — llm-mailroom 7cc14da (notebooks/PLAN.md + README suite-plan section + CHANGELOG entry; suite 410 passed, baseline held); board mirror this repo 9645ec0 (governance files only, KANBAN-094 lane untouched). Implementation proceeds notebook-by-notebook under #15 per the plan’s build order (bench → 01 → 00 → 02/03 → 04/05 → 06/07/08). IMPLEMENTATION SHIPPED 2026-08-24 — CARD CLOSED — llm-mailroom 9b119a6 (pushed b39cc2d..9b119a6): all nine notebooks 00–08 with stored headlessly-reproducible outputs on the REAL graph; bench gains sequence-scripting + genuine provider-shaped flaky seam + idempotent LabSandbox.open() (double-open env leak found & fixed by the new guards); src/tests/test_notebook_suite.py enforces all four PLAN duties; README rescoped; CHANGELOG entry same commit; full suite 459 passed (410 baseline + 49 guards); issue #15 CLOSED with full-trail closing comment. Discovered scope spawned BEFORE close per §5(f): KANBAN-098 (llm-mailroom #17) — Lane B composed-path trap demonstrated live in notebook 03
KANBAN-094 — (board-only) done Single-source-of-truth repair: contracteval side-log merged into the canonical experiment log + derived-tree pruning (self-discovered defect 2026-08-24) — the Posit site pins (test_rendered_pages_committed, test_quarto_render_is_deterministic_and_clean) have been failing as ‘chronic baseline’ since the 2026-08-18 contracteval batch: 8 runs wrote per-run JSONs into docs/data/runs/ but their log rows landed in an untracked side file reports/experiment_log_contracteval.jsonl, breaking the reports/experiment_log.jsonl ⇒ build_site.py ⇒ docs/data/runs/{n:03d}.json derivation chain (195 log rows vs 203 run files; deep-link count pin 203≠195). One more side row (qwen3.7-flash_contracteval_v5) never produced a run file at all (run aborted on key-limit exhaustion). Fix: merge all 9 side rows into the canonical append-only JSONL (hazard-sanitized, chronological), regenerate experiment_log.md + site data from THE single source, teach build_site.py to prune orphaned run files so the tree is always exactly {001..N}, add network-free pins for the single-log law. Side file removed after verified absorption; v5 stays honestly absent until its rerun. hermes (2026-08-24) — TRUTH COMMITTED 2026-08-24 — 6127d5f: canonical JSONL 195→204 rows (side file deleted after absorption proof), experiment_log.md + site data + pre-render includes + rendered Posit pages regenerated from THE source ({001..204}, deep-links 204=204); build_site.py orphan-prune + –check file-drift detection; _pre-render.py re-execs into the repo venv when driven by bare python3; CHANGELOG [Unreleased] entry; NEW tests/test_kanban094_single_source_truth.py (5 network-free pins). SHIPPED 2026-08-24 — posit verification suite 14 passed / exit 0 (test_rendered_pages_committed, test_quarto_render_is_deterministic_and_clean, + 5 new pins all green): the chronic pair is closed at root, derived tree provably self-healing
KANBAN-096 #50 done Modal-hosted vLLM serving capability for llm-entity-extraction (cross-repo parity with mailroom KANBAN-064) — human directive 2026-08-24 (Discord): integrate + utilize Modal+vLLM across BOTH pipelines INCLUDING ALL CONFIG FILES + LOCAL INFRA. Mailroom side already ships complete (064 / v0.4.1: deploy/modal_vllm.py + [deploy] extra + runbook + 12 tests; this card verifies it under mailroom #16). Entity scope: deploy/modal_vllm.py sibling Modal app (same env-knob contract: MODEL/GPU/QUANTIZATION/MAX_MODEL_LEN/API_TOKEN/HF_TOKEN, persistent HF-cache volume); provider seam DOCUMENTED end-to-end through the existing env-overridable OPENROUTER_BASE_URL chokepoints (agents/base_agent.py::llm() LangChain eval path + src/openrouter_utils.py raw path incl. the vision classifier) — shared VLLM_BASE_URL/VLLM_API_KEY contract so ONE deployment backs BOTH repos; [deploy] extra + requirements/deploy.txt honoring the KANBAN-081 dependency-manifest law; config/environments/.env.example deploy blocks; deploy/README.md runbook (setup/deploy/smoke/flip/teardown/cost) + smoke script; network-free guard suite (stubbed-modal app load, pure argv builder, manifest parity, runtime-tree-stays-deploy-clean census, env-seam pins). Non-goals: switching any default serving path (OpenRouter stays primary), live GPU deploys from CI. hermes (2026-08-24) — issues #50 + mailroom #16; CHANGELOG [Unreleased] both repos. PROOF 2026-08-24: capability suite 21/21 network-free green (dotenv-regression pin included); entity suite 685 passed / 3 failed / 4 skipped — failures = documented kanban076 hub-sha chronic + PRE-EXISTING docclass drift (39b3d5a Aug-18 added attorney_demand to taxonomy without listing it in sorter_docclass_v0; zero files of this card touched); mailroom verify-only pass 12/12 + issue #16 CLOSED verified (c31bfd5 pushed there); seam repaired to call-time resolution across base_agent/llm_chain/classifier
KANBAN-092 #48 done Overhaul The-Mailroom root README — human directive 2026-08-24 (Discord): the visualizer’s README is bland (102 lines, zero badges, near-zero sister-repo references) despite the repo being a fully governed family member. Surfaces: The-Mailroom README.md overhaul (factual static badge row ONLY — no release/license/CI badges: no v0.2.0 tag/release exists remotely, repo has neither LICENSE nor workflows, per KANBAN-089 honesty precedent; governed-constellation family section with YOU-ARE-HERE marker covering llm-mailroom upstream / llm-entity-extraction prompt loop / llm-dojo-scoring engine / corpus feeds / atticus-investigation eval sibling / graph sites; new trace-contract & schema-mirror-duty section deferring to AGENTS.md as authority; structure/dividers; ALL existing operational content kept), same-commit CHANGELOG [Unreleased] entry per their mandatory release law, wiki/Home.md family paragraph (mirror spirit; published by their release train). Docs-only; suite holds pre-edit baseline 84 passed (captured on this machine via out-of-tree venv per their no-venv-in-repo rule). hermes 2026-08-24 — (changelog-only, ships in Exios66/The-Mailroom) claimed 2026-08-24 — issue #48 created + claim comment posted (comment 5391489974); recon complete against cloned tree @ 4e53830: house rules read in full, badge honesty gaps verified, suite baseline green. SHIPPED 2026-08-24 (The-Mailroom 435bb04, pushed 4e53830..435bb04, closes #48): README rebuilt 102 → 177 lines — factual static badge row (version 0.2.0 / python 3.11+ / data source: Langfuse only; NO release/license/CI badges — none exist to honestly badge), “The governed constellation” section (YOU-ARE-HERE ASCII diagram + seven-repo at-a-glance table linking llm-mailroom’s canonical sister-repos map), new “The trace contract & the mirror duty” section (same-window mirror rule + visible-by-design breakage map, AGENTS.md as authority), [!IMPORTANT] schema-cache restart alert, dividers, honest “No license published yet” close; every operational byte preserved (17 load-bearing strings verified present); wiki/Home.md constellation paragraph added (published by their release train per their law). Same-commit CHANGELOG [Unreleased] entry. Proof: fences balanced (10), single H1, tables clean, internal link to AGENTS.md resolves, suite 84 passed before and after (0.34s) — docs-only, zero behavior change
KANBAN-091 #47 done Add The-Mailroom visualizer to the umbrella docs — human directive 2026-08-23 (Discord): The-Mailroom (pixel-art Langfuse-driven visual engine, v0.2.0) appears in ZERO umbrella surfaces despite being a fully governed family member (own AGENTS.md with schema-mirror sync duty back to llm-mailroom, own semver release train, own wiki). Surfaces: mailroom docs/sister-repos.md (constellation diagram + at-a-glance row + dedicated section), mailroom README umbrella row, mailroom wiki Home, entity docs/sister-repos.md, entity README prose, BOTH CHANGELOGs (one card / one issue / both changelogs, KANBAN-061 precedent). Docs-only; suites hold baselines. hermes 2026-08-23 — (changelog-only) claimed 2026-08-23 — issue #47 created + claim comment posted; recon done against the cloned The-Mailroom tree. SHIPPED 2026-08-24 (closes #47): entity 76567a5 + llm-mailroom 5cf8b60 — all six surfaces live: mailroom sister-repos (constellation diagram node, At-a-glance ‘Downstream visualizer’ row, dedicated ‘The-Mailroom — the visual engine’ section incl. the schema-mirror duty), mailroom README Umbrella row, mailroom wiki Home paragraph (pushed to the live wiki d441b74), entity sister-repos (diagram visualizer: line + downstream table row), entity README working-surfaces prose, both CHANGELOGs [Unreleased]. Docs-only; suites hold baselines (mailroom 405 passed; entity 654 passed / chronic-only failures).
KANBAN-090 #46 done Dedicated docclass prompt variants for every classification-chain role — human directive 2026-08-23 (Discord): developed + deployed docclass prompts for all 7 specialists (contracts, corporate_records, due_diligence, correspondence, compliance, court_opinions, insurance_claims), the sorter reviewer, the judge trio (completeness/classification/correctness), the arbiter, the boss, and the docclass sorter — each under a SEPARATE prompt module in BOTH repos: entity src/prompts_docclass.py (re-exporting the existing sorter_docclass v0–v6 family byte-identical + derived specialist/boss/judge variants off real bases) and mailroom src/langchain_agents/prompts_docclass.py (mirrored from local bases); wired for deployment through the Langfuse sync seam under distinct mailroom-*-docclass names; runtime defaults unchanged (additive-only doctrine). hermes 2026-08-23 — (changelog-only) issue #46 created + claim comment posted 2026-08-23; card claimed same-session per Phase 0/1. SHIPPED 2026-08-23 (closes #46): entity 3df3747 — src/prompts_docclass.py with 21-key DOCCLASS_PROMPT_VERSIONS (8 sorter-docclass re-exports byte-identical + 10 single-anchor .replace() derivatives of real base constants + 3 authored-fresh V0s: reviewer / arbiter / insurance_claims, provenance-commented), merged into PROMPT_VERSIONS at the prompts_archive-style tail (versions 103 → 116, zero collisions asserted at import); mailroom 79b0126 — src/langchain_agents/prompts_docclass.py, 13 variants, every one a PURE APPEND (variant.startswith(base) in full) of mailroom’s own production bases incl. house reviewer/arbiter/insurance constants; deployment seam: entity registration-IS-deployment via the eval Langfuse sync (all 21 keys mirrored like every family), mailroom OPT-IN ONLY — new scripts/sync_prompts.py --docclass pushes namespaced mailroom-docclass-<key> prompts, the thirteen agent-pinned production templates provably untouched (count pin held, negative assertions on every template). Shared DOCCLASS ARM CONTEXT block byte-compatible across repos (extended 8-class set incl. insurance_claim+merger_agreement; doc_subclass dims: CUAD contract subtypes / merger consideration types / title-derived record types); per-role rules: routing label = state not ground truth, claim-documentation + M&A leakage read-through, judge trio subclass-specific support requirements + cross-family leakage checks, classification judge grades the extended set with family discriminators, boss routes classification-fault conflicts to human review. Runtime defaults unchanged both sides (opt-in by key only). Guards: entity 6 network-free tests (tests/test_kanban090_docclass_prompts.py), mailroom 4 (src/tests/test_kanban090_docclass_prompts.py). Proof: targeted prompt suites green (entity 70 = 64 existing + 6 new; mailroom 9 = 5 existing + 4 new); FULL suites: mailroom 405 passed (20.5s, zero regressions vs 401 baseline), entity 654 passed / 2 failed / 4 skipped with the 2 failures proven pre-existing by stash-to-clean-HEAD replay (the documented chronic derived-site posit pair) and kanban076 hub-sha deselected (needs live localhost:6006). CHANGELOG [Unreleased] entries in the same commits.
KANBAN-089 #45 done llm-mailroom root README polish retrofit — human directive 2026-08-23 (Discord): aesthetic/enhancement pass tying the README together cleanly on top of the shipped KANBAN-082 docile skeleton — <details> toggle sections for long operational blocks, an at-a-glance facts table under the badge row, GitHub alert callouts for genuinely important operational warnings, honest badge cleanup (no license/CI badges — repo has neither), visual-rhythm/divider pass. Docs-only; zero behavior change; suite must hold the 401-passed baseline. Every number re-derived from the tree before writing. hermes 2026-08-23 — (changelog-only) claimed 2026-08-23 — issue #45 created + claim comment posted (comment 5390723579); work ships in llm-mailroom per cross-repo convention. SHIPPED 2026-08-23 (mailroom f9da346): at-a-glance facts table under the badge row; green release badge v0.4.1 (pyproject+tag matched); GitHub alert callouts [!NOTE] (zero-DB quick-start) + [!IMPORTANT] (taxonomy-cache restart gotcha); three <details> toggles (config cookbook, Ollama shortlist, deployment runbook); dividers between the four major parts. Honesty rules held: NO license/CI badges (repo has neither), license row states “not yet published”. Pre-push proof: fences balanced (38), all 33 internal anchors resolve under GitHub slug math, summaries single-line closed, at-a-glance table 2 cells/row, suite 401 passed (20.2s) — docs-only, zero behavior change; CHANGELOG [Unreleased] entry in the same commit
KANBAN-088 #44 done Family-wide JSONL line-boundary hazard guard — sweep remaining ensure_ascii=False writers (carve-out from KANBAN-087 residue): Hub-bound/line-oriented writers (enron publishers, legalbench pack/streamers, docclass builders merged/v5/pilot) adopt export_bt_to_hf.sanitize_line_boundary_chars; purely-local intermediates get documented exemptions; pins extended network-free. Published artifacts verified hazard-free in the 2026-08-24 staging census — prevention, not incident response. hermes 2026-08-24 — (changelog-only) issue #44 created + claim comment posted 2026-08-24; queued behind in-flight family work. SHIPPED 2026-08-24 (closes #44, entity 95d8bf0): census 15 sites / 11 files; canonical scripts/datasets/_jsonl_safety.py (verbatim extraction from the exporter) + safe_jsonl_line; exporter delegates via object-identical re-exports (kanban087 pins hold unchanged); 9 row-writer sites across 7 files adopted (backfill_extraction_kpis, build_docclass_merged, build_legalbench_full_pack ×2, publish_enron_correspondence(+dedup), stream_legalbench_tasks_to_bt ×2, build_docclass_v5); 5 sites exempted with justification markers (3 field-value dumps, 2 CSV-cell flattens, braintrust_utils hash input); guards tests/test_kanban088_jsonl_safety_sweep.py (5 network-free incl. repo-wide no-unmarked-hazard-sites scan). Suite 668 passed (663 + 5), failures byte-matching documented baseline.
KANBAN-084 #43 done Stratified pilot sample of docclass-merged — full type×subtype coverage — human directive 2026-08-23 (Discord): hyper-tailored, cleanly distributed pilot dataset derived from Lucius-Morningstar/docclass-merged covering every doc type and every subtype present in the parent, published to HF as its own family repo; feedstock for The Mailroom pipeline visualizer debugging + per-agent pilot eval. hermes 2026-08-23 — (changelog-only) claimed 2026-08-23 — issue #43 created + claim comment posted; recon: parent = 810 GT rows / 4 doc_types / 46 subclass strata (stratum sizes 1–57); honest gap declared (parent holds 4 of mailroom’s 7 classes — coverage = all types/subtypes present in source); plan: deterministic stratified draw, parent split preservation, two-config blind/GT publish per KANBAN-079 doctrine. SHIPPED 2026-08-23 (closes #43): docclass-pilot live — 138 rows / 48 strata (quota=3), blind+GT configs, deterministic sha256-within-stratum draw, datasets-server green; parent evolved to schema v5 (1,210 rows, +400 insurance_claim w/ InsuranceClaimExtraction GT contract; clause-level answer keys ground_truth-only: CUAD 509/509 contracts with 13,753/13,753 spans verified at exact char offsets, MAUD 152/152 mergers; subclass canon 28→26 contract subclasses killing the Affiliate/Endorsement duplicate-bucket skew, distinct buckets pinned unmerged); 10 network-free pins (test_kanban084_pilot_sample.py), suite 613 passed with residuals proven pre-existing on a pristine detached-HEAD worktree baseline; scope addenda (claims class + masterlabels-derived clause GT + subclass normalization) directed by Jack in-band same-day and folded into this card.
KANBAN-083 #41 done Root folder consolidation — nest the content/governance dirs — human directive 2026-08-23: finish the KANBAN-080 job; 13 visible root dirs overwhelm newcomers. hermes 2026-08-23 changelog-only SHIPPED 2026-08-23 (49a66c4, closes #41): root dirs 13 -> 10 — board/+discussion/ -> governance/, wiki/ -> docs/wiki/ (matches mailroom’s convention), site/ -> docs/posit-src/. Unmoved by design: reports/, data/, package dirs (25–49 live refs each). Lockstep census-driven updates: _quarto.yml output-dir deepened (../../docs/posit), _pre-render.py ROOT depth + governance reads, build_site.py, render_message_board_qmd.py, test_posit_site.py (SITE_DIR, div-balance read, git pathspec, yml-contract re-pin), README layout block (+ stale root-memos row from 080 removed), docs/README, wiki mirror pages, .gitignore rules. Portal render PROVEN from new home; wiki synced via docs/wiki/sync-wiki.sh (123efda). Baseline discipline: pristine-HEAD worktree baseline 617p/9f captured first; shipped tree 633p effective / 2f — both survivors are the documented chronic classes (explorer deep-link drift 195-vs-203, kanban076 hub-sha data check); ZERO move-caused failures. Residue RESOLVED same-day: AGENTS.md path-doctrine rows (9 lines) APPROVED by Jack in-band and landed post-consent; sibling’s live .qmd insert mid-write caused a tag-splice, repaired line-level with assertions before commit.
KANBAN-087 — (board-only) done Hub dataset broken: mailroom-cuad-contracts-full DatasetGenerationError (human report 2026-08-23) — ROOT CAUSE: JSONL was structurally valid (510/510 parse) but ONE record (line 73) carried 16 literal U+2028 LINE SEPARATOR chars inside .input.doc_text; Hub worker parses batches via str.splitlines() which treats U+2028 as a record break INSIDE the row and shreds it into invalid fragments (ujson_loads ValueError), while local datasets 5.x splits BYTES and loads happily — version-dependent landmine. SHIPPED: exporter guard sanitize_line_boundary_chars() (U+2028/U+2029/NEL escaped losslessly at write time); staging sha-matched Hub bytes (cac0c845…), sanitized with the exporter’s own function, per-row semantic round-trip asserted (0 mismatches), worker-shape A/B proven (526 shredded pieces → 510 intact); republished under canonical names overwriting broken blobs (after evicting misnamed _repaired duplicates from the first upload attempt); verified tree census=4 files, fresh-download round-trip sha=c8beefd6…, manifest consistent, datasets-server pending: [] / failed: []; 6 network-free pins in tests/test_kanban087_jsonl_hazards.py. hermes 2026-08-23 — (changelog-only) FIXED + republished clean 2026-08-24; commit lands exporter guard + pins + CHANGELOG [Unreleased] bullet; honest residue: 12+ sibling ensure_ascii=False writers in scripts/datasets lack the same guard — carve-out for a future sweep card if the family grows another Hub-bound writer
KANBAN-086 — (board-only) done Cold-suite interpreter pin: posit pre-render spawned bare python3 — both spawn sites in tests/test_posit_site.py (_write_include + quarto-determinism test) invoked _pre-render.py via bare "python3", so any suite run without the repo venv on PATH died in-subprocess with ModuleNotFoundError: llm_dojo_scoring → 5 phantom failures masquerading as content regressions (rode along in KANBAN-080’s documented baseline; KANBAN-081 recorded them as “7 chronic posit renders”). Fix: sys.executable at both sites (inherits pytest’s interpreter). A/B proof in identical stripped-PATH cold env (PATH=/usr/bin:/bin:/usr/sbin:/usr/local/bin, PYTHONPATH scrubbed): unpatched 7 failed / 2 passed → patched 7 passed / 2 failed, residual pair = documented derived-site chronic class, unrelated to interpreter resolution. Test-only. hermes 2026-08-23 — (changelog-only) SHIPPED same-day: commit lands tests/test_posit_site.py + CHANGELOG [Unreleased] entry (same-commit rule); board-only card per §8 (small/single-session/test-only).
KANBAN-085 #42 done validate_pipeline.py FIXTURE_EXPECTATIONS keys never match — intrinsic fixture expectations silently dead (carve-out from KANBAN-082) — rel = path.relative_to(REPO_ROOT) compares against repo-root-relative paths but keys say tests/fixtures/… / examples/sources/… while reality is src/tests/fixtures/… / docs/examples/sources/…, so every standalone fixture’s intrinsic (doc_class, subtype) expectation silently skips. Fix: correct key prefixes, add network-free regression test asserting every key resolves/matches ≥1 file, run validate_pipeline.py end-to-end to confirm expectations engage. Functional code — carved out of the docs-only KANBAN-082 ship. Numbering note: opened as KANBAN-083, renumbered twice after colliding with concurrently-shipped cards (#41 → KANBAN-083, #43 → KANBAN-084). hermes 2026-08-24 next release sweep claimed 2026-08-24 by hermes per human sweep directive (issue comment 5391361131). SHIPPED 2026-08-24 (closes #42): llm-mailroom 1401f95 — all 14 FIXTURE_EXPECTATIONS keys corrected to repo-root-relative truth (src/tests/fixtures/…, docs/examples/sources/…); NEW 5-guard network-free regression suite src/tests/test_kanban085_fixture_expectations.py (no stale prefixes / literal keys on disk / globs match ≥1 fixture / matcher provably engages per entry / registry pinned at 14); e2e proof validate_pipeline.py --fixtures --sources: expectations engage, per-class accuracy populated across all 7 classes, 21/22 matched-expected. Honest residue flagged: evidence-based mock classifies insurance_claim/sample_claim.txt as contract/other (0/1) — mock-evidence limitation, invisible while the map was dead, documented in CHANGELOG + issue. Suite 410 passed (405 + 5).
KANBAN-082 #40 done llm-mailroom modern README (docile-style) + docmd local doc-renderer integration — human directive 2026-08-23: clean, modern root README with badges, inline images, embedded links to associated repositories, and agent-organization/architecture maps (mermaid), design modeled on rossumai/docile; plus docmd integrated into llm-mailroom as the local markdown doc renderer over docs/. Docs/tooling-only; code ships in llm-mailroom. hermes 2026-08-23 — (changelog-only) SHIPPED 2026-08-23 — llm-mailroom commit 5030d7d pushed (5 files, +169/−65): README rebuilt docile-style — pixel-art banner docs/assets/banner.png (owl-mailroom masthead; gibberish AI-label text erased via scripted pixel surgery + vision QA rounds), badge row, repo-consists-of list, grouped TOC, NEW agent-organization mermaid map (15 agents / 7 doc classes read fresh from taxonomy.yaml; old diagram’s “6 specialists” corrected); API truth-fixed (phantom POST /ops/pause removed, real GET /queue documented, bearer-token guard stated per-route) with the same fixes + corrected “no auth” claim in src/api/README.md; Quick Start fixture path corrected to src/tests/fixtures/...; README.md#installing anchor added (repairs 4 vendored openrouter-* skill links); umbrella table linking the repo constellation (sourced from docs/sister-repos.md). docmd (@docmd/core v0.9.4, Node ≥20) integrated as zero-config local docs renderer: npx @docmd/core dev/build documented in README §Browsing the Docs Locally, site/ gitignored, local build proof 27 pages in ~4s (sidebar nav, offline search, llms.txt). Suite: 401 passed = shipped KANBAN-078 baseline byte-identical (docs-only change; PYTHONPATH-scrubbed run). Integrity gates green: all 26 TOC anchors resolve, all 12 relative link/image targets exist, mermaid/fence balance, repo-wide stale-ref sweep clean (frozen audit-report prose left as history). Mailroom CHANGELOG [Unreleased] entry in same commit. Honest residue carved out as KANBAN-085/#42: validate_pipeline.py FIXTURE_EXPECTATIONS keys never match — intrinsic fixture expectations silently dead. Issue #40 closed manually post-archive (cross-repo closes-#N no-op).
KANBAN-081 #39 done Modular dependency batches — evidence-derived install profiles — human directive 2026-08-23: split deps into purposeful installable batches so users tailor their footprint instead of installing everything. hermes 2026-08-23 — (changelog-only) SHIPPED 2026-08-23: core floor frozen at exactly 8 packages (pyproject dependencies ≡ root requirements.txt, parity-tested); extras [tracing] [evals] [datasets] [reporting] [embeddings] [dev] [all] + legacy [pdf] alias mirror requirements/<batch>.txt 1:1. Defects fixed: tracing stack was MISSING from pyproject entirely (bare pip install -e . could not import the default sink src/tracing.py); openpyxl + huggingface_hub were imported-but-undeclared; pandas/pyarrow dead pins removed (AST census: zero imports, no notebooks exist). Guard: NEW tests/test_dependency_manifests.py — 6 network-free pins incl. live AST census of agents/+src/ against a module→batch owner map (batch modules may use core ∪ their batch only). Proof-before-push: fresh venv (python3.13 / pip 26.1.2) core-only install from clean tree copy → src.prompts (103 prompt versions) + agents.sorter_agent import GREEN with phoenix/langfuse/opentelemetry/braintrust/huggingface-hub/sentence-transformers VERIFIABLY ABSENT; -e ".[tracing]" add-on → src.tracing imports and initializes the live Phoenix tracer. Suite delta vs pristine-HEAD worktree baseline: 611→627 passed (+6 new pins; FAILED set byte-identical: 7 chronic posit renders + kanban076 hub-sha check). Honest residue: matplotlib/openpyxl still reach core installs transitively — llm-dojo-scoring v0.7.0 declares them in its own install_requires (upstream dojo slim-down owed); AGENTS.md quickstart dep lines approved-and-landed same day (in-band Jack approval).
KANBAN-080 #38 done Repo housekeeping: docs currency + root de-clutter/nesting + graphify graph & links + v0.20.0 proof wrap-up — human directive 2026-08-23: (1) fresh-install archive proof re-run on pushed tag v0.20.0 (forensics 2026-08-23: langchain-core 1.x IS live on default PyPI 1.0.0→1.6.0 — the interrupted train’s “unresolvable floor” was a stale index view); (2) nest loose root files with ALL references updated in lockstep (SCORING.md→docs/SCORING.md, deploy_phoenix.sh→scripts/deploy/, root memos/ merged into docs/memos/ w/ build_site.py reader updated, legacy root board symlinks removed — all programmatic readers already canonical); (3) graphify knowledge graph for THIS repo (skill vendored since KANBAN-065, no build existed) + derived-artifact Pages site mirroring llm-mailroom-graph + links wired into README/docs/wiki both directions; (4) docs currency sweep to post-v0.20.0 code truth per mailroom doc conventions. Immovable-at-root documented: pyproject/requirements/CHANGELOG/README/AGENTS/dotfiles. hermes 2026-08-23 changelog-only SHIPPED 2026-08-23 — entity-extraction commit 1ac476c pushed (38 files): fresh-install archive proof GREEN on pushed tag v0.20.0 (langchain-core 1.6.0 live on default PyPI — the train’s floor failure was a stale index view; src.prompts imports clean from the archived tree, 103 prompt versions); root de-clutter shipped with all live references updated in lockstep — SCORING.md→docs/SCORING.md, deploy_phoenix.sh→scripts/deploy/, root memos/ merged into docs/memos/ (site memos tab now serves all 34 memos incl. v34–v39, was silently 22), tracked .bak cruft removed, legacy root board symlinks deleted; frozen history untouched. Graphify built code-only (3,402 nodes / 7,252 edges / 151 communities, graphify 0.9.48) and published at llm-entity-extraction-graph (verified HTTP 200 on / and /report.html); NEW docs/sister-repos.md umbrella map; links wired README/wiki Home/wiki _Sidebar + mailroom reciprocal commit b1ef37a. Derived artifacts regenerated (meta.json scoring_md, memos.json 22→34, benchmarks refreshed live 1,438 rows, posit pages re-rendered). Targeted suite (test_posit_site + test_graphify_skill): 12 passed, failures byte-identical to the documented chronic pair (test_rendered_pages_committed, test_quarto_render_is_deterministic_and_clean — pre-existing derived-site classes). Issue #38 auto-closed COMPLETED via closes #38. Same-day follow-up: all 5 AGENTS.md path-doctrine lines APPROVED by Jack and landed (file-map row, scoring pointer, checklist cite, memo cite, memos-section header) — zero stale memos//SCORING.md refs remain in agent doctrine.
KANBAN-079 #37 done Enron dedup GT enrichment — content-topic + sentiment labels, two-config GT separation — human directive 2026-08-23: add content_topic (+ evidence) and sentiment_score/sentiment_label GT to Lucius-Morningstar/enron-correspondence-dedup; split the repo into TWO card-declared configs — default = blind rows (filename/text/subject/split/metadata), ground_truth = answer keys joined on filename — so mailroom agents load blind by default and the viewer keeps both for human auditing; monolithic all-columns jsonl leaves the repo root; labeler lives in Enron-Evaluation-Environment beside the shared subclass module; hermes 2026-08-23 v0.20.1 SHIPPED 2026-08-23 — entity-extraction commit bd8b2c7 pushed: publisher v2 rewrite (enrichment stage importing Enron-Evaluation-Environment labelers content_topics.py + sentiment_scorer.py @ c3bb908, never forked; two card-declared configs; legacy monolithic jsonl DELETED from the Hub root; per-file sha verify LFS ≥10MB + round-trip below); determinism proven (two independent builds byte-identical across ALL FOUR data files); Hub Lucius-Morningstar/enron-correspondence-dedup republished 247,523 rows / 222,572+24,951 per config — all four files sha-verified GREEN; datasets-server conversion GREEN (4 splits, pending 0 / failed 0) and first-rows proof: default/train exposes ZERO GT keys, ground_truth/train serves all nine, both lead allen-p/_sent_mail/1. (join integrity live); distributions: topics general_business 74.2% / energy_market 8.7% / …all 11 keys populated, sentiment 167,964 neutral / 51,668 positive / 27,891 negative; honest gaps on card+manifest (lexicon sentiment = weak labels, single-topic ~2000-char head window, exact-hash dedup only); pins 3 deliberate re-pins + 10 new test_kanban079_gt_separation.py; suite 621 passed (+19 vs pristine HEAD 602), failure set byte-identical (8 chronic posit renders)
KANBAN-078 — (board-only) done Mailroom dataset browser notebook — human directive 2026-08-23: implement a Docile-style dataset browser in llm-mailroom under a dedicated notebooks folder (docile/tools/dataset_browser.ipynb pattern: thin notebook + real tool module). Plan: notebooks/dataset_browser.ipynb + reusable loader module; browses the pilot sample set via docs/examples/samples/manifest.csv (30 rows: provenance CUAD/external/synthetic, expected class/stage/fields) joined with pipeline catalog state (data/mailroom.db) when present. hermes 2026-08-23 — (changelog-only) SHIPPED 2026-08-23 — mailroom commit c9ea57a pushed: notebooks/dataset_browser.ipynb (thin, Docile-pattern) + reusable notebooks/dataset_browser.py + folder README; ground truth = docs/examples/samples/manifest.csv (30 rows, provenance rule CUAD|external = REAL), observed layer = read-only catalog join (data/mailroom.db, URI mode=ro; missing/schema-less degrade to empty); ipywidgets picker behind new [notebooks] extra, plain-text fallback on core install; kernel-cwd-proof bootstrap; pdfplumber first-page preview (verified live: contract_01 → 2,527 raw chars embedded); HTML escaping pinned; 10 network-free pins src/tests/test_dataset_browser.py; full suite 401 passed (391 baseline + 10). Notebook executed cell-by-cell headlessly incl. hostile-cwd run from notebooks/ itself
KANBAN-077 #36 done llm-mailroom wiki + docs currency pass — sister-repos umbrella map — human directive 2026-08-23 (“comb through the LLM-MAILROOM wiki and update it with all appropriate documentation, references to sister repositories, and references to all governed repositories that additionally fall under the llm-mailroom umbrella of influence”). (a) Canonical docs/ refresh; (b) NEW docs/sister-repos.md; (c) wiki refresh via docs/wiki/sync-wiki.sh. hermes 2026-08-23 — (changelog-only) SHIPPED 2026-08-23 — mailroom b7d2f79 (15 files, +307/−52): architecture.md 13 nodes + real edge map from routing.py + MemorySaver-default checkpointer truth; agents.md §8 Insurance Claims Specialist (schema table + honest gap) + vendored-prompt lineage corrected (sorter alias→V13); configuration.md +MAILROOM_CHECKPOINTER/MAILROOM_JUDGE_VERIFY/VLLM_API_KEY; testing.md+README node-count fixes; sister-repos umbrella map (entity-extraction, dojo-scoring @v0.7.0, Enron-Eval-Env, claims-data-eda, atticus-investigation, llm-mailroom-graph). Wiki push 92341ef: 13 pages (+1801/−522), sync script now refreshes all 7 mirrors from canonical docs at sync time. Verified: suite 391 passed (= documented baseline, docs-only), wiki pages all HTTP 200, Agents page renders Insurance Claims Specialist + SORTER_PROMPT_V13 + claims-data-eda. AGENTS.md staleness cured under the protection gate with user approval. Changelog [Unreleased] entry in ship commit; closes #36
KANBAN-076 — (board-only) done HF family sync finish: Hub .json* loader landmine ROOT-CAUSED via live canaries + all repos repaired + deduplicated Enron correspondence published — human directive 2026-08-23 (“finish up the sync of the datasets to the huggingface hub, including the new Enron dataset which has been deduplicated and cleaned”). (a) Surgical repair of enron-correspondence + docclass-merged: root manifest.json is ingested by the Hub’s JSON loader as a 1-row data table (CastError “column names don’t match” on reconvert — KANBAN-073/074 lesson); canaries (kanban076-canary{1..4}) proved the sharper rule — ANY path whose name contains .json is ingested as data rows (.json.txt and subdirs included; only an extension with no json substring survives) — so manifests ship as manifest.txt ONLY (round 2), corpus blobs asserted untouched across every hop; (b) NEW enron-correspondence-dedup: exact-duplicate removal per Enron-Evaluation-Environment scripts/dedupe.py body_hash rule (md5 over UTF-8 text; first occurrence wins by maildir-path order; empty bodies never deduped against each other) applied to the sha-verified staged export (local sha256 == hub LFS 0554a5973935…); splits recomputed+asserted via family assign_split() on dedup’d filenames (0 mismatches; 222,572 train / 24,951 test); schema guard + provenance card + manifest.txt ONLY; (c) network-free regression pins: no publisher stages ANY path containing .json, plus builder/publisher metadata-uniformity guards (round 3: MAUD’s nested maud_categories dict — present only in later row-groups — crashed the loader’s struct cast; round 4: CUAD’s list-typed applicable_categories needed JSON-string treatment too; every metadata value is now a plain string on all rows, fingerprint unchanged cd652e77…). hermes 2026-08-23 — (changelog-only) SHIPPED 2026-08-23 — final datasets-server verification ALL THREE GREEN: enron-correspondence 3 parquet files / 517,390 rows / 8 typed features; enron-correspondence-dedup 2 / 247,523 / 8 (hub LFS sha == local e2f7241f…-built artifact); docclass-merged 1 / 700 / 7 after rounds 3–4 (normalize_metadata_rows(): uniform key union, EVERY value a plain string — nested dicts AND lists as sorted-key JSON strings; fingerprint unchanged cd652e77…; final blob af0a5324bb65… local==hub). Canary repos deleted post-verdict. 15 network-free pins in tests/test_kanban076_hf_sync_finish.py. Changelog [Unreleased] entry in ship commit
KANBAN-075 #35 done Adapt Karpathy coding guidelines into core AGENTS.md doctrine — human directive 2026-08-22: integrate multica-ai/andrej-karpathy-skills insights into this repo’s AGENTS.md AND apply the core insights to the operator’s Hermes agent files. Adapt-don’t-vendor (KANBAN-066 recipe); upstream pin 2c606141936f1eeef17fa3043a72095b4765b9c2; four principles (Think Before Coding / Simplicity First / Surgical Changes / Goal-Driven Execution) mapped onto house doctrine in a new AGENTS.md section, subordinated to board governance + append-only prompt versioning; provenance sidecar .opencode/agents/CODING_GUIDELINES_PROVENANCE.md; network-free mechanics pins in tests/test_coding_guidelines_agent_file.py. Overlap audit: existing “surgical” mentions are test-scope selection only — additive doctrine, no duplication. hermes 2026-08-22 — (changelog-only) SHIPPED 2026-08-22 — AGENTS.md guidelines section + sidecar + 9 pins; suite 597✓ / 7 documented posit fails unchanged / 4 skip (= documented 588 baseline + exactly the 9 new pins); changelog [Unreleased] entry in ship commit; closes #35
KANBAN-070 — (board-only) blocked Populate empty BT datasets: mailroom-maud-contracts + mailroom-s1-corporate-records — residue carve-out from KANBAN-069 (issue #34, closed): both datasets exist in Braintrust but hold zero rows because the upstream streaming runs never completed. Run the existing streamers (stream_maud_to_bt.py pulls MAUD v1 from Zenodo/HF mirror; stream_s1_exhibits.py discovers corporate-record exhibits via SEC EDGAR full-text search), verify row counts + dispositions in the live catalog, then extend the HF mirror via the KANBAN-069 export→publish pair (skill: mlops/bt-hf-dataset-mirror). MAUD = CC BY 4.0 (Atticus Project); S-1 exhibits = public EDGAR filings. hermes 2026-08-22 — (changelog-only) BLOCKED 2026-08-22: BT write path dead — org plan quota num_log_bytes_calendar_months violated on every logging batch (dataset insert reported “152 inserted, 0 failed” but live catalog shows 0 rows — the write rolled back when telemetry quota killed the session; verified via read-only BTQL 2026-08-22). MAUD 152-row local dump survives at data/maud/contracts.jsonl. Unblock: BT plan upgrade OR skip-BT entirely — superseded by KANBAN-071 which publishes these corpora direct-to-HF from upstream sources
KANBAN-069 #34 done Mirror Braintrust evaluation datasets → Hugging Face Hub — sync the eval ground-truth datasets from Braintrust to HF Hub (Lucius-Morningstar) for universal agent/eval-runner access. Braintrust stays READ-ONLY (AGENTS.md: BRAINTRUST_LOGGING=disabled preserved — live-catalog GET + BTQL reads only; streamer defaults absent from the catalog recorded as skipped, never created). One HF dataset repo per dataset with provenance cards (CUAD CC BY 4.0, source corpus + BT linkage); exports staged gitignored (GH001); docs in data/README.md. hermes 2026-08-21 — (changelog-only) issue #34; SHIPPED 2026-08-22 — 3 repos live + verified: mailroom-cuad-contracts (50 rows + 546 page PNGs, 596 row→image refs resolve), mailroom-cuad-contracts-full (510 rows, LFS sha256 byte-identical), mailroom-lb-hearsay (5 rows); tooling scripts/datasets/export_bt_to_hf.py (read-only BTQL export) + publish_hf_mirror.py (cards+upload+sha verify); honest gaps in EXPORT_SUMMARY.json (MAUD/S-1 exist-but-empty upstream, LB-classification names never created — populate upstream, re-run pair); suite 571✓ (+8 network-free pins tests/test_kanban069_hf_mirror.py), failure set = documented 7 posit renders unchanged. Carve-out: populating the empty upstream datasets is follow-on work for their streamer owners, not this card
KANBAN-068 #33 in_progress Dedicated Jupyter notebooks per scoring process (carve-out of KANBAN-067, human directive): six fully functional notebooks showcasing every feature of each scoring process end-to-end. hermes 2026-08-24 — (changelog-only) EXEMPLAR SHIPPED 2026-08-24 per human sweep directive (entity 8bf6e49): notebooks/03_doc_type_bundles.ipynb (doc-type bundles = the v0.7.0 headline process) — DOC_TYPE_BUNDLES walkthrough + get_doc_bundle honesty-resolver tour + closing honest-gap table derived from REAL experiment_log data (195 records / 19,642 scored doc rows: contract ×16783, correspondence ×407, merger_agreement ×335, corporate_record ×69, compliance_filing ×44, due_diligence ×4 have real benchmark rows; court_opinion + insurance_claim genuinely declared-pending). Thin-notebook KANBAN-078 pattern: kernel-cwd-proof bootstrap, stdlib-only cells, zero network/LLM. Guards: tests/test_kanban068_bundles_notebook.py (4 network-free incl. hostile-cwd headless nbclient execution asserting the gap summary). Install: new [notebooks] extra ⇔ requirements/notebooks.txt under the KANBAN-081 parity contract, joined into [all]/all.txt. notebooks/README.md scopes all six + conventions. REMAINING: 01 classification, 02 typed-field extraction, 04 audit/verification, 05 chained pipelines, 06 report/aggregation — same pattern. Suite 663 passed.
KANBAN-067 #32 done Full-coverage scoring for every agent & document type — human directive 2026-08-21: llm-dojo-scoring must provide scorers for ALL mailroom agents AND final outputs that vary by processed document type (contracts, merger agreements, correspondence [emails, attorney demands, client correspondence], insurance claim documentation, court opinions + other native types). Honest-gap mandate: where adequate metrics don’t exist for a type/agent, say so in the implementation and keep it fully modular for new scorers as the scoring repo grows. Plus dedicated Jupyter notebooks per scoring process (fully functional, showcasing all features). Gap audit: profiles cover 22/22 agents (dojo v0.6.0); missing = doc-type-aware suites (insurance_claim not a native class), profile doc-bundle resolution, all notebooks. hermes 2026-08-21 v0.7.0 (dojo) + changelog-only consumers issue #32; SHIPPED 2026-08-21 across all three repos — Phase 1 mailroom 99536d8: insurance_claim 7th first-class class at every court_opinion-parity surface (schema/registry, taxonomy, specialist agent, graph node, classifier+sorter vocab=7, derived prompts SORTER_PROMPT_V13 from production-alias V0 + SORTER_VISION_PROMPT_V1 insurance-check-#5 specific-before-generic, predecessors byte-frozen w/ immutability tests, synthetic FNOL fixture); suite 391✓ (was 376). Phase 2 dojo v0.7.0 @ 51822bc: doc_bundles.py DOC_TYPE_BUNDLES (8 doc classes, separate doc: namespace), honesty resolver resolve_doc_bundle() -> (bundle, used_fallback) (explicit flag, never silent default; raises when fallback disabled), insurance_claims_specialist = 23rd profile (exact-set pin re-pinned deliberately + preexisting-22-profiles regression); HONEST GAPs declared in-bundle (MAUD merger / Enron correspondence / DE-SynPUF claims scorers PENDING; CUAD contracts + LegalBench court_opinions real today); suite 209✓/5 skip (was 193). Phase 3 consumers changelog-only: mailroom 08f9bd7 pin→v0.7.0 (provenance verified via direct_url.json, 391✓ unchanged); entity 61c57b3 pins pyproject+requirements→v0.7.0 (fixed pre-existing requirements.txt drift: was v0.4.0/comment-v0.1.2 vs pyproject v0.6.0; bridge imports verified; 561✓/6 skip unchanged). Notebooks carved out to #33 (KANBAN-068)
KANBAN-066 #31 done True-GEPA upgrade of the prompt-engineer agent — human directive 2026-08-21: rewrite .opencode/agents/prompt-engineer.md’s GEPA sections to be source-true to gepa-ai/gepa @ pinned commit: minibatch acceptance criteria (StrictImprovement default / ImprovementOrEqual lateral-move variant), candidate-selection strategies (Pareto default / CurrentBest / EpsilonGreedy / TopK), frontier types (instance / objective / hybrid / cartesian), component selection (RoundRobin vs All), ASI reflective datasets + separate reflection_lm + skip_perfect_score, system-aware merge preconditions (common ancestry, validation-support disjointness, accept iff score ≥ max(parents)), budget/stop/eval-cache mechanics — governed workflow (version-key identity, same-surface A/B, noise floor, chunked surfaces) preserved. Plus PROMPT_ENGINEER_GEPA_PROVENANCE.md sidecar + network-free consistency test. Tooling/docs-only. hermes 2026-08-21 — (changelog-only) issue #31; SHIPPED 2026-08-21 @ cb61919: agent-file GEPA sections source-true to gepa-ai/gepa @ b265bf9ca77f — 9-step loop (Pareto parent selection w/ CurrentBest/EpsilonGreedy/TopK alternatives, epoch-shuffled seeded minibatches, full-trace ASI capture, reflection-dataset construction w/ separate reflection_lm + skip_perfect_score=True, component-scoped proposal RoundRobin-vs-All, same-minibatch child eval, StrictImprovementAcceptance gate w/ ImprovementOrEqual lateral-move variant, frontier recompute across instance/objective/hybrid/cartesian, scheduled system-aware merge) + Phase 3.5 four source-true merge preconditions (common ancestry, validation-support disjointness merge_val_overlap_floor=5, composable shared components, accept iff >= max(parents)) + Phase 0/5 per-cell dominance (get_pareto_front_mapping) + rejected-mutation ledger. Sidecar PROMPT_ENGINEER_GEPA_PROVENANCE.md (pin/license/source-map/re-sync recipe) + tests/test_prompt_engineer_gepa.py 12 network-free pins green — full suite 563✓, failure set = the documented 7 pre-existing posit-site renders (unchanged). Governed workflow markers asserted intact by test.
KANBAN-065 #30 done Vendor Graphify agent skill into BOTH repos — human directive 2026-08-21: install the graphify knowledge-graph skills from Graphify-Labs/graphify for future agent use. Upstream’s official opencode skill (graphify/skill-opencode.md + 8 references/ sidecars, Apache-2.0/MIT) copied verbatim to .opencode/skills/graphify/ in llm-entity-extraction AND llm-mailroom, plus a PROVENANCE note (source tag/commit/license) and network-free consistency tests pinning both copies identical. Tooling-only: no pipeline/prompt/dependency changes. hermes 2026-08-21 — (changelog-only) issue #30; SHIPPED 2026-08-21: skill vendored in both repos (byte-identical, diff-verified vs upstream v8 @ b2cd362) + 5+5 network-free tests green + mailroom suite 376✓ / entity 551✓ (7 pre-existing site-render fails, KANBAN-060 lane — proven identical on pristine HEAD via temp worktree)
KANBAN-062 #28 done Lane A — Sorter Review (agent second opinion after classification) — architecture-alignment build (human-approved 2026-08-21, full-build option 1): upstream llm-dojo-scoring v0.6.0 registry additions (sorter_reviewer + contract_auditor + five companion specialist auditors + arbiter, audit/classification bundles), mailroom sorter_reviewer agent + graph wiring (medium-band exhaustion → agent review before human escalation; reviewer label wins at high confidence), state additions, network-free tests. Attacks the 217-item PENDING human-review backlog (KANBAN-006). hermes 2026-08-21 v0.19.1 issue #28; SHIPPED 2026-08-21: mailroom v0.4.0 — Lane A live (review_classify node: blind second opinion, winning-label application, both-opinions preservation; 33 new lane tests; mailroom suite 355✓)
KANBAN-063 #29 done Lane B — Judge in-pipeline + Arbiter escalation — architecture-alignment build (human-approved 2026-08-21): wire JudgeAgent in-graph after extract gated by needs_judge_review / ambiguous-band confidence (zero added LLM calls on clean runs), new arbiter agent on judge failure (accept-with-caveats / bounded retry_extract / human_review), routing + state additions, traced generations scored via the unified engine’s audit bundle. Depends on KANBAN-062 upstream release. hermes 2026-08-21 v0.19.1 issue #29; SHIPPED 2026-08-21: mailroom v0.4.0 — Lane B live (gated judge_verify + bounded arbiter; cost contract: judge fires only in the 0.70–0.85 ambiguous band, MAILROOM_JUDGE_VERIFY=off kill-switch; fail-safe on every error path; mailroom suite 355✓)
KANBAN-061 #27 done Unified scoring layer — metric registry, T0-T3 tiers, agent profiles/bundles, unified emitter + llm-mailroom field-scoring de-duplication (proposal: ~/Desktop/Cold_Storage/scoring-refactor-proposal.md, human-approved 2026-08-21, option C) — four phases, zero breaking changes: (P1) upstream llm-dojo-scoring registry.py (YAML-backed metric definitions, every existing function mapped to tiers — f1/binary_metrics→T0, precision/recall/f2/jaccard/field_presence/laziness/cost→T1, confusion/failure-modes/bootstrap→T2, raw logs→T3) + bundles.py (classification/extraction/extraction_open/cost/factuality/laziness_detection/transcription/audit/reporter) + profiles.py (agent profile registry: sorter, 6 specialists, judge, boss, pdf_transcriber, image_extractor, archivist, audit_agent⭐NEW — first-class metric identity for KANBAN-060’s audit pass); ship v0.5.0, network-free tests; (P2) emitter.py (unified score emitter: Langfuse/local sinks) + pruning.py (tier-based dashboard filtering); entity-extraction src/score_emitter.py bridge + dep re-pin; (P3) mailroom: add dojo dep, swap 12 import sites off observability/field_scoring.py (1,273-LOC duplicate), wire taxonomy→package Settings, consolidate 37 flat SCORE_CONFIGS into the registry, deprecate-not-delete; (P4) docs/CHANGELOGs ×3 repos, tier-filtered dashboards. Constraint: calculations untouched (Hungarian matching, embedding rescue, bootstrap CI, CUAD equivalences); old APIs keep working. hermes 2026-08-21 v0.19.1 issue #27; SHIPPED: llm-dojo-scoring v0.5.0 (bb3a78c, repaired 6ced4d7) + v0.5.1 (ebcfc68) — both tagged w/ GitHub releases; llm-mailroom v0.3.2 (42c8624) migrated onto the engine; this repo: pin @v0.5.1 + src/score_emitter.py bridge (+5 tests green). Suites: dojo 187✓ / mailroom 326✓ / entity 540✓ (7 pre-existing site-render fails, KANBAN-060 lane). FOLDED DONE 2026-08-21 per human directive, closed alongside the KANBAN-062/063 train: dojo v0.6.0 (review/audit profiles) + mailroom v0.4.0 + entity v0.19.1 re-pin @v0.6.0; issue #27 closed in the same pass.
KANBAN-038 — (board-only) backlog Docclass vision full benchmark + PDF retention for MAUD/S-1 — † numbered KANBAN-038 after a collision — KANBAN-037 was already claimed by the Posit portal card (2026-08-16, opencode) when this follow-on (reserved as “KANBAN-037” inside KANBAN-033’s close-out entry) was posted; reconciliation post on the discussion board. (1) full-pages vision benchmark (--vision-pages all, larger sample) on the merged docclass surface with sorter_docclass_vision_v0 (the pilot validated page-1 vision only, n=8); (2) retain MAUD (Zenodo archive) + S-1 (EDGAR) PDFs locally so the vision arm covers merger_agreement + corporate_record rows, not just CUAD contracts; (3) data-side subclass GT repair (MAUD consideration backfill for the 57 GT-“other” rows; S-1 streamer label fixes) — the unlock for the GT-bound subclass metric (56/69 full-676 subclass misses are GT gaps). unclaimed v0.19.0 memo memos/docclass_v3_merged_benchmark.md §4/§uncertainties
KANBAN-005 #4 backlog Mirror sync → llm-mailroom (cross-repo) — apply the v22/v23 champion prompts to the llm-mailroom pipeline project (Langfuse key file drop-in + sync_langfuse_prompts.py --env-file); regenerate its synced experiment log. unclaimed v0.19.0 AGENTS.md “Langfuse projects” / “Mirror sync”; scripts/eval/sync_langfuse_prompts.py
KANBAN-006 #5 backlog HITL annotation queue processing — work the pending llm-dojo queue items (extraction < 0.85 + sorter failure queue — 217 PENDING at last count): adjudicate, feed corrections into the next prompt iteration. Tooling fix (bounded status scan) landed 2026-08-14. DEFERRED 2026-08-16 — moved back to backlog (stale 2 days, owner released); 217 PENDING items documented on the discussion board; needs a dedicated processing session (Langfuse adjudication + corrections feed). unclaimed v0.19.0 scripts/eval/run_annotation_queue.py status; wiki Annotation-Queues.md
KANBAN-064 — (board-only) done Modal+vLLM offline serving capability (cross-repo: llm-mailroom) — human-directed 2026-08-21: framework-in-place so llm-mailroom can be served locally/offline later via a Modal-deployed vLLM exposing an OpenAI-compatible /v1 API. API calls (OpenRouter) REMAIN the primary serving path — this is a configuration capability only; zero behavior change for API-mode runs. Scope: deploy/modal_vllm.py Modal app (env-configurable MODEL/GPU/quantization/max-len/token auth), vllm provider hardening in llm/providers.py (optional VLLM_API_KEY, env-overridable base URL already present), .env.example + docs, network-free tests. hermes 2026-08-21 mailroom v0.4.1 SHIPPED 2026-08-21: mailroom v0.4.1 (ed4f576, tagged + GH release) — deploy/modal_vllm.py Modal app (env-config MODEL/GPU/quant/max-len/bearer token, persistent HF cache volume), optional VLLM_API_KEY provider hardening, [deploy] extra, deploy/README.md runbook; 12 new network-free tests, mailroom suite 367✓; OpenRouter unchanged as primary
KANBAN-011 #9 backlog Post-v23 model sweep (gated OPEN) — run v22/v23 prompts × {deepseek-v4-flash, deepseek-v4-pro} on the same 50 docs to quantify the remaining model-bound segmentation gap (the v18 sweep proved scope-fidelity is model-agnostic; confirm the ko 0.85→0.89 plateau closes at the newest prompts). unclaimed v0.19.0 memo model_sweep_v18.md; V16_PROPOSITION.md §9.3/§15
KANBAN-072 #25 in_progress Contract specialist improvements — verbatim CUAD/MAUD category questions + agent inquiry schemas (human-filed issue 2026-08-21) — integrate The Atticus Project’s verbatim CUAD description prompts (Change Of Control, Anti-Assignment, Notice Period To Terminate Renewal, Termination For Convenience, Price Restrictions, Most Favored Nation, Governing Law, … full 41) and MAUD verbatim deal-point queries (MAE/MAC definition + scope exclusions, ordinary-course covenants + consent restrictions, No-Shop/Go-Shop + fiduciary-out, breakup/termination fees, Outside/Drop-Dead Date) into the contracts specialist via the unified agent-inquiry schema from the issue (agent_inquiry_templates: per-category prompt + target_dataset). Complementary to the landed registry config/clause_categories.yaml (#26, closed under KANBAN-061): the YAML holds names/class_ids/answer_types/confidence; THIS card wires the verbatim question texts + inquiry templates into prompt surfaces/agent routing and measures on the extraction surfaces (v34→v39+audit lineage, KANBAN-054..060). Scope guard: same-surface A/B per governed workflow; no champion change without a paired win. hermes 2026-08-24 v0.19.5 issue #25; config/clause_categories.yaml; GEPA lineage memos
KANBAN-054 — (board-only) done Extraction agent: anti-collapse prompt v34 + ContractEval-rubric KPIs (human request 2026-08-19) — two-part: (1) prompt contracts_specialist_v34 must NEVER collapse expected fields or clause groups: R1 field-presence self-check (v32@510 presence: contract_value 0.39, renewal_terms 0.37, effective_date 0.88, term_length 0.83), R2 category-level completeness (the 32 canonical CUAD YES/NO categories as a post-extraction checklist — present category ⇒ ≥1 item + ≥1 canonical-tagged reasoning entry; never fabricate absent), R3 verbatim quoting at the GT span grain (mapping memo: verbatim 9.2% vs ≥0.7 containment 42.7% @v32 — paraphrase penalty dominates; GT labels are the clause’s own text). (2) KPIs: ContractEval-rubric F1/F2/Jaccard/false-nr + semantic coverage bands become CORE per-run metrics (scores.contracteval_kpis via src/contracteval.py::run_kpis, injected in log_experiment_to_repo for both extraction runners), rendered in the experiment log, charted in site trends (F2-lead per human decision), backfilled over historical records (eval machine). Decisions (human 2026-08-19): add alongside existing metrics (F2 leads), verbatim-at-span-grain formulation, A/B on the 50-doc chunked surface only (full-corpus KPI baseline stays v32). 2026-08-19 (code complete, in_review): v34 prompt + run_kpis + log renderer + site trends chart + backfill script all landed; 132 surgical tests green (incl. test_extraction_kpis_land_in_record, test_run_kpis_block, test_contracts_v34_anti_collapse_rules); KPI block verified on the stored v32@510 record (F1 0.1579 / F2 0.1049 / Jaccard 0.2129 / false-nr 0.6818 / semantic ge0.7 0.4332); CHANGELOG entry in same commit. 2026-08-19 (HALF-CORPUS A/B LANDED, human directive replaced the 50-doc plan):** qwen3.7-flash_contracts_specialist_v34_extraction_chunked_half (255 docs, seed 42, chunked, Langfuse llm-dojo, ~$0.23) = overall 0.8738 / presence 0.9711 / verified-prec 0.9904 / schema 1.0 + KPIs (F1 0.1331 / F2 0.0963 / Jacc 0.2803 / recall 0.0813 / false-nr 0.4775 / laziness 0.8632 / semantic verbatim 8.2% · ge0.7 38.8%); KPI block + laziness metric extended (scores.contracteval_kpis.laziness = ContractEval §III-D no-related-clause rate, alias of no_related_rate; backfill refreshes blocks lacking it).** opencode 2026-08-19 v0.19.0 src/prompts.py v34; src/contracteval.py; run_extraction_eval.py; memos/contracts_specialist_v34.md
KANBAN-055 — (board-only) done One-pass extraction: item-level category split → contracts_specialist_v35 (the THIRD anti-collapse lever) — the user directive (2026-08-19, after the V35 rename): keep opencode’s v34 (R1/R2 structural), add the missing item-level lever. v33 (RETAG) fixed umbrella tags; opencode’s v34 (KANBAN-054) added R1 field-presence self-check + R2 category-level completeness. v35 closes the THIRD collapse mode neither targets — ITEM-LEVEL CATEGORY COLLAPSE: a single key_obligations item holds duties from TWO different canonical categories (e.g. “Neither Party shall assign this Agreement nor use its trademarks” folds Anti-Assignment INTO Non-Disparagement / IP Ownership), routing to one category and scoring 0 on the other; measured ~15,516/33,312 umbrella-tagged entries on the stored v31/v32 corpus. v35 = v34 + ONE surgical append: one ENTRY per distinct category’s duty within a clause (a two-category clause emits one entry per duty, each tagged with its OWN canonical name), and EXACT-category tagging only (never a sibling / family / generic ‘IP’). Registered CONTRACTS_SPECIALIST_PROMPT_V35 in PROMPT_VERSIONS; test test_contracts_v35_item_level_category_split green (60 prompt tests). Rebased cleanly onto KANBAN-054 (prior v34 name-collision resolved by rename). 2026-08-19 (HALF-CORPUS A/B LANDED — human directive replaced the 50-doc plan): qwen3.7-flash_contracts_specialist_v35_extraction_chunked_half (255 docs, seed 42, chunked, Langfuse llm-dojo, ~$0.23) = overall 0.8670 / presence 0.9709 / verified-prec 0.9909 / schema 1.0 + KPIs (F1 0.1408 / F2 0.1024 / Jacc 0.2955 / recall 0.0866 / false-nr 0.4437 / laziness 0.8554 / semantic verbatim 9.0% · ge0.7 39.8%). Paired A/B vs v34 (identical 255 docs, scripts/reporting/ab_paired_compare.py, n_boot 2000 seed 42): Δ +0.0068 (v34−v35), CI [−0.0034, +0.0169], P(win) 0.909 — INSIDE the noise band → LOGIC REPAIR, NO champion change; v34 stays the half-corpus leader. Field-level: v34 BEATS on term_length (+0.0523, CI [0.004, 0.102], P 0.982); v35 directionally ahead on governing_law/parties (inside band); key_obligations tied (0.7612 vs 0.7618). KPIs show the v35 item-split lever improved the semantic bands (verbatim 8.2→9.0%, ge0.7 38.8→39.8%, laziness 0.863→0.855, false-nr 0.478→0.444) without an aggregate win. hermes 2026-08-19 v0.19.5 CHANGELOG [Unreleased]; src/prompts.py CONTRACTS_SPECIALIST_PROMPT_V35
KANBAN-060 — (board-only) done Absent-family recall mechanism → runner-level AUDIT PASS (human directive 2026-08-20: “most fitting and appropriate methods…”) — 2026-08-20 (RUNNER LANE IMPLEMENTED by opencode: contracts_audit_v0 prompt constant + ContractsSpecialist.audit_extraction() (same windows, missed-category feedback, union merge, never-remove) + --audit runner flag + parameters.audit record field; 14 new tests green (7 unit + 7 smoke); dry-run verified; run name RESERVED: qwen3.7-flash_contracts_specialist_v39_audit_extraction_chunked_half, manifest data/manifests/extract_v39_audit_half.jsonl; AUDIT RUN LAUNCHED — same 255/seed-42/chunked surface vs v39). 2026-08-20 (A/B LANDED — RECALL/F2 CHAMPION; prefix-cache consolidation implemented):** run KPIs — recall 0.3627 / F1 0.4605 / F2 0.3963 / precision 0.6306 / false-nr 0.2388 / verbatim 0.371 / laziness 0.799 (vs v39: R 0.2833 / F1 0.4146 / F2 0.3244 / P 0.7727 / false-nr 0.3643). Paired per-doc gate (252 shared, seed 42, 2000 boots): recall +0.0637 BEATS (P 1.000), F2 +0.0489 BEATS (P 1.000), F1 +0.0258 inside band (P 0.950), precision −0.0942 LOSES (P 0.000). Mechanism direct: absent positive pairs 612→399 (−34.8%); Post-Termination 59→19. Audit-added: 1,139 clauses/227 docs = 55 TP + 797 in GT-present cats (357 overlap a GT label, 440 real sibling sentences GT never sampled — CUAD partial-GT) + 342 GT-absent (fp). Cost: 12.2M prompt tokens $0.49 (audit +108% prompt vs v39 5.84M $0.28) → prefix-cache consolidation implemented + tested: audit call reuses the extraction system prompt + byte-identical user prefix (OpenRouter cache-read = 20% of input price) → next audit run ≈ $0.33-0.35. Verdict: audit = new recall-side champion (F2-lead); v39 = precision champion; Pareto = {v39 (P), audit (R/F2)}.** — 2026-08-20 (DIAGNOSIS COMPLETE by prompt-engineer — mechanism = EMISSION-STAGE OMISSION, prompt levers cannot fix it; handed to opencode’s runner lane, NO v40 constant written): in-text verification over ALL 645 v39 absent pairs (Braintrust corpus fetched read-only, NFC+whitespace-collapse+casefold both sides): absent 645 = all-in-text 551 (85%) + some-in-text 40 + none-in-text 54 (8%, GT debt) + chunk-invisible 2 (0.3% — architectural hypothesis refuted); model behavior on the 591 in-text pairs: no-output 523 (81%) + partial 82 + sibling 18 + other 22; 429/591 in SINGLE-WINDOW docs (220/254 docs single-window — no chunking confound); median 15 items/doc (no global under-emission — category-selective skipping: Covenant 21/21, Competitive Restriction 32/33, Volume 27/28 zero-output). v37 scan-family / v38 re-scan / v39 completion all measured flat (absent 636→645): a single generation cannot re-read — the fix is a SECOND CALL with feedback. SPEC (runner lane, opencode): post-extraction audit call per doc/chunk — window text + canonical-tagged extraction + 32-category list → missing_obligations[{category, verbatim full-sentence clause}]; union-merge with dedupe, ADDING-only, never-fabricate; recommend versioned contracts_audit_v0; cost ≈ +1 call/doc ≈ +$0.25-0.30/run; A/B same surface (255/seed 42/chunked) v39+audit vs v39; projection +105-210 TP → recall 0.35-0.41, F1 0.48-0.52. Side findings: Warranty Duration taxonomy gap (164 GT labels, field: None → never scored — taxonomy lane); 54 none-in-text pairs = GT re-validation (data lane); 2 chunk-invisible labels (overlap edge). Run name (v39+audit) to be reserved by opencode; manifest data/manifests/extract_v39_audit_half.jsonl suggested. No paid run launched. opencode 2026-08-20 v0.19.5 —
KANBAN-059 — (board-only) done GEPA iteration 3 → contracts_specialist_v39 (human directive 2026-08-20: “maximize recall + precision + F1 + F2 in the next iteration”) — post-KANBAN-058 corrected-scorer state (255-doc half-corpus, seed 42, chunked): champion v36 = run F1 0.4073 / F2 0.3243 / recall 0.2855 / precision 0.7107 (per-doc paired bootstrap vs v34 P 1.000; vs v37 Δ +0.0100 / vs v38 Δ +0.0095 both inside band); v37 (payment content) = highest run-level F1 0.4170 / F2 0.3382 / recall 0.3004 / precision 0.682; v38 (sparse-family re-scan) = regression on F1 (0.4111). 2026-08-20 (CONSTANT SHIPPED by prompt-engineer, A/B pending): corrected-scorer diagnostics → fp audit (TFC 53 fp = largest, genuine errors, NO enumeration entry; Uncapped fee-cap confusion; Revenue service-fee confusion; Third Party = GT noise, not suppressed) + near-miss decomposition (556 = 371 multi-label under-quote — 35% of positives carry ≥2 GT clause sentences — + 88 sibling + 63 leading-phrase drop + 19 paraphrase + 15 dash-GT). CONTRACTS_SPECIALIST_PROMPT_V39 = v37 (embeds v36 + payment fold) + 4 surgical .replace() edits: entry 27 = Termination For Convenience WITHOUT-CAUSE boundary (precision lever), R2 money-family boundary clarifications, WITHIN-CATEGORY COMPLETION in the grain rule (every distinct clause sentence per category, first-word through final period), R2 one-item-per-clause-sentence strengthen (recall lever; fp-neutral). Test test_contracts_v39_payment_fold_precision_and_completion green (64 prompt + 19 sweep + 12 smokes); dry-run accepts. Run name confirmed reserved qwen3.7-flash_contracts_specialist_v39_extraction_chunked_half; manifest data/manifests/extract_v39_half.jsonl; champion gate vs v36. 2026-08-20 (v39 A/B LANDED — PRECISION CHAMPION, Pareto-frontier): v39 = F1 0.4146 / F2 0.3244 / recall 0.2833 / precision 0.7727 (best ever) / verbatim 29.1% / false-nr 0.3643 / laziness 0.8514 / overall 0.8748. Paired per-doc gates (254 shared, corrected scorer, seed 42): v39 BEATS v36 on precision (Δ −0.0429, CI [−0.0832, −0.0034], P 0.983) and BEATS v37 on precision (P 1.000); F1 Δ −0.0009 (P 0.523) / F2 / recall all inside the band vs v36 AND v37 → statistically tied; v36 BEATS v37 on precision (P 0.018). Pareto frontier = {v39} (strictly best precision, weakly dominant elsewhere); v36 = F1-tied recall-side incumbent (per-doc R 0.3278 highest); v37 dominated. Residual recall mass untouched (laziness 0.85, ~536 absent-family pairs) — next iteration’s lever. opencode 2026-08-20 v0.19.5 —
KANBAN-058 — (board-only) in_review GT/scorer artifact fix → ContractEval KPIs re-scored (human directive 2026-08-19: “ensure this fix is applied”) — 37% of the KPI FN mass is GT-storage artifact, not extraction failure: 493/1,686 whitespace FN (1,416 cells with \n/multi-space runs) + 242 <omitted> FN (695 spans carry redaction markers). Fix applied in src/contracteval.py: new _clean_span() collapses whitespace + strips <omitted>/[omitted] at GT-load time (shared llm-dojo-scoring package untouched); backfill_extraction_kpis.py --refresh added (documented re-scoring pass); tests test_clean_span_*/test_load_master_gt_normalizes_artifact_spans/test_cleaned_gt_span_matches_model_output_verbatim. Effect (re-scored stored records, zero LLM spend): v34 F1 0.1331→0.1740, v35 0.1408→0.1777, v36 0.3277→0.4073 (recall 0.2855, precision 0.7107), v37 0.3256→0.4170, v38 0.3108→0.4111; re-scored per-doc paired bootstrap re-confirms v36 champion (BEATS v34 P 1.000; v36 vs v37/v38 inside band, v36 numerically ahead). KPI blocks backfilled (–refresh), log md + site regenerated, CHANGELOG [Unreleased] Fixed entry added. Residual: 18 literal-newline cells (GT data debt); F1 headroom now genuinely prompt-side. opencode 2026-08-19 v0.19.5 src/contracteval.py; scripts/reporting/backfill_extraction_kpis.py; tests/test_contracteval.py
KANBAN-057 — (board-only) done GEPA iteration 2 → contracts_specialist_v38 (human directive 2026-08-19): next F1 mutation — v36 = F1 champion (0.3277 @255; 2.5× vs v34, P 1.000); v37 = logic repair (payment content measured-improving: contract_value presence 0.396→0.441, false-nr 0.478→0.364, but precision dropped 0.653→0.613 and F1 inside band). v38 mandate: FIGURE OUT how to achieve a real F1 gain. 2026-08-20 (CONSTANT SHIPPED by prompt-engineer, A/B pending): KPI-level fn decomposition of v36 (1686 positive pairs) → v36 FN 1319 = 493 whitespace-artifact (GT-side, scorer card KANBAN-058) + 242 <omitted>-label + 48 genuine-near + 536 ABSENT; the prompt lever = the 536 absent pairs. CONTRACTS_SPECIALIST_PROMPT_V38 = v36 + 2 surgical .replace() edits — enumeration entries 27-29 (Warranty Duration, Competitive Restriction Exception, Volume Restriction — absent/nameless families with real-clause shapes + measured stats) + UNDER-QUOTED FAMILY RE-SCAN sentence in the R2 block (named absent-heavy families, 536/1686 stat). Crossover decision: v37’s payment block NOT folded (F1-flat + precision regression on current scorer; revisit post-KANBAN-058). Test test_contracts_v38_sparse_family_shapes green (63 prompt + 46 sweep); dry-run accepts. Run name reserved qwen3.7-flash_contracts_specialist_v38_extraction_chunked_half; manifest data/manifests/extract_v38_half.jsonl; champion gate vs v36. opencode 2026-08-19 v0.19.5 —
KANBAN-058 — (board-only) done ContractEval KPI scorer/GT defect — whitespace + <omitted> artifacts (flagged by prompt-engineer 2026-08-20, scoring lane) — 493/1319 v36 FN = GT-whitespace artifacts (502/1686 positive labels carry \n/multi-space runs; contracteval_classified lacks whitespace normalization → character-complete quotes fail the TP predicate); 695 GT labels carry literal <omitted>/[omitted] placeholders (242 on positive pairs, unfixable by any model). Fix: whitespace-normalize in upstream llm-dojo-scoring contracteval_classified (or repo-side pair builder src/contracteval.py::evaluate_record) + master-CSV label cleaning; projected F1 0.3277 → ~0.63 with no LLM calls (re-score stored records via run_contracteval_report.py). Prompt iterations must NOT compensate for these artifacts. unclaimed v0.19.5 KANBAN-057 diagnostic
KANBAN-056 — (board-only) done GEPA iteration → contracts_specialist_v36 (human directive 2026-08-19): recall-first mutation on the half-corpus surface + full Pareto-curve selection — directive: run the GEPA prompt-improvement cycle with the prompt-engineer agent after the v34/v35 A/B, targeting (1) recall on truly identified clauses, (2) missed expected fields (v34@255 presence: contract_value 0.396, renewal_terms 0.337, effective_date 0.886, term_length 0.800), (3) verbatim extraction at the GT span grain (v34@255 verbatim 8.2% / ge0.7 38.8% / ge0.3 73.6% — the paraphrase penalty dominates; ContractEval recall 0.081, laziness 0.863), (4) one-or-two-pass capability. Methodology per directive: full GEPA Pareto-curve (score/cost/robustness + instance-level frontier, complementary-lesson crossover when disjoint) with ONLY the critical champion-decision MC sims (paired-bootstrap champion selection on the shared surface — NOT the full sim suite). Champion baseline to beat: v34 = 0.8738 @255 (noise band CI [−0.0034, +0.0169] vs v35; v32@510 full-corpus 0.8807 as the generalization anchor). Tooling landed with this card: scripts/reporting/ab_paired_compare.py (paired same-surface A/B + GEPA verdict; tests green) + laziness KPI (scores.contracteval_kpis.laziness = ContractEval §III-D no-related-clause rate) + backfill refresh. 2026-08-19 (v36 SHIPPED by prompt-engineer, A/B pending): CONTRACTS_SPECIALIST_PROMPT_V36 = v35 + 7 surgical .replace() edits — full-clause-sentence span grain replacing the v10-era ATOMIC-FRAGMENT 10-25-word rules (the rule_contradiction behind 146/448 pure-truncation near-misses + 265 condensed overlaps on the v34/v35 sim-matrix), v35’s item-level guard re-cast to full-sentence quoting, term_length duration-only guard (16/208 MISS), effective_date blank-placeholder carve-out (5/16 fabricated fills; null scores 1.0 on blank GT). Test test_contracts_v36_full_sentence_grain green (61 prompt tests); run name reserved qwen3.7-flash_contracts_specialist_v36_extraction_chunked_half; memo memos/contracts_specialist_v36.md; CHANGELOG [Unreleased] entry in the same commit. 2026-08-19 (v37 DESIGN frozen by prompt-engineer, constant HELD for v36’s verdict): payment/monetary capture + canonical tag discipline — data: contract_value never GT (0/255 expected; predicted 101/255; 113/255 payment-GT docs null), payment family fn = 297/801 = 37% of the laziness mass (Price 0.000 F1, Uncapped 1/46, Volume 3/35), mechanism = 78/255 field-collapse docs (115 fn; 50/78 have emitted-but-untagged money items) + 182 genuine scan gaps; run name reserved qwen3.7-flash_contracts_specialist_v37_extraction_chunked_half; design memos/contracts_specialist_v37_design.md; base = v36 constant in every verdict scenario. 2026-08-19 (v36 A/B LANDED — F1 CHAMPION): results = ContractEval F1 0.3277 (v34 0.1331 → 2.5×), recall 0.2187 (2.7×), precision 0.653, F2 0.2523, Jaccard 0.4373, false-nr 0.3856, verbatim 22.2%, ge0.7 61.9%; aggregate overall 0.8698 vs v34 0.8736 (paired Δ +0.0037 CI [−0.0084, +0.0169], inside band, no regression); paired per-doc F1 bootstrap (254 shared, seed 42): Δ −0.1691 (v36−v34), CI [−0.2016, −0.1398], P(v36 beats) 1.000 → BEATS the F1 metric, champion change v34→v36. v37 crossover now in flight (builds ON v36). 2026-08-19 (v37 CONSTANT WRITTEN by prompt-engineer, tests green, A/B NOT launched): CONTRACTS_SPECIALIST_PROMPT_V37 = v36 + 4 surgical .replace() edits (+3,981 chars; v36 byte-identical) per the frozen design — PAYMENT TERMS & MONETARY CLAUSES mandatory scan family (10 money shapes at v36’s grain, inlined measured examples) + canonical tag discipline (never field-level key_obligations entry; royalty ≠ License Grant; 78/255 stat) + contract_value trigger extension (payment schedule/royalty/minimum commitment = visible consideration; 113/255 stat) + Uncapped/Liquidated enumeration appends; contradiction check passed; test_contracts_v37_payment_monetary_capture green (62 prompt + 15 sweep tests; dry-run accepts). Run qwen3.7-flash_contracts_specialist_v37_extraction_chunked_half (reserved) — command returned to human. 2026-08-19 (v37 A/B LANDED — LOGIC REPAIR on F1; CHAMPION CONFIRMED = v36): v37 = F1 0.3256 (v36 0.3277; paired per-doc Δ +0.0129 v36−v37, CI [−0.0119, +0.0387], P(v37 beats) 0.159 → inside band, no F1 change), recall 0.2217 / precision 0.6129 / F2 0.2541 / Jaccard 0.4561 / false-nr 0.3641 / laziness 0.8398 / verbatim 22.6% / ge0.7 63.0%; aggregate 0.8669 (paired Δ +0.0029 vs v36, inside band); payment-capture axis (human directive) measured-improving: contract_value presence 0.396 → 0.441, false-nr 0.478 → 0.364 (v34→v37), laziness 0.863 → 0.840, key_obligations presence 0.988; CHAMPION = contracts_specialist_v36 (F1 0.3277 @255; 2.5× vs v34, P 1.000). GEPA cycle complete: v36 beat, v37 repair. Next if the F1 chase continues: v38 = v36 + payment content + precision-recovery guard (v37 precision 0.613 vs v36 0.653 is the one regression the payment lever introduced). opencode 2026-08-19 v0.19.5 scripts/reporting/ab_paired_compare.py; tests/test_ab_paired_compare.py; src/contracteval.py
KANBAN-052 #22 blocked Directly-mirrored ContractEval task (arXiv 2508.03080) — replicate their experiment, compare to Table III, run our prompt loop — build the first-class task: dataset builder (scripts/datasets/build_contracteval_testset.py → data/contracteval/contracteval_test.jsonl, the CUAD cuad-qa TEST split = 4,128 (contract, question) pairs / 102 contracts / 41 categories, faithful full-context) + versioned prompt contracteval_v0 (ContractEval’s exact system prompt) + canonical ContractEval evaluator UPSTREAM in llm-dojo-scoring (new contracteval task kind: verbatim-containment TP, F1/F2/acc/prec/recall, token-set Jaccard over positives, false-‘no related clause’ rate with the paper’s hardcoded 1244 denominator + per-category breakdown; bump v0.4.0, re-pin) + dedicated runner scripts/eval/run_langfuse_contracteval_eval.py (full-context faithful: one call per pair, temp 0, max_tokens 5000, input cap disablable; manifest resume; Langfuse traces; task=contracteval experiment-log records) + report tooling scripts/reporting/run_contracteval_report.py (offline: our runs vs the full 19-model Table III + per-category). Decisions (human 2026-08-18): our OpenRouter models first; faithful full-context; staged build-first-run-later; scorer upstream. Milestone 1 (build, network-free) in this card; runs + prompt iteration = follow-on cards. 2026-08-18 progress: runner debugged (dep install, reasoning-effort kill, --tracing-backend phoenix, resume recompute, contract-grouped cache-friendly dispatch); src/phoenix_tracing.py now a REAL sink (OTel root+agent spans, output/score events, OpenAI SDK instrumentation — hide_input_text Goldilocks payload bound, native annotations.* correct/incorrect CODE annotations); pilot 100 pairs F1 0.4578 / F2 0.5053 / Jaccard 0.3433 / false-nr 0.1143 (~$0.07); FULL v0 4,182-pair benchmark COMPLETE: F1 0.5541 / F2 0.6164 / Jaccard 0.5058 / false-nr 0.0289 (paper 0.0289), TP 829/TN 2019/FP 919/FN 415, 0 errors, $2.39 — Table III: F1 #4, F2 #3, Jaccard #1 (tied gemini-2.5-pro), false-nr #2** (reports/contracteval_benchmark.md); Phoenix spans annotated post-hoc (contracteval+jaccard, correct/incorrect); prompt verified verbatim in live traces (extraction, not classification); GEPA iterations 2–4 → contracteval_v2/v3/v4 (v2 trigger/span decoupling; v3 verbatim quote fidelity — F1 0.5550 best; v4 = v3 minus doubt-bias + v2’s smallest-complete-span rule, quote VERBATIM and SMALL, one surgical clause replace, tests 63 green, dry-run clean); gpt-4.1-mini × contracteval_v1 COMPLETE: F1 0.6562 / F2 0.6675 / Jaccard 0.4674 / false-nr 0.0844 (beats the paper’s own gpt-4.1-mini 0.644; manifest data/manifests/contracteval_gpt41mini_v1_full.jsonl); v4 COMPLETE: F1 0.5619 / F2 0.6169 / Jaccard 0.5329 / false-nr 0.0346, TP 821/FP 857/FN 423 (smallest-span rule: 100 FP→TN, J 1.0 wins on worst bloaters, but 1,310/1,567 quotes stayed whole-sentence — the sentence-granularity tail was the J drag); GEPA iteration 5 → contracteval_v5 (fragment permission + every-word/every-part guards — the 19 partial multi-span drops fixed; projected F1 0.568–0.571 / J 0.616–0.636; tests 64 green, dry-run clean); RUNNING: qwen3.7-flash × contracteval_v5 on the REGULAR OpenRouter key (per human directive — gpt-4.1-mini burned ~$17.39 vs qwen ~$2.3, funding key not used for v5), manifest data/manifests/contracteval_qwen_v5_full.jsonl, name qwen3.7-flash_contracteval_v5_contracteval_langfuse reserved.); LOCAL-VLLM CROSS-MODEL (2026-08-18, opencode): qwen3-8b × contracteval_v0 COMPLETE on the node’s single RTX A5000 (GPU 1): F1 0.5646 / F2 0.5313 / Jaccard 0.1454 / false-nr 0.1367, 4,182/4,182, 0 errors, $0 local inference, TP 899/TN 2305/FP 769/FN 345 — Table III: F1 beats the paper’s own qwen3-8b 0.530 and qwen3-8b-thinking 0.540; Jaccard 0.145 = the known v0 over-quote bloat** (reports/contracteval_benchmark.md, record qwen3-8b_contracteval_v0_contracteval_langfuse, manifest data/manifests/contracteval_full.jsonl).** 2026-08-22 (ZOMBIE DIAGNOSIS by hermes — run stalled, NOT running): manifest audit found the v5 run died 2026-08-18 ~15:54 UTC at 3,917/4,182 unique completed (265 missing) with 4749 manifest rows (831 error rows = 564 retryable OpenRouter 429s + 265 fatal 403 Key limit exceeded (weekly limit) on the REGULAR key + 2 transient disk-full). Regular key re-tested 2026-08-22: STILL weekly-capped. Unblock: resume via the SAME runner command (--manifest data/manifests/contracteval_qwen_v5_full.jsonl auto-skips completed rows) once either (a) the regular key’s weekly limit resets, or (b) the human authorizes the research-funding key (~$0.70 remaining cost est.). No experiment-log record exists for v5 until then. opencode (2026-08-18) v0.19.0 issue #22; upstream v0.4.0; CHANGELOG [Unreleased]

Sweep rule: when a release ships, re-target every non-done card to the new [Unreleased] version and move landed cards to the Archive. (Last sweep: v0.18.0 2026-08-15 — KANBAN-004/017/018/019/020 archived below; open cards re-targeted to v0.19.0.)

Discussion board

The full discussion log — every claim, lane move, decision, result, blocker, handoff, and reopening, newest at top — lives in MESSAGE_BOARD_DISCUSSION.qmd (a Quarto document: color-coordinated entries per agent, agent profiles + color legend, inline references/citations to issues, commits, memos, and repo paths). It was moved out of this file on 2026-08-15 (from inline bullets to YAML, then restyled as .qmd on 2026-08-16) to keep the board lean; the log itself is UNCHANGED and remains append-only — never edit a past entry, post a correction.

Post to the discussion board by appending a new entry at the TOP of entries in MESSAGE_BOARD_DISCUSSION.qmd (the ::: {.entry ...} block template in its “How to append” section — date + agent + card + subject + body), and mirror lane changes on the card table above.

Archive (completed work — kept for auditability)

Card | Shipped in | Commit / tag | Result |
KANBAN-097 | [Unreleased] (2026-08-24) | close-out commit (this one); foundation 9483437; concurrent lane 1afa047; closes #51 | Docclass role-prompt mutation iteration testing + per-role eval tasks — DONE. Eval tasks for every roster role (edge/judge-mutation/conflicts/blind-classification), deterministic scoring, –dry-run gates, append-only agent_bench records; silent no-op judge pilot_v1 repaired; insurance_claims_specialist_v1 mutation validated (template-fabrication class eliminated); judge v1 = FPR-dominant Pareto champion (0.00 vs 0.88); reviewer 20/20 after GT-label fix; boss 1.00 / arbiter 0.70. Residue -> KANBAN-100; log-truncation incident -> KANBAN-099. Key results: judge-mutation matched arm n=48 seed-42; conflicts n=10/role; reviewer+insurance edge n=20 each. |
KANBAN-074 | — (changelog-only, 2026-08-22) | this close-out commit; datasets live on the Hub (commitless artifacts) | HF family completeness: deterministic splits + cleaned Enron correspondence corpus — human directive 2026-08-22 (changelog current; metadata/labels/GT/splits across the HF family, incl. cleaned Enron). SHIPPED 2026-08-22 — enron-correspondence NEW: FULL cleaned CMU corpus, 517,390 rows / 150 custodians / zero dropped, GT = shared 10-key labeler with on-row label_evidence (email 505,929 · memo 3,568 · press_release 2,520 · notice 2,842 · letter 2,077 · demand 315 · meeting_request 135 · attorney_demand 4), splits 465,570/51,820 via the family rule md5(filename)%10==0→test; LFS sha local==hub 0554a5973935…; datasets-server GREEN. docclass-merged → schema v3: per-row split column added to v2’s subclass+filename, all three non-null on 700 rows (split_coverage 628/72 in manifest); republished, LFS sha local==hub af7705368c83…, row 0 serves Co_Branding/train. One split implementation (assign_split()), imported by both publishers — no forked rule; builder+publisher guards extended to require it. Family audit recorded honestly: legalbench-full ships native upstream train/test TSVs (complete); 3 BT mirrors are whole-gold eval pools (splits not applicable — documented, not manufactured). New tooling scripts/datasets/publish_enron_correspondence.py; +4 pins (split determinism/single-source, dump split coverage, enron guard pins). Suite 588✓ / 7 documented posit fails unchanged / 4 skip |
KANBAN-073 | — (changelog-only, 2026-08-22) | commit: KANBAN-073 close-out (this commit); dataset live on the Hub (commitless artifact) | docclass-merged schema v2 — contract subclasses + file names on every row; Hub viewer string→null cast-crash fixed — SHIPPED 2026-08-22 (human directive 2026-08-22). docclass-merged now carries non-null string expected_subclass + filename on ALL 700 rows: contracts use CUAD’s own grouping (metadata.category, 28 groups — Marketing, Maintenance, License_Agreements, …) with filename = source-PDF basename; MAUD consideration types + S-1 record subclasses already carried both. Root cause of the reported viewer crash (DatasetGenerationError: Couldn't cast array of type string to null): the old file was written CUAD-first with all-null subclass/empty filename → JSON loader inferred null-typed columns from opening batches → crashed casting later string batches. Fix + guardrails: builder refuses to write any row lacking either field; publisher gained a pre-upload schema guard refusing partial-null uploads (this bug class can never ship again); manifest records schema_version: 2 + coverage (700/700 filenames, 700/700 subclasses, 28 contract groups); 4 new network-free pins (tests/test_kanban071_hf_pack.py → 12). Republished + verified GREEN: LFS sha256 local == hub (3bd9d74de9f1…), deterministic fingerprint cd652e77…, datasets-server serves clean — no pending/failed conversions, first-rows types both label columns as string with real values on row 0. Bonus side effect: the docclass eval runner grades expected_subclass when present, so contracts now participate in subclass scoring (previously skipped) — strictly richer eval surface, zero runner changes. Suite 585✓ / 7 documented posit fails unchanged / 4 skip (+4 = exactly the new pins). |
KANBAN-071 | — (changelog-only, 2026-08-22) | commit: KANBAN-071 close-out (this commit); datasets live on the Hub (commitless artifact) | Full LegalBench subsets + docclass merged corpus → Hugging Face, CUAD-quality label enrichment — SHIPPED 2026-08-22 (human directive 2026-08-22). Two verified Hub datasets under Lucius-Morningstar: legalbench-full — ALL 162 upstream task dirs of HazyResearch/legalbench fetched verbatim (160 with data: 856 train rows; 16 test splits = 10,219 rows; 2 honestly EMPTY), plus the CUAD enrichment layer for all 38 cuad_* tasks: every {excerpt, Yes/No} row re-joined to CUAD_v1.json expert annotations via whitespace-flexible excerpt location (199 exact + 1 fuzzy / 20 span-unmatched / 8 unknown-contract = 228/228 dispositioned, none dropped) with char offsets + overlapping clause questions + expert spans attached, and an on-row category_audit cross-checking LB’s label vs CUAD’s highlights ON THE EXCERPT (192 agree / 8 SUSPECT / 0 mismatch — flags ride along, labels never rewritten). docclass-merged — 700 rows = CUAD 509 (local staging-export reuse, BT untouched) + MAUD 152 + S-1 39, deterministic fingerprint 5b682f62…. Verification GREEN twice (pack re-upload = independent reproduction): legalbench-full 379/379 files byte-proven via git-blob OID vs Hub tree + aggregates round-trip hash; docclass-merged LFS sha256 local == hub (c8faf0ab6ed8…). Tooling: scripts/datasets/build_legalbench_full_pack.py + scripts/datasets/publish_kanban071.py; build_docclass_merged.py local-first loader. Side effects: restored missing data/maud/classification.jsonl (25,827 rows) and fixed a latent FigureCanvasBase.get_renderer crash in the EDA footer-collision test helper (test was skip-guarded until this session’s CUAD download un-skipped it). Suite 573✓/7 documented posit fails unchanged + 8 new network-free pins tests/test_kanban071_hf_pack.py. Honest gaps in ENRICHMENT_REPORT.json: 20 excerpts unlocatable, 8 unknown contracts, 8 SUSPECT audit rows — flagged on-row, never silently dropped or rewritten. |
KANBAN-053 | v0.19.0 (2026-08-19) | commits d92c30a (phase 1: untrack) + be4bbec (close-out, post-rewrite) | Repo storage optimization COMPLETE (human request) — (1) untracked stale root-level Quarto renders (MESSAGE_BOARD.html + MESSAGE_BOARD_DISCUSSION.{html,md}, ~2.6 MB; files stay on disk, now ignored); (2) history purged with git-filter-repo — every blob of the already-gitignored reports/experiment_log.jsonl (69 blobs ≈ 1.98 GB) + .md (94 blobs ≈ 641 MB) + .phoenix/ (~29 MB) removed → pack 64.4 MiB → 24.7 MiB (.git 80 MB → 25 MB), ~2.65 GB dead blobs gone from every future clone + GitHub storage; (3) legacy gh-pages branch deleted locally + on GitHub (Pages serves /docs from main); (4) AGENTS.md “After every run” snippet corrected (only docs/data is committed). Critical data untouched: experiment log files (already gitignored, local-only), docs/data/* + docs/posit/* (public record, still tracked), data/cuad/master_clauses.csv (GT). Verification: 514 passed / 7 skipped; the 7 test_posit_site failures are pre-existing env gaps (need the local-only reports/experiment_log.jsonl, absent from this checkout). Caveats: all commit SHAs changed — historical git_snapshot SHAs in log/site records are cosmetic labels now; every existing clone MUST be re-cloned; backup of the pre-rewrite history: /tmp/opencode/llm-entity-extraction-backup.bundle (64 MB, delete after the new history is confirmed stable). |
KANBAN-051 | v0.19.0 (2026-08-17) | commit e2650e6 (KANBAN-051 work) + upstream llm-dojo-scoring@v0.3.0; issue #21 CLOSED | key_obligations scoring bottleneck fixed + ContractEval mapping benchmark (issue #21) — (1) disaggregation: run_extraction_eval.py preprocesses key_obligations/termination_clauses through the upstream disaggregate_clause_spans before score_extraction + score_category_presence (stored predicted keeps the raw model output; disaggregated_counts recorded for audit) — merged multi-clause items no longer dilute the 0.6 bipartite match below threshold; (2) reasoning-trace RETAG: contracts_specialist_v33 = v32 + the RETAG RULE (obligation reasoning.entries[].field = canonical CUAD category name, 32-category vocabulary enumerated, key_obbligations misspelling guarded) — the stored corpus had 15,516/33,312 entries umbrella-tagged; (3) upstream v0.3.0: score_category_presence routes each YES/NO category to its reasoning-trace-tagged entry (else the disaggregated spans of its mapped field) and matches by token containment (≥0.7) or embedding (≥ new presence_embedding_threshold 0.7); dep re-pinned. ContractEval mapping benchmark (src/contracteval.py + scripts/reporting/run_contracteval_mapping.py, offline): champion qwen3.7-flash v32 F1 0.164 / F2 0.109 / Jaccard 0.215 / false-nr 0.670 (llama-4-scout v31 F1 0.034) vs ContractEval GPT-4.1 F1 0.641 — and the coverage_bands companion shows the gap is a paraphrase penalty, not missing extraction: verbatim 9.2% vs containment ≥0.7 42.7%. Precision is structurally 1.0 (one-pass extractor). Tests 264 surgical + all runner smokes green. Follow-on: the KANBAN-052 slot was redirected by human directive (2026-08-18) to the directly-mirrored ContractEval task (issue #22) — the quote-faithful v34 arm idea from this card’s memo is folded into that task’s prompt-iteration phase. |
KANBAN-009 | v0.19.0 (2026-08-17) | scripts/reporting/rescore_manifests.py + tests/test_rescore_manifests.py (2 green) + reports/same_scorer_scores.json (tracked, covers v13–v23); issue #7 closed | Score-drift hygiene verified complete — issue #7 closed — the same-scorer rescore pipeline already satisfies both deliverables: (1) extends beyond the 50-doc series — rescore_manifests.py accepts arbitrary repeatable --manifest args (--auto-50 is only the convenience flag for the seed-42 series), (2) reports/same_scorer_scores.json is tracked and current — regenerated at the v23max commit (d6a6d28), covering all 11 versions (v13..v23, 50 docs each). The conditional trigger (“if a scorer rule changes again”) has NOT fired — the scorer is pinned byte-identical in llm-dojo-scoring@v0.2.0 (KANBAN-044). Tests green; the 50-doc manifests themselves are gitignored (present only locally), so --auto-50 regenerates from the local manifests and the committed JSON is the artifact record. |
KANBAN-008 | v0.19.0 (2026-08-17) | decision already recorded in README “Recommended production configuration (extraction)” §87-107 + AGENTS.md “Production decision (KANBAN-008)” §768-773 + V16_PROPOSITION.md §15.1 + memo docs/memos/contracts_specialist_v23.md; issue #6 closed | v23×max ko arm — production decision recorded: SPLIT — the recommended production config is the overall arm (reasoning_effort=none, default; champion line contracts_specialist_v31 0.8737 @ full-509), with the ko arm (--reasoning-effort max, v23×max: ko 0.8510, 0 parse errors, lowest ellipsis 18.7%) as a documented opt-in for compliance/covenant-heavy reviews at 2.6× cost — NOT v19×max (1/50 parse error + worst overall 0.9135). Max reasoning buys +2.2pp ko at 2.6× cost and −1.5pp overall. Documentation was already complete in the tree (README/AGENTS/V16_PROPOSITION §15.1/memo); this pass verified the evidence and closed the card + issue #6. |
KANBAN-049 | v0.19.0 (2026-08-17) | scripts/reporting/monte_carlo_gepa.py + reports/monte_carlo/gepa-champion-contender-*.md + tests/test_monte_carlo.py + memos/monte_carlo_gepa.md; issue #17 closed | Monte Carlo simulations folded into the GEPA loop as a champion-contender selection layer + half-corpus effectiveness pilot (human directive, issue #17) — monte_carlo_gepa.py selects the MC champion contender per task via corpus-wide paired-bootstrap prompt ablation (mean Δ, 95% CI, P(win) ≥ 0.9 + CI-excludes-zero noise-floor contract) + committee-voting robustness @ K + a 25/50/75/100% document-count sweep. Pilot (qwen3.7-flash, seeded 42): subtype full corpus → sorter_v15 (0.9506, tied with v13 at 7 wins, tiebreak by accuracy), the seeded 50% sample (254 shared docs) recovers the same champion, 25% (127 docs) → plateau (P(win) 0.021, CI touches zero) → the sample-efficiency boundary is between 25% and 50%; docclass → plateau at every fraction (v6-vs-v3 full-676 +0.0015, CI [+0.0000,+0.0044], P(win)=0.637 — the gate correctly refuses to crown a noise-floor delta). Tests grow to 14 (13 passed + 1 skip; gepa scenario smoke + clear-winner selection + committed-report drift guard). Memo memos/monte_carlo_gepa.md; CHANGELOG [Unreleased] Added; issue #17 closed. |
KANBAN-050 | v0.19.0 (2026-08-17) | SCORING.md + wiki/Scoring.md + docs/slides/12-task-aware-and-robustness-metrics.md + README/src-README/AGENTS/wiki-Architecture/wiki-FAQ/slides-02-04/deck-index/data-judgments-README | Scoring documentation refreshed to the current pipeline state (human request) — SCORING.md rewritten (new §0 maps the llm-dojo-scoring@v0.2.0 outsourcing: six re-export shims + dojo_config/dojo_compat + package module map + CLI) and all new metrics documented: subtype metrics (strict/equiv, per-family, failure modes, CIs), docclass hierarchical metrics (doc_type/subclass + equiv + per-subclass + input modes + failure modes), task-aware dispatcher (score_task: MAUD consideration strict/equiv, LegalBench binary P/R/F1, multiclass macro/micro, court opinions, chained 0.25/0.75), judge calibration, chained ablation, cost scoring, failure-mode taxonomy, Monte Carlo robustness (KANBAN-048), run-sink/tracing; wiki/Scoring.md mirrored; new slides deck 12; CHANGELOG [Unreleased] Changed; docs-only, no code change (18+79 surgical tests green, site data current, SCORING.md≡wiki/Scoring.md); wiki pushed via ./wiki/sync-wiki.sh. |
KANBAN-048 | v0.19.0 (2026-08-16) | src/monte_carlo.py + scripts/reporting/monte_carlo_{corpus,ensemble,prompt_ablation,failures,exemplars,verify}.py + reports/monte_carlo/* + tests/test_monte_carlo.py | Monte Carlo simulation suite ported from the RVL-CDIP-classifier (issue #17, KANBAN-048) — zero-spend what-if analysis over the joint reasoning corpus (17,691 rows: 16,162 subtype + 1,442 docclass + 80 chained + 7 sorter, 99.7% with reasoning; corpus.jsonl gitignored, scenario outputs tracked). Scenarios + headline findings: ensemble voting — subtype 0.9209 → 0.9513 @ K=25 (weak lever, ~4pp ceiling), doc_type saturated 0.9928 (no gain); confidence-gated escalation — subtype +0.44 pp @ alpha 0.15 to a 0.95 model (1.3x cost); docclass escalation loses (baseline > 0.95); paired-bootstrap prompt ablation — 156 subtype + 12 docclass pairs; sorter_v10/v11 vs v3 +14.1 pp P(win)=1.000, docclass v5 loses on the diag-30 slice; failure pipeline — observed 0.2374% failure rate, fallback pass is the lever (0.004% vs 0.202% at max_tries=1), ~0 failures at 320K; exemplar mining — 6 subtype + 4 docclass near-miss exemplar appendices (development→license +25.0 expected flips). Memo memos/monte_carlo_robustness.md; 12 tests green. No model spend. |
KANBAN-047 | v0.19.0 (2026-08-16) | commit 98d4ed7 + llm-dojo-scoring @ v0.2.0 (upstream 2a7e37b, tag v0.2.0) | llm-dojo-scoring covers the additional document hierarchy (issue #19) — new llm_dojo_scoring.tasks module with a score_task() dispatcher + task-aware normalization: MAUD (merger-agreement doc-class + consideration-type subclass strict/equiv scoring + per-question classification), LegalBench (binary Yes/No exact-match + per-class + P/R/F1), multi-classification (macro/micro + confusion), court opinions (court_opinion doc-class), chained evaluation runs (chained_composite/chained_summary, sorter+extractor weighted 0.25/0.75); config.py task registries (DOC_CLASS_KEYS, MAUD_CONSIDERATION_*, LEGALBENCH_BINARY_LABELS, COURT_OPINION_CLASS, TASK_KINDS); 10 upstream tests (144 suite green), README task-coverage section; dep re-pinned in pyproject.toml + requirements.txt. Issue #19 closed. CHANGELOG [Unreleased] Added |
KANBAN-046 | v0.19.0 (2026-08-16) | commit 98d4ed7 | Full Phoenix local trace sink documented + cost-efficiency configuration cemented (issue #18) — new wiki/Phoenix-Tracing.md (linked from Home + _Sidebar; wiki pushed): the local Arize Phoenix sink (Langfuse-primary resolution with Phoenix fallback via resolve_tracer, OTLP spans, SQLite, discard-by-delete) + the resume/checkpoint/queue/cache configuration (ManifestStore resume + header contract, append-only experiment-log checkpoint, HITL annotation queue, embedding-cache reuse + usage accounting, --dry-run/assert_production_run/--research-funding-key gates). .env.example Phoenix section + AGENTS.md pointer + drift-guard test test_env_example_documents_phoenix_sink. Issue #18 closed. CHANGELOG [Unreleased] Added |
KANBAN-045 | v0.19.0 (2026-08-16) | scripts/eda/explore_pipeline_sources.py + data/eda/{maud,s1,docclass,legalbench}/ + tests/test_pipeline_sources_eda.py | Full EDA suites on the new pipeline sources (human request) — one reproducible script (--source all\|maud\|s1\|docclass\|legalbench, --no-figures) writes per-source data/eda/<source>/{report.md, findings.md, figures/} for every post-CUAD source: MAUD (152 merger agreements, 54.1M chars, median 338k chars — ALL 152 over the 90k chunk window; consideration-GT all_cash 57 / other 57 / all_stock 24 / mixed_cash_stock 13 / election 1; 25,827 per-question rows across 22 families / 7 categories, MAE 8,548 rows largest) + S-1 corporate records (15 EDGAR exhibits, EX-3.1×4 / EX-3.2×3 / EX-4.x; subclasses articles_of_incorporation 8 / rights_instrument 6 / bylaws 1) + merged doc-class (676 = CUAD 509 + MAUD 152 + S-1 15; GT-other gap cluster 57, subclass-None 509; doc_type contract 75.3% / merger_agreement 22.5% / corporate_record 2.2%) + LegalBench (hearsay train 5 / test 94 + 10 CUAD subtask 6-row surfaces). Reports regenerate byte-identically (reproducibility pinned by test). 3 tests green; CHANGELOG [Unreleased] Added entry. |
KANBAN-044 | v0.19.0 (2026-08-16) | commit 1f9881b (integration) + close-out commit; llm-dojo-scoring @ v0.1.2 | llm-dojo-scoring integration completed (KANBAN-044) — scoring/error-analysis/export library as the pinned single source shared with llm-mailroom: dep llm-dojo-scoring @ git+…@v0.1.2 in pyproject.toml + requirements.txt; src/dojo_config.py maps config/taxonomy.yaml → package Settings (incl. ambiguous_band→tuple / partial_gt_fields/containment_fields→set coercion + cost_models list-form + load_env() first); src/dojo_compat.py (classify_failure None-on-ok); the 6 local scoring modules become thin re-export shims (llm-mailroom pip install -e . imports unchanged); export_experiment_results.py re-exports llm_dojo_scoring.export; export_sweep_results.py STAYS local (KANBAN-040 reference-format Notes contract); dojo-analyze/dojo-export/dojo-sync CLIs verified + python -m llm_dojo_scoring.cli entry dispatch added upstream (v0.1.0→v0.1.2, pushed + tagged). All 11 dojo-integration tests green (3 wrong expectations corrected to the real contract); full suite green. Memo memos/llm_dojo_scoring_integration.md; CHANGELOG [Unreleased] Added/Changed. |
KANBAN-043 | v0.19.0 (2026-08-16) | runs llama-4-scout_sorter_v13_subtype_langfuse (0.8880 @06:36) + gpt-4o-mini_sorter_v13_subtype_langfuse (0.9312 @11:50); manifests data/manifests/{llama4scout,gpt4omini}_sorter_v13_509.jsonl | Champion-sweep extension — llama-4-scout + gpt-4o-mini on sorter_v13, full-509 (human directive) — two funded full-corpus subtype evals (reasoning medium, temp 0.1, seed 42, research-funding key, Langfuse-primary tracing): llama-4-scout subtype 0.8880 (equiv 0.9077, CI [0.8605, 0.9136], 57 fails) and gpt-4o-mini 0.9312 (equiv 0.9352, exact 0.9961, CI [0.9096, 0.9528], 35 fails). Both 509/509 in the log; the KANBAN-036 sweep table + workbook absorb them; experiment-log md regen (195 records). CHANGELOG [Unreleased] entry. |
KANBAN-036 | v0.19.0 (2026-08-16) | verified sweep table in reports/experiment_log.{jsonl,md} + reports/sheets/Sorter_Model_Sweep_Results.xlsx (22 rows) | Sorter subtype model sweep — 8 models on the champion sorter_v13 prompt, full-509, verified table — deepseek-v4-pro 0.9528 (only model to beat the qwen champion; cross-model significance NOT claimed), qwen3.7-flash 0.9430 (champion), gpt-4o-mini 0.9312, deepseek-v4-flash 0.9332 (canonical clean rerun), gpt-5-nano 0.8978, llama-3.3-70b-instruct 0.8900 (canonical), llama-4-scout 0.8880, gpt-4.1-nano 0.8782. Duplicate-run caveats verified against the log; workbook regenerated covering all models. |
KANBAN-033 | v0.19.0 (2026-08-16) | run_langfuse_docclass_eval.py + merged dataset data/datasets/docclass_merged.jsonl (676, fp 5602b71f) + sorter_docclass_v6 + src/tracing.py | MAUD + EDGAR S-1 corporate-record wiring + hierarchical doc-class eval task COMPLETE — MAUD utilized dataset (152 contracts + 25,827 per-question rows), new merger_agreement class (7-class sorter, shared 6-class surface untouched), run_langfuse_docclass_eval.py, EDGAR S-1 exhibit ingestion (15 corporate records); prompt iteration CLOSED at v6 = docclass text champion (full-676 A/B with noise control: v6 0.8935 vs v3 0.8905, subclass +1.19pp, 0 regressions); scoring depth (bootstrap CIs, per-subclass tables, subclass_accuracy_equiv, input-mode counts); tracing Langfuse-primary. Follow-on reserved: full-pages vision benchmark + MAUD/S-1 PDF retention (KANBAN-038). Memo memos/docclass_v6.md. |

|—|—|—|—| | KANBAN-042 | v0.19.0 prep (2026-08-16) | runs deepseek-v4-pro_sorter_v13_subtype_langfuse + meta-llama-llama-3.3-70b-instruct_sorter_v13_subtype_langfuse; reports/sheets/Sorter_Model_Sweep_Results.xlsx (8 rows) | Sorter model-sweep expansion — deepseek-v4-pro + llama-3.3-70b-instruct on champion sorter_v13, full-509, research-funding key — two funded full-corpus subtype evals (reasoning medium, temp 0.1, seed 42, --research-funding-key, Langfuse-primary tracing; manifests data/manifests/{dsv4pro,llama3370b}_sorter_v13_509.jsonl): deepseek-v4-pro subtype 0.9528 (CI [0.9332, 0.9705]), exact 0.9961, equiv 0.9548, 24 fails, 6.97M tokens ≈ $3.15 est. — highest in the sweep to date (vs qwen champion 0.9430; cross-model significance NOT claimed); llama-3.3-70b-instruct subtype 0.8782 (CI [0.8487, 0.9057]), exact 0.9941, equiv 0.8998, 62 fails, 6.66M tokens, cost None (OpenRouter-billed). Both n_ok=509, full reasoning + per_subtype; LangSmith 429 noise non-fatal. Sweep workbook regenerated (8 rows, Notes for both) + copied to ~/Downloads; deck slide 16 updated (llama sorter run table; “no llama sorter runs” note superseded) + recopied; experiment-log md regenerated (178 records). CHANGELOG [Unreleased] Added; uncommitted — commit left to the human | | KANBAN-041 | v0.19.0 prep (2026-08-16) | reports/sheets/contract_specialist_v32_and_sorter_v14_deck.xlsx (regenerated, 19 slides) + reports/sheets/llama4scout_v31_extraction_langfuse.json | Slides deck regenerated per human request — qwen sorter lineage v3→v13 + llama runs — deck 16 → 19 slides: (14) qwen lineage summary v3→v13 (best full-surface per version, 509-doc chain v5 0.8585 → v6 0.9312 → v8 0.9018 → v9 0.9175 → v12 0.9293 → v13 0.9430), (15) all 30 qwen v3→v13 runs (degraded rows kept: v11 0.0000, v13 0.7741), (16) llama runs — exhaustive search (log 0 hits, Langfuse model+session filters, Braintrust 403, LangSmith HEARSAY-only) found NO llama sorter runs; the ONLY llama run is llama-4-scout × contracts_specialist_v31 EXTRACTION in Langfuse llm-dojo (509 traces / 20 scored, truncated; overall 0.6627 vs qwen v31 0.8737, n=20 not comparable) — fetched via langfuse-cli with rate-limit backoff, record in llama4scout_v31_extraction_langfuse.json (deck --llama-json). Lineage slides switched to raw-list loader (load_records_list — name→record dedup dropped same-name reruns). Verified: 19 slides banner/footer/landscape fit-to-page, values cross-checked (v6 0.9312@509, v13 0.9430@509, llama 0.6627); copied to ~/Downloads. CHANGELOG [Unreleased] Added entry; uncommitted — commit left to the human | | KANBAN-040 | v0.19.0 prep (2026-08-16) | scripts/reporting/export_sweep_results.py + reports/sheets/Sorter_Model_Sweep_Results.xlsx | Sorter model-sweep workbook (per human request, reference-format) — xlsx in the EXACT format of Sorter_Experiment_Results.xlsx (114 columns + trailing Notes; Eval Results + Codebook sheets; header/styling/freeze/autofilter reused verbatim from export_experiment_results.py; 114 shared headers byte-identical to the reference). Rows = every subtype_classification run of the champion prompt (sorter_v13, --prompt overridable), chronological: DEGRADED first qwen v13 run (93 connection errors, flagged in Notes) → champion clean rerun (qwen3.7-flash 0.9430) → gpt-5-nano 0.8978 (KANBAN-035) → deepseek-v4-flash + gpt-4.1-nano 1-doc smokes → deepseek-v4-flash 0.9253 full-509 (KANBAN-036; gpt-4.1-nano full-509 still pending). Per-subtype strict/equiv + cell sizes, failure modes, CI, tokens/cost all populate from the log records. Tests tests/test_export_sweep_results.py (4, network-free). CHANGELOG [Unreleased] Added entry; card archived; commit left to the human | | KANBAN-039 | v0.19.0 prep (2026-08-16) | scripts/reporting/export_slides_deck.py + reports/sheets/contract_specialist_v32_and_sorter_v14_deck.xlsx | Slides-style xlsx deck export (per human request) — ONE Google-Slides-formatted workbook, 16 sheets = 16:9 slides (landscape fit-to-page, dark banner + footer, stat cards): Part A = contracts specialist v32 510-full-clean (overall 0.8807 CI [0.8689, 0.8913], field_presence 0.9701, schema_valid 1.0, verified_precision 0.9799, per-field + error decomposition + entity-list + MAE/R² diagnostics, v31 comparison with memo verdict) + full extraction codebook (9 fields/types/scoring class + rubric thresholds); Part B = sorter v14 subtype (exact 0.9961, subtype 0.9371 CI [0.9155, 0.9568], equiv 0.9411, 32 failures by mode + examples, per-subtype strict/equiv table from row-level results) + full sorter codebook (25 subtypes, 4 equivalences, failure modes, scoring rules); Part C = docclass v5 diag-30 bonus (doc_type 0.8333 / subclass 0.5263) + sources. Script reads the experiment log + taxonomy.yaml + sorter_agent at build time; verified by workbook reload + value cross-check; CHANGELOG [Unreleased] Added entry in the same working tree (commit left to the human) | | KANBAN-037 | v0.19.0 prep (2026-08-16) | site/ → docs/posit/ + tests/test_posit_site.py | Posit Cloud integrated portal (complementary to the SPA) — Quarto website (site/ sources) rendering into docs/posit/ (the SAME docs/ tree GH Pages serves; zero Actions — Pages deploys from branch, Posit Cloud deploys via quarto render site + publish). Three integrated sections at one URL prefix: experiment log (generated from reports/experiment_log.jsonl by site/_pre-render.py with the canonical render_full_log renderer; full run index + per-run metadata/scores/tokens/diagnostics + per-run deep links into ../index.html#/run/{n}), kanban board (live MESSAGE_BOARD.md copy), discussion board (live MESSAGE_BOARD_DISCUSSION.qmd copy, agent colors preserved); portal↔︎SPA navbar links both ways. Custom blue→teal gradient theme (cosmo light / darkly dark, toggle, navbar, search, TOC). Rendered output committed; .gitignore Quarto/Posit section. Repairs: 6 discussion entries missing ::: closers (append-only preserved, pinned by test); KANBAN-037/038 number-collision reconciled. 8+1 tests green (incl. skip-if-no-quarto determinism); SPA browser audit green. Docs: site/README (new), docs/README, README, wiki/Site.md | | KANBAN-035 | v0.19.0 (2026-08-16) | gpt-5-nano × sorter_v13, full-509 | GPT cheapest-model benchmark on the sorter subtype surface — openai/gpt-5-nano ($0.05/M prompt + $0.40/M completion — smallest & cheapest GPT on OpenRouter) on the champion sorter_v13 prompt, full 509-doc corpus, reasoning medium, temp 0.1, 0 errors, Phoenix sink, --research-funding-key (human directive). strict 0.8978 (CI [0.8703, 0.9234]) vs the qwen3.7-flash champion 0.9430 = −4.5pp — far outside the ±0.006 noise band → nano does NOT match the champion; a cost-floor frontier arm only (equiv 0.9018, doc_type 0.9941; 52 fails: 41 family_confusion / 6 other_fallback / 3 function_over_form / 2 equivalent_family). Run cost ≈ $0.48 (6.94M tokens). Run gpt-5-nano_sorter_v13_subtype_langfuse; manifest data/manifests/gpt5nano_sorter_v13_509.jsonl | | KANBAN-034 | v0.19.0 (2026-08-15) | --research-funding-key gate (close-out commit) | Externally-funded OpenRouter key behind a production-only flag — RESEARCH_FUNDING_OPENROUTER_API_KEY in .env reachable ONLY via --research-funding-key on all 10 eval runners (default key untouched); src/env_utils.py resolve_openrouter_key() / assert_production_run() / add_research_funding_flag() — the gate HARD-REFUSES dry-runs and pilot-scale samples (<100 rows, or less than the full dataset when smaller) with SystemExit before any LLM call and prints a funding banner; 12 new tests + all runner smoke suites green. Code swept into a6964c8; changelog/board/docs in the close-out commit | | KANBAN-030 | v0.19.0 (2026-08-15) | commit e0f758e | Contract-specialist v1..v16 archived to src/prompts_archive.py (FROZEN pre-documentation lineage, imported back into prompts.py) — file 227 KB → 155 KB (−32%), all 32 version strings byte-identical, every version key resolvable; archive rule pinned by tests (never edit an archived constant — a change = a new version key); 384 tests green | | KANBAN-029 | v0.19.0 (2026-08-15) | contracts_specialist_v32 | effective_date rule_contradiction repair — CLEAN full-510 A/B (495-row intersection): v32 0.8799 vs v31 0.8746 = +0.0053, CI [−0.0052, +0.0159], P(Δ≤0)=0.1715 — INSIDE the ±0.011 noise band → LOGIC REPAIR, v31 stays champion (the first candidate run’s +0.0115 was survivorship bias from 52 transient errors); effective_date field +0.0171 (23 improved / 11 regressed, 16/23 on the diagnosed target cluster) → v32 = effective_date field specialist; never-null over-fire cluster deterministic → banked as the v33 carve-out (stated-FULL-date requirement). Memo memos/contracts_specialist_v32.md | | KANBAN-026 | v0.19.0 (2026-08-15) | LegalBench hearsay + CUAD-subtask prompt-iteration series (arms 1-7) | Arm 1: v0 baseline @5 = 1.0 (saturates the train surface). Arm 2: 94-row official test split wired (--test). Arm 3: Langfuse dataset mirror + LangSmith retargeted to project HEARSAY. Arm 4: legalbench_task_v1 @94 = 0.8511 (80/94) vs v0 0.7766–0.7872 (directional win, P=0.0905). Arm 5: legalbench_task_v2 @94 = 0.8830 (83/94) vs v1-control 0.8617 (logic repair, P=0.3345). Arm 6: legalbench_task_v3 7/7 tasks 42/42 = 100% vs v0 36/42. Arm 7: legalbench_task_v4_* CUAD subtask series — hygiene base + CRE (0.8333→1.0) + CNTS (→1.0) operative rules, 5 subtasks at ceiling. Memos memos/legalbench_task_v1/v2/v3/v4.md; runs ..._test + ..._subtask_* in the experiment log | | KANBAN-024 | v0.19.0 (2026-08-15) | LangSmith tracing + live-error analysis | LANGSMITH_API_KEY + LANGSMITH_TRACING=true + LANGSMITH_PROJECT=llm-mailroom in .env/.env.example; AGENTS.md env docs; verified every LangChain LLM call auto-traces to the project, coexisting with Braintrust’s setup_langchain. Live-error analysis (langsmith SDK): 100 runs / 14d, all 10 errors are OpenRouter-exported OpenRouter Request spans — qwen via Alibaba 429, ~230ms burst rate-limit, OpenRouter failover recovered 90/100. Braintrust span ingestion failing on plan limit (num_log_bytes_calendar_months, org UW-Madison-Capstone) — LangSmith is the reliable span sink. Docs-only; 359 tests green | | KANBAN-027 | v0.19.0 (2026-08-15) | d6c6c9d | Repository streamlining + navigation pass — README Table of Contents + Layout-tree repair (src modules mis-nested under wiki/, stale script list + test counts 223→375, prompt tables to sorter_v12/contracts_specialist_v31, BRAINTRUST_LOGGING conditional documented); src/README.md (+braintrust_logging/eval_shims/master_labels/metrics), memos/README.md (table bug + v28/v30/v31/sorter_v10_v11 rows), scripts/README.md (all runners + eda/) refreshed; scripts/backfill_cost_estimates.py nested under scripts/reporting/ (live refs updated, CHANGELOG history untouched). Docs-only + safe nesting; 375 tests green, site render audit clean, no new release-gate failures. CHANGELOG [Unreleased] Changed entry in the same commit | | KANBAN-032 | v0.19.0 (2026-08-15) | sorter_v14 (rule 30) | Sorter marketing-title strengthening arm — LOGIC REPAIR, NOT a win (v13 stays champion). Full-509 A/B (fp c2341957…, seed 42, temp 0.1, reasoning medium): v14 0.9371 vs v13-clean 0.9430 = −0.0059, paired CI [−0.0177, +0.0059], P(Δ≤0)=0.8765 — inside the ±0.006 noise band, negative direction. Marketing cell 14/17 → 16/17 (Audible + PACIRA recovered, rule-30 reasoning pinned; Zounds resists even its own literal example). Flagged counterfactual FIRED: Playboy “CONTENT LICENSE, MARKETING AND SALES” regressed license→marketing — carve-out (a) too narrow (cited only the exact “Content License Agreement” phrase) → banked v15 lesson: widen to any license-PRIMARY title. 4/6 other regressions are untouched-family noise. Memo memos/sorter_v14.md | | KANBAN-031 | v0.19.0 (2026-08-15) | sorter_v13 (rule 29) + Phoenix sink | Sorter maintenance title-wins arm — AGGREGATE WIN. Full-509 A/B (fp c2341957…, seed 42, temp 0.1, reasoning medium): v13 0.9430 vs the v12 rerun control 0.9293 = +0.0137, paired CI [+0.0020, +0.0255], P(delta<=0)=0.0090 — OUTSIDE the ±0.0059 identical-prompt noise band → new aggregate champion (v9 → v12 → v13). Maintenance cell 30/34 → 34/34 (SUNTRONCORP/WELLSFARGO/PRIMEENERGY/AtnInternational recovered, rule-29 reasoning pinned); recovered 8 / regressed 1 (ImperialGarden = pre-existing rule-24 outsourcing variance flip, NOT rule 29 → banked v14 lesson). 0-risk counterfactual (34/34 maintenance-titled docs GT maintenance). First v13 run degraded (93/509 connection-error defaults) → clean rerun before any claim. run_langfuse_subtype_eval.py selects Phoenix tracing by default (tracing_backend="phoenix"). Memo memos/sorter_v13.md | | KANBAN-023 | v0.19.0 (2026-08-15) | sorter_v12 (rule 28) | Sorter strategic_alliance title-wins arm — the first banked KANBAN-013 cluster. Full-509 A/B (fp c2341957…, seed 42, temp 0.1, reasoning medium): v12 0.9234 vs the clean v9 rerun 0.9175 = +0.0059, paired CI [−0.0098, +0.0216], P(Δ≤0)=0.251 — INSIDE the noise band → logic repair, NOT an aggregate win (identical-prompt v9 rerun itself moved +0.0059); the strategic_alliance cell is deterministically fixed 28/32 → 31/32 (Iovance/Giggles/Adaptimmune recovered with rule-28 reasoning pinned; Intricon remains — license carve-out didn’t override the substance read), recovered 9 / regressed 6 (none rule-28-driven; 2 equiv-recovered). v9 stays aggregate champion; v12 joins the frontier as the strategic_alliance field specialist. NOTE: the first v9 @509 control rerun was degraded (42 transient errors) — replaced by a clean rerun. Memo memos/sorter_v12.md | | KANBAN-028 | v0.19.0 (2026-08-15) | master_clauses CSV + loader | Master ground-truth CSV added to the repo + repo-local default — data/cuad/master_clauses.csv (510 CUAD contracts × 40 normalized -Answer columns) committed; DEFAULT_MASTER_LABELS resolves to the repo-local copy first (MASTER_LABELS_CSV env wins; sibling llm-mailroom path fallback); loader normalizes the stray-space Notice Period To Terminate Renewal- Answer header variant so that category’s answer loads (was silently dropped). 377 tests green. KANBAN-027 = ATOM’s separate completed repo-streamlining card | | KANBAN-025 | v0.19.0 (2026-08-15) | run-sink swap commit | Run sink = Langfuse + LangSmith; Braintrust logging OFF by default (BRAINTRUST_LOGGING=disabled) — the run_*_eval.py runners skip braintrust.Eval and use the local scoring loop (src/eval_shims.py), run_langfuse_*_eval.py become the documented primary path. Verified: disabled-path pilot (tracing_backend=none, langsmith=True, trace tree in LangSmith, zero Braintrust 400s). 365 tests green | | KANBAN-013 | v0.19.0 (2026-08-15) | sorter_v10 + sorter_v11 | Sorter tail-sampling iteration: the “1-off long tail” plateau reading superseded by cluster analysis — marketing cell 0.5/10 (243) & 7/17 (509), worst family on both surfaces, unchanged since v6. v10 = rule 26 MARKETING TITLE WINS (title-wins doctrine mirror of R23/R24); v11 = rule 27 AFFILIATE IS NOT MARKETING (fixes the measured rule-26 over-fire). 243-doc same-surface A/B: champion rerun 0.9259→0.9300 (±1 doc noise floor); v10/v11 0.9342 (equiv 0.9424), paired CI [−0.0247, +0.0165], P=0.710 — inside the band → logic repair; v9 stays champion. Rule-driven: 4 deterministic recoveries + 2 affiliate restorations, 1 R27-wording regression (equiv-recovered); marketing cell 0.5→0.8. Memo sorter_v10_v11.md; banked lessons → KANBAN-023; issue #11 closed | | KANBAN-022 | v0.19.0 prep (2026-08-15) | commits f417227 (wiring) + live-run follow-up | LegalBench hearsay task wired end-to-end: mailroom-lb-hearsay synced from the ACTUAL task data (5 train rows, binary Yes/No, 5 slices, CC BY 4.0, Neel Guha) + classes manifest data/legalbench_classes.jsonl; upload_text_dataset now inserts deterministic content-addressed ids (_deterministic_record_id — Braintrust’s insert() assigned fresh UUIDs per call, so reruns APPENDED duplicates, observed 2×5; reruns now upsert); run_langfuse_classification_eval.py gains --prompt-mode task (mirror previously hardcoded the sorter path and dropped the row’s prompt) — hearsay/LegalBench tasks trace into llm-dojo with one legalbench_task observation per row; BaseAgent._call_llm captures usage/cost (task-mode records previously had tokens 0 / cost 0); LegalBench-task docs updated with the actual task data; README Credits (LegalBench, CUAD/The Atticus Project, MAUD, GEPA, LangChain/Braintrust/Langfuse); first benchmark: exact_match 1.0 (5/5, per-class 1.0/1.0) on qwen3.7-flash × legalbench_task_v0, replicated on the Langfuse mirror (5 traces/session verified); 10 new tests, 359 total green; issue #13 closed (reopened for the live run, re-closed). NOTE: Braintrust ORG at monthly log-bytes plan limit — experiment row data does not upload to Braintrust until billing is addressed | | KANBAN-021 | v0.19.0 prep (2026-08-15) | contracts_specialist_v31 | token-efficiency refactor (−8.0% = 2,679 chars, 7,700 system tokens −5.7%) with every operative constraint preserved; full-corpus 509-doc A/B: v31 0.8737 vs v28 0.8622 (+0.0116, CI [+0.0005, +0.0236], P=0.021) — Pareto win; re-baseline: 50-doc surface overstates champion ~6pp (0.9228→0.8622); 7,250-entry reasoning-trace corpus; unblocked after a new OpenRouter key; memo contracts_specialist_v31.md | | KANBAN-020 | v0.18.0 (2026-08-15) | contracts_specialist_v29 + contracts_specialist_v30 + runner chunking + GEPA agent | Follow-up arm with a noise-floor control: identical-prompt rerun band ±0.03 overall (v28 0.9228 → rerun 0.8935) — v29 (CoC-definition carve-out, Ediets 0.692→0.769) and v30 (chunk-mode scalar quoting) ship as UNMEASURED logic repairs inside the band; v28 remains champion (re-validated +0.0448 vs v26, P=0.004). Per-span diffs: Ediets rule-driven (fixed), 3 others noise; renewal_terms = 1 doc; Gridiron = 1-off; run_extraction_eval.py gains --chunked + confound warning; prompt-engineer agent now runs the full GEPA reflective loop (frontier, noise floor, Pareto selection); memo contracts_specialist_v30.md | | KANBAN-004 | v0.18.0 (2026-08-15) | contracts_specialist_v27 + contracts_specialist_v28 | key_obligations span residual: sim-matrix diagnosis (wrong-span at sentence level in multi-requirement family sections; ~60–70% NEAR 0.35–0.59) → multi-item family-section rule (v27) sharpened (v28: operative-vs-definitional + additive re-scan). 50-doc chunked A/B: v28 0.9228 vs v26 0.8780 (+4.48pp, CI [+0.0094, +0.0907], P=0.004); ko +11.4pp, 20 recovered / 4 regressed; term +0.040; tokens +6.7%. Truncation confound on sample5 surface documented; memo contracts_specialist_v28.md; issue #3 closed | | KANBAN-017 | v0.18.0 (2026-08-15) | contracts_specialist_v26 | term_length containment fixed through TWO iterations: v25 additive-prefix + worked example recovered Ediets but leaked the example template (Ritter/Phasebio containment 0.7059/0.2222); v26 (opener variants + “never reuse the instructions’ wording”) — overall 0.9447 best of the arm (v23 0.9366 / v24 0.9336 / v25 0.9154), term_length 1.0000, all three term docs containment 1.0, no leakage | | KANBAN-018 | v0.18.0 (2026-08-15) | commit 1fcc734 | prompt-engineer agent shipped (.opencode/agents/prompt-engineer.md, mode all, verified via opencode agent list): master diagnostic evaluator & prompt engineer — sole role reviews all traces/reasoning/failures/errors/results and produces data-backed prompt mutations (new version keys); full diagnose→root-cause→mutate→same-surface A/B→land loop, failure taxonomy, plateau/overfit doctrine (clusters not outliers, generalization test, evidence floor), board+CHANGELOG+memo close-out; AGENTS.md “Agents (this repo)” section | | KANBAN-016 | v0.17.0 (2026-08-15) | commit 6f77615 | Contracts specialist v24: required per-field reasoning trace (schema reasoning object first, chunked merge unions entries, never scored, rides into log + Langfuse) + metrics-aligned format discipline (canonical duration phrase leads term_length, plain currency contract_value, ISO dates — format only, no master-CSV leakage). A/B seed 42 n=5: overall 0.9336 vs 0.9366 (noise), key_obligations +10.2pp, reasoning 5/5 rows, tokens +2.5%; term containment dip on 1 doc documented; issue #12 closed | | KANBAN-015 | v0.17.0 (2026-08-15) | commits 91392ea + follow-up | Extraction regression diagnostics shipped end-to-end: R² + MAE tracked for dates/durations (date_r2/duration_r2 = 1 − SS_res/SS_tot, negative kept), money MAE (USD), span-count drift (MAE + signed mean), field error decomposition, pair counts — all in scores.diagnostics (JSONL + dedicated md-log section + GH Pages run-detail diagnostics card); src/metrics.py + src/master_labels.py (curated CSV preferred, raw clause-text fallback), --master-labels/MASTER_LABELS_CSV on both extraction runners; scoring-method slide decks docs/slides/ (7 decks + index, worked examples incl. real pilot block pilot_diag_v22_sample2); SCORING.md §4 + README + AGENTS.md + wiki updated; 31+4 new tests, 337 total green | | KANBAN-014 | v0.17.0 (2026-08-15) | commit 2fe4103 | Full-corpus CUAD EDA shipped: scripts/eda/explore_cuad.py + data/eda/{report.md,findings.md,figures/01–10} all git-tracked. Headlines: median 33,425 chars (max 338,211), 17.1% over the 90k chunk window, key_obligations scope mean 16.0 spans/doc (49 null docs), 131 docs with [***] redaction markers, Anti-Assignment+Change Of Control 98% co-occurrence. Follow-on (2026-08-16, commit 6128722): EDA figures 01–10 regenerated with CUAD dataset citations + self-contained data paths — _add_citation() footer on every figure, figure 01 retitled “CUAD contract subclass distribution (25-family taxonomy, n=509)”, repo-local data/cuad_pdfs/CUAD_v1.json corpus path, verified experiment-log subtype distribution (SUBTYPE_FALLBACK sums to 509), title-derived per-subtype lengths (503/510 matched); refined further by 0a75c76 (dedicated citation footer band) | | KANBAN-012 | v0.16.0 (2026-08-13) | commit 6697ea9 | sorter_v9 A/B landed: v9 wins (+2.88pp strict, 0.9259) — promotion/outsourcing/customization-schedule clusters eliminated, 25→18 fails, v6→v9 +5.8pp; ~0.93 practical plateau → follow-on KANBAN-013; issue #10 closed | | KANBAN-010 | v0.16.0 (2026-08-13) | commit 25aa942 | Resolved by decision — site cost telemetry intentionally REMOVED (costs meta + per-run cost gone from docs/data/); “restore cost accounting” superseded; issue #8 closed | | KANBAN-003 | v0.16.0 prep (2026-08-12) | commit cbb5b93 | sorter_v7 250-doc A/B landed: v7 wins (+0.82pp strict, 0.8765) — promotion cluster fixed; issue #2 closed on sweep | | (user-run, no card) — sorter_v8 A/B | v0.16.0 prep (2026-08-12) | commit 43ef2ab | v8 wins (+2.06pp strict, 0.8971) on the 243-doc stratified surface — development & IP clusters eliminated; memo addendum + proposition §17 | | KANBAN-007 | v0.15.0 (2026-08-12) | tag v0.15.0 → 4b6ad5f; commits 93eb938, 0a4051e, 0afdf2e | Release finalized: changelog dedup-repaired, tag pushed, GitHub release with dedicated notes published; release.py --check green (303 tests) | | KANBAN-002 | v0.16.0 prep (2026-08-12) | commit ac156a5 | Dirty tree landed: experiment log regen (57 runs), site data regen, sorter_v7 registration + test, AGENTS.md board governance, changelog [Unreleased] entry for sorter_v7 | | KANBAN-001 | v0.16.0 prep (2026-08-12) | commit ac156a5 | SORTER_PROMPT_V7 constant + PROMPT_VERSIONS["sorter_v7"] + test_sorter_v7_data_backed_rules landed (18 prompt tests green). Evaluation tracked in KANBAN-003 | | (v0.15.0 content) | v0.15.0 (2026-08-12) | tag v0.15.0 | All v0.15.0 changelog entries (v18 sweep → v19 → v20 → v21 → v22 → v23 → v23×max; scorer fixes; annotation queues ×2; memos tab + 6 memos; wiki Langfuse-Traces + Annotation-Queues; rescore pipeline; two-project Langfuse strategy; prompt-store cleanup) — cataloged in CHANGELOG.md v0.15.0, each with its own commit in git log v0.14.0..v0.15.0 |

When moving a card here: fill this table AND leave the lane row visible in the Key Kanban table only if the card is still open; closed cards live ONLY here (the table above holds open work, the archive holds the record).