llm-entity-extraction — Posit Cloud Portal

llm-entity-extraction

The prompt-experiment loop for the llm-mailroom legal document pipeline — every evaluation run, every agent card, every decision, in one place.

This portal brings the three working surfaces of the repository to a single URL — the experiment log (every eval run, scored), the agent kanban board (open work, lanes, archive), and the discussion board (claims, decisions, results, blockers).

Open the experiment log Kanban board Discussion board Interactive explorer

In one glance

8eval runs

1models

6prompt versions

6.2Mtokens billed

12open cards

51archived cards

216discussion entries

Tasks covered: docclass_classification, docclass_specialist_extraction. Latest run: 2026-08-26.

The three working records

Experiment log — historical record of every eval run (sorter, specialists, subtype, docclass, LegalBench tasks): run index, per-run metadata, scores and breakdowns, token usage, diagnostics. Each run deep-links into the interactive explorer.

Read the experiment log →

Kanban board — the living cross-agent task board: backlog / in_progress / blocked / in_review lanes, GitHub-issue links, the completion protocol, and the full archive of shipped work.

Read the kanban board →

Discussion board — the append-only inter-agent communication log: claims, lane moves, decisions, results, blockers, handoffs, reopenings — newest at top, color-coded per agent.

Read the discussion board →

Latest runs

# Experiment Model Prompt(s) Headline score Rows Tokens
1 qwen_qwen3.7-flash_insurance_claims_specialist_docclass_v1_docclass_ab120_s42 qwen/qwen3.7-flash insurance_claims_specialist_docclass_v1 extraction 0.6904 — 48609
2 qwen_qwen3.7-flash_insurance_claims_specialist_docclass_v0_docclass_ab120_s42 qwen/qwen3.7-flash insurance_claims_specialist_docclass_v0 extraction 0.6957 — 44376
3 qwen_qwen3.7-flash_contracts_specialist_docclass_v1_docclass_ab120_s42 qwen/qwen3.7-flash contracts_specialist_docclass_v1 extraction 0.8444 — 655412
4 qwen_qwen3.7-flash_contracts_specialist_docclass_v0_docclass_ab120_s42 qwen/qwen3.7-flash contracts_specialist_docclass_v0 extraction 0.6884 — 346327
5 qwen_qwen3.7-flash_sorter_docclass_v7_docclass_ab120_s42 qwen/qwen3.7-flash sorter_docclass_v7 exact_match 0.6833 — 1999237
6 qwen_qwen3.7-flash_sorter_docclass_v6_docclass_ab120_s42 qwen/qwen3.7-flash sorter_docclass_v6 exact_match 0.5833 — 1873843
7 qwen_qwen3.7-flash_sorter_docclass_v7_docclass_merged_ab qwen/qwen3.7-flash sorter_docclass_v7 exact_match 0.6667 — 652282
8 qwen_qwen3.7-flash_sorter_docclass_v6_docclass_merged_ab qwen/qwen3.7-flash sorter_docclass_v6 exact_match 0.6333 — 621037

The full, filterable run index lives in the experiment log; interactive per-run detail (per-document scores, traces, confusion matrices) lives in the explorer.

How this site is built

  • Quarto website (site/): rendered to docs/posit/ — the same docs/ tree GitHub Pages serves, so one URL prefix hosts both this portal and the explorer (../index.html). Deployable from Posit Cloud with quarto render + publish; no GitHub Actions anywhere.
  • Derived, never hand-edited: every page body regenerates from reports/experiment_log.jsonl, MESSAGE_BOARD.md, and MESSAGE_BOARD_DISCUSSION.qmd via site/_pre-render.py on every render; the experiment-log page uses the same renderer (src/experiment_log.py::render_full_log) as reports/experiment_log.md.
  • Same-sample honesty carries over: headline scores show raw values with sample sizes; significant deltas must come from the paired A/B runs in the log, not from cross-sample eyeballing.
Tip

One URL, two sites. docs/ is served as a whole: the Posit Cloud portal at posit/ and the interactive experiment-log explorer at the root (index.html). They link to each other from the navbar.