🦉 llm-mailroom · graph report

Knowledge Graph Report — llm-mailroom

built 2026-08-27 commit 13346a92 code-only AST production src/ · tests/notebooks/.opencode excluded graphify 0.9.50
0
nodes
0
edges
0
communities
70 shown · 24 thin omitted
0
source files
production Python only

↗ Open the architecture map 🌳 Module tree

What changed

Rebuilt from 13346a92 (was 30ff6874, mailroom PRs #30–#36: dojo 0.9.0/0.10.0 scoring suites, intake clerk, tagged dojo pin, honesty gaps, local eval packs, and parking objective misses for review before catalog write).

This build indexes production src/ only (112 files) so the map follows the live architecture: specialist get_suite() extras, honesty-gap metadata, local eval packs, reconsideration.py parking, Hub CUAD/MAUD inventories, and the 13-node conveyor. Tests, notebooks, and .opencode/skills are excluded on purpose.

Navigate the architecture map as layers → modules → symbols. The 13-node conveyor strip at the top jumps to ingest_node, classify_node, Lane A/B, Boss, catalog, and archive. Comments (rationale nodes) are hidden until you uncheck them. Double-click a symbol to isolate its neighborhood.

Community Hubs (click to filter)

God Nodes — most connected core abstractions

load_config()
37
BaseAgent
35
build_graph()
33
BaseAgent
32
get_langfuse_client()
31
get_managed_prompt()
28
_execute_run()
25
ensure_schema()
25
DocumentState
23
main()
23

Surprising Connections (cross-community bridges)

Import Cycles

✓ None detected — the dependency surface is acyclic.

Communities (94 total · 24 thin omitted)

#000

Graph nodes & state

90 nodes
cohesion
0.06

DocumentState, DocumentManifest, PipelineStage, apply_intake(), arbiter_node(), archive_node(), boss_escalation_node(), _build_checkpointer() (+82 more)

#001

HF pilot & honesty gaps

83 nodes
cohesion
0.06

_denial_reasons(), determination_consistency_is_quality(), honesty_trace_metadata(), insurance_determination_consistent(), insurance_determination_issues(), insurance_expected_set_is_homogeneous(), insurance_gt_is_homogeneous(), _norm_determination() (+75 more)

#002

Eval mocks & validation

59 nodes
cohesion
0.05

FakeLangChainLLM, _FakeStructuredRunner, chat, _Choices, completions, _EvalLangChainLLM, _HintedEvalLangChainLLM, _MockClient (+51 more)

#003

Inbox bins

58 nodes
cohesion
0.09

get_queue(), _move_rejected_to_failed(), _ensure_dirs(), accepted_extensions(), archive_dir(), classified_dir(), clear_ingestion_paused(), failed_dir() (+50 more)

#004

Arbiter & Boss

56 nodes
cohesion
0.06

ArbiterAgent, BossAgent, ComplianceSpecialist, CorporateRecordsSpecialist, CorrespondenceSpecialist, InsuranceClaimsSpecialist, .arbitrate(), .system_prompt() (+48 more)

#005

Tracing backends (5)

55 nodes
cohesion
0.06

configure(), flush_braintrust(), instrument_openai_client(), is_configured(), _apply_taxonomy_settings(), client_kwargs(), flush_langfuse(), get_trace_id() (+47 more)

#006

Routing & reconsideration

54 nodes
cohesion
0.08

after_arbiter(), after_boss(), after_classify(), after_extraction(), after_extraction_gated(), after_human_review(), after_judge(), after_report() (+46 more)

#007

Catalog & audit trail

49 nodes
cohesion
0.10

AuditLogRecord, DocumentRecord, MatterRecord, Base, _all_chains(), main(), get_audit_chain(), get_latest_audit_hash() (+41 more)

#008

Sorter classification (8)

41 nodes
cohesion
0.07

SorterReviewerAgent, SorterAgent, .review(), .system_prompt(), build_structured_schema(), format_sorter_subclass_catalogs(), sorter_subclass_catalog(), valid_sorter_subclasses() (+33 more)

#009

LangChain specialists

35 nodes
cohesion
0.09

ComplianceFilingSpecialist, CorporateRecordsSpecialist, CorrespondenceSpecialist, _SpecialistBase, get_prompt(), .system_prompt(), .system_prompt(), .system_prompt() (+27 more)

#010

LangChain BaseAgent

33 nodes
cohesion
0.10

BaseAgent, .augmented_system_prompt(), ._build_user_content(), ._call_llm(), ._call_structured(), ._call_vision(), ._call_vision_multi(), ._check_deadline() (+25 more)

#011

FastAPI intake

33 nodes
cohesion
0.11

_check_database(), get_audit_trail(), get_document_status(), get_matter(), health(), lifespan(), ops_resume(), ops_status() (+25 more)

#012

Hub subclass inventories

33 nodes
cohesion
0.12

clause_handoff(), skip_conflict_field(), coerce_gt_value(), _compact(), enrich_extraction(), _normalize(), normalize_claim_type(), normalize_communication_type() (+25 more)

#013

Run limits & budgets

28 nodes
cohesion
0.10

RunBudgetExceeded, RunDeadlineExceeded, _bounded(), compute_run_metrics(), check_token_budget(), estimate_cost(), get_call_timeout_seconds(), get_deadline_seconds() (+20 more)

#014

classifier

25 nodes
cohesion
0.11

classify_image(), clean_prediction(), extract_confidence(), extract_reasoning(), extract_runner_up(), _valid_classes(), build_text_messages(), build_vision_messages() (+17 more)

#015

Vision rendering

24 nodes
cohesion
0.16

_resolved_models(), agent_uses_vision(), _any_specialist_uses_vision(), is_vision_capable(), max_pages(), pipeline_uses_vision(), render_document_pages(), render_image() (+16 more)

#016

Quality scores

24 nodes
cohesion
0.12

ensure_field_score_configs(), score_and_log_extraction(), _client(), emit_pipeline_scores(), ensure_score_configs(), is_enabled(), langfuse_score_name(), _score_data_type() (+16 more)

#017

PDF & image transcription

23 nodes
cohesion
0.19

ImageExtractor, extract_text_from_image(), .extract(), ._extract_with_vision(), ._fallback_extract(), .system_prompt(), ._llm_transcribe(), compile_matter_record() (+15 more)

#018

LegalBench runner

22 nodes
cohesion
0.18

RunResult, build_parser(), main(), log_run(), _model_name(), print_summary(), run_task(), _tokens_summary() (+14 more)

#019

Langfuse evaluator sync

22 nodes
cohesion
0.16

_build_evaluator_request(), _build_output_definition(), _build_rule_request(), _client(), _current_evaluator_prompt(), _ensure_llm_connection(), _existing_rule_ids(), main() (+14 more)

#020

Experiment log

21 nodes
cohesion
0.19

append_record(), build_record(), default_log_path(), default_sibling_root(), git_snapshot(), _inside(), regenerate(), _run_python() (+13 more)

#021

Field scoring calibration

21 nodes
cohesion
0.14

get_field_types(), warm_embedding_model(), main(), _perturb_date(), _perturb_entity_list(), _perturb_free_text(), _perturb_money(), _perturb_name() (+13 more)

#022

CUAD corpus loaders

21 nodes
cohesion
0.19

_contracts_from_annotations(), _contracts_from_txt(), _download(), download_all(), _list_hf_files(), _load_subtype_taxonomy(), main(), _normalize_category() (+13 more)

#023

Grounded pilot

20 nodes
cohesion
0.18

_attach_field_scoring(), diff_report(), filter_real_samples(), _ground_truth_scores(), _ingest_scores(), main(), misfile_candidates(), _parse_expected_fields() (+12 more)

#024

Watcher ingest

19 nodes
cohesion
0.20

InboxHandler, Watcher, claim_file(), ._infer_matter_id(), .__init__(), ._is_processable(), .on_created(), ._process() (+11 more)

#025

LLM providers

18 nodes
cohesion
0.22

ProviderConfig, .__init__(), _check_llm_provider(), get_llm(), get_llm_client(), get_llm_model(), instrument_client(), _build_providers() (+10 more)

#026

CUAD/MAUD inventories

18 nodes
cohesion
0.20

as_clause_lines(), enrich_contract_extraction(), flatten_cuad_clause_labels(), flatten_maud_clause_labels(), infer_merger_consideration(), normalize_consideration(), parse_json_obj(), _as_meta() (+10 more)

#027

config

17 nodes
cohesion
0.15

ContractsSpecialist, ContractsSpecialist, .__init__(), .__init__(), .system_prompt(), get_extraction_schema(), get_all_doc_types(), get_doc_class() (+9 more)

#029

LegalBench data

17 nodes
cohesion
0.16

CorpusUnavailable, Sample, _fingerprint(), load_cuad_qa(), load_family_rows(), _normalize_prediction(), _extract_binary(), data.py (+9 more)

#028

LLM retry

17 nodes
cohesion
0.20

_is_retryable_error(), _is_json_mode_400(), _is_retryable(), _retry_after_seconds(), _retry_config(), retry_sleep_seconds(), _status_code(), check_run_deadline() (+9 more)

#030

ops monitor

16 nodes
cohesion
0.17

OpsMonitor, _main(), ._analyze_metrics(), ._gather_metrics(), .__init__(), .is_paused(), .pause_info(), ._query_catalog() (+8 more)

#031

audit

16 nodes
cohesion
0.23

AuditLogEntry, archive_document(), _file_sha256(), build_audit_entry(), compute_audit_hash(), compute_audit_hash_v1(), verify_chain(), _verify() (+8 more)

#032

LegalBench agent

16 nodes
cohesion
0.14

LegalBenchAgent, .answer_binary(), .augmented_system_prompt(), .classify_family(), .__init__(), .system_prompt(), .usage(), agent.py (+8 more)

#033

Langfuse tracing

16 nodes
cohesion
0.16

attach_run_scores(), ensure_score_configs_if_enabled(), _environment(), legalbench_trace(), question_observation(), is_enabled(), pipeline_trace(), langfuse_tracing.py (+8 more)

#034

Agent toolkit & memory

15 nodes
cohesion
0.21

AgentTool, ._tool_context(), .__init__(), .run(), _build_toolkit(), get_tools(), render_tools(), _tool_field_types() (+7 more)

#035

Tracing backends

15 nodes
cohesion
0.15

_NoopLangfuse, _NoopSpan, .create_trace_id(), .flush(), .get_current_trace_id(), .set_current_trace_io(), .shutdown(), .start_as_current_observation() (+7 more)

#036

Specialist scoring suites

15 nodes
cohesion
0.22

attach_single_doc_extras(), _numeric_extra(), score_and_log_intake(), score_intake_suite(), score_with_suite(), unwrap_suite_result(), suite_scoring.py, Any (+7 more)

#037

Langfuse model sync

15 nodes
cohesion
0.21

default_environment(), load_env(), _client(), _cost_models(), _existing_by_name(), main(), _match_pattern(), _prices_match() (+7 more)

#038

Pipeline guards

15 nodes
cohesion
0.20

apply_classification_guard(), apply_extraction_guard(), guard_classification(), guard_extraction(), _has_substantive_content(), _is_valid_confidence(), _valid_subtypes(), guards.py (+7 more)

#039

Dataset sync

15 nodes
cohesion
0.26

_escape(), generate_pdf_from_text(), _load_manifest(), prepare_samples(), _client(), _doc_text(), _ensure_dataset(), main() (+7 more)

#041

Mailroom BaseAgent

14 nodes
cohesion
0.23

BaseAgent, ._build_multimodal(), ._call_llm(), ._call_structured(), ._configured_max_input_chars(), ._configured_max_tokens(), ._configured_reasoning_effort(), .system_prompt() (+6 more)

#042

Extraction schemas

14 nodes
cohesion
0.29

ComplianceFilingExtraction, ContractExtraction, CorporateRecordExtraction, CorrespondenceExtraction, InsuranceClaimExtraction, Matter, get_extraction_schema(), judge.py (+6 more)

#044

Dashboard sync

14 nodes
cohesion
0.29

WidgetSpec, _client(), _existing_placements(), json_dumps(), main(), _placement_kwargs(), _score_widget(), _spec_to_request() (+6 more)

#040

Langfuse log sync

14 nodes
cohesion
0.24

main(), _slug(), _client(), main(), _parse_since(), sync_logs(), _trace_basics(), _trace_stage() (+6 more)

#043

LegalBench scoring

14 nodes
cohesion
0.25

equivalent_subtypes(), _binary_f1(), _ece(), _mean(), _safe_div(), score_binary(), score_multiclass(), scoring.py (+6 more)

#045

logging

12 nodes
cohesion
0.24

_RotatingFileSink, .__call__(), .__init__(), setup_logging(), _base_env(), main(), run_config(), RotatingFileHandler (+4 more)

#046

Quality judges

12 nodes
cohesion
0.23

CompletenessJudge, ._field_list(), .judge_classification(), .judge_completeness(), .judge_extraction_correctness(), .system_prompt(), ._taxonomy_spec(), ._truncate() (+4 more)

#047

Agent toolkit & memory (47)

12 nodes
cohesion
0.26

_memory_dir(), _memory_path(), recent_context(), record_outcome(), stats(), _tool_memory(), memory.py, Path (+4 more)

#048

Docclass prompt arm

12 nodes
cohesion
0.20

list_prompts(), PROMPT_TEMPLATES(), docclass_prompts_enabled(), langchain_prompt_version(), managed_prompt_lookup(), langchain_agents/prompts.py, docclass_mode.py, List all available prompt versions. (+4 more)

#049

Field scoring & metrics

12 nodes
cohesion
0.24

_date_pair_days(), extraction_diagnostics(), _mean(), _median(), parse_duration_days(), _r2(), metrics.py, Run-level diagnostic metrics for extraction scoring. Ported from ``llm-entity-… (+4 more)

#051

mock

11 nodes
cohesion
0.27

MockLegalBenchModel, _hash(), .answer_binary(), .classify_family(), .__init__(), .last_usage(), .usage(), legalbench/mock.py (+3 more)

#050

Graph nodes & state (50)

11 nodes
cohesion
0.18

_latest_audit_hash(), _persist_provenance(), _persist_scores(), _run_coro(), _touch_heartbeat(), _write_catalog_record(), Run a coroutine from a sync context: schedule it on the running loop when one…, Best-effort fetch of the last entry_hash for this doc_id (the previous link of… (+3 more)

#052

bootstrap

11 nodes
cohesion
0.29

bootstrap_ci(), _clean(), delta_significance(), _resample_means(), bootstrap.py, Any, Random, Bootstrap confidence intervals and small-sample delta testing. Ported verbatim… (+3 more)

#053

Quality judges (53)

11 nodes
cohesion
0.31

create_trace_score(), is_real_sample(), _dim_summary(), _ingest(), judge_one(), main(), print_summary(), _raw_text_for() (+3 more)

#054

External samples

11 nodes
cohesion
0.33

_caption_from_text(), _download(), fetch_atticus(), fetch_legalbench(), fetch_pileoflaw(), main(), _stream_pol_records(), fetch_external_samples.py (+3 more)

#055

Managed prompts

10 nodes
cohesion
0.31

_langchain_prompt(), prompt_templates(), get_langfuse_client(), _client(), _current_production(), main(), sync_one(), sync_prompts.py (+2 more)

#056

Prompt cutover

10 nodes
cohesion
0.40

cutover_agent(), cutover_all(), list_agents(), list_local_models(), load_config(), main(), recommend_cutover_order(), save_config() (+2 more)

#057

PDF & image transcription (57)

9 nodes
cohesion
0.33

PDFTranscriber, ._extract_raw_text(), ._looks_clean_text_pdf(), .system_prompt(), .transcribe(), transcribe_pdf(), BaseAgent, Path (+1 more)

#060

tasks

9 nodes
cohesion
0.25

LegalBenchTask, _binary_classes(), _call_binary(), _call_family(), _extract_family(), _family_classes(), _family_labels(), tasks.py (+1 more)

#058

env utils

9 nodes
cohesion
0.28

bool_env(), get_env(), load_env(), require_env(), env_utils.py, Load ``braintrust.env`` then ``.env`` into the environment (idempotent).…, Validate that all given environment variables are set and non-empty. Returns…, Get an environment variable with a default fallback. (+1 more)

#059

prompts docclass

9 nodes
cohesion
0.28

_append(), _build_versions(), _rules(), _specialist_rules(), prompts_docclass.py, Docclass prompt variants for every mailroom classification-chain role.…, Pure-appended docclass variant: base is a STRICT PREFIX of the result., Derive every variant from the live production template of that role. (+1 more)

#061

Pilot reports

9 nodes
cohesion
0.42

build_report(), _clean_extracted(), _field_score_for(), _fmt_usd(), _json_block(), _load_config(), main(), _manifest_rows() (+1 more)

#062

Sorter classification

8 nodes
cohesion
0.29

SorterAgent, .classify(), .classify_json(), .__init__(), _LangChainSorterAgent, Mailroom-configured sorter. - Model/budget defaults come from ``taxonomy.yaml``…, Classify a document, optionally with page images attached. Returns ``(doc_type,…, Structured classify used by the graph (includes ``doc_subclass``).

#063

Grounded pilot (63)

7 nodes
cohesion
0.29

_check_cost_watchdog(), _fetch_openrouter_prices(), _price_for(), _record_langchain_response(), Warn at $0.15, abort the run at $0.20 (cumulative across all samples)., Mirror _wrap_client's usage/cost accounting for a LangChain response., Fetch live OpenRouter pricing (per-token), normalized to $/M tokens. The…

#064

prompts

6 nodes
cohesion
0.40

family_classification_prompt_v1(), get_prompt(), legalbench/prompts.py, Versioned LegalBench task prompts. Prompt version = experiment identity in the…, Fill the 25-family list into the multiclass prompt (called per run so the…, Resolve a prompt version to its system-prompt text.

#065

compare runs

6 nodes
cohesion
0.60

_aggregate(), _cell(), main(), _print_table(), _scores_of(), compare_runs.py

#066

Graph nodes & state (66)

5 nodes
cohesion
0.40

_chunk_config(), _extract_contracts(), _run_chunked_extraction(), Chunked-extraction config from taxonomy.yaml (`chunking:` block). Chunking…, Run a specialist extraction, chunking long documents (v15+ pass).…

#067

Graph nodes & state (67)

4 nodes
cohesion
0.50

_prompt_versions(), _bound_prompt_versions(), Prompt versions bound during the run (best-effort; Langfuse-managed prompts…, Version keys currently wired into production / agent defaults. Used for catalog…

#068

Field scoring & metrics (68)

4 nodes
cohesion
0.50

field_is_ambiguous(), get_type_bands(), Per-field-type ambiguous-band overrides from ``field_scoring.type_bands``.…, Is this field score in the (possibly type-specific) ambiguous band? Band check…

#069

FastAPI intake (69)

3 nodes
cohesion
0.67

_require_token(), Request, Dependency: reject requests without the bearer token (audit L-2).

No communities match.

Knowledge Gaps

Rationale comments and a handful of stdlib/unattributed symbols sit in thin communities. They are hidden on the map by default. Isolated __init__.py package markers are expected.

Suggested Questions

Why does `BaseAgent` connect `LangChain BaseAgent` to `LegalBench agent`, `HF pilot & honesty gaps`, `Eval mocks & validation`, `Agent toolkit & memory`, `Sorter classification (8)`, `LangChain specialists`, `LangChain BaseAgent (72)`, `LangChain BaseAgent (78)`, `LangChain BaseAgent (79)`, `Grounded pilot`, `LLM retry`?

High betweenness centrality (0.090) - this node is a cross-community bridge.

Why does `load_config()` connect `Vision rendering` to `Graph nodes & state`, `Inbox bins`, `Tracing backends (5)`, `Routing & reconsideration`, `Sorter classification (8)`, `FastAPI intake`, `Run limits & budgets`, `PDF & image transcription`, `Langfuse evaluator sync`, `Field scoring calibration`, `config`, `LLM retry`, `Agent toolkit & memory`, `Langfuse model sync`, `Extraction schemas`, `Quality judges`, `PDF & image transcription (57)`, `Graph nodes & state (66)`, `Field scoring & metrics (68)`?

High betweenness centrality (0.067) - this node is a cross-community bridge.

Why does `load_env()` connect `Langfuse model sync` to `HF pilot & honesty gaps`, `compare runs`, `Inbox bins`, `Eval mocks & validation`, `Dataset sync`, `Langfuse log sync`, `Catalog & audit trail`, `FastAPI intake`, `Dashboard sync`, `logging`, `Langfuse evaluator sync`, `Field scoring calibration`, `CUAD corpus loaders`, `Grounded pilot`, `Watcher ingest`, `LLM providers`, `Quality judges (53)`, `Managed prompts`?

High betweenness centrality (0.040) - this node is a cross-community bridge.

Are the 8 inferred relationships involving `BaseAgent` (e.g. with `SorterAgent` and `_SpecialistBase`) actually correct?

`BaseAgent` has 8 INFERRED edges - model-reasoned connections that need verification.

Are the 25 inferred relationships involving `build_graph()` (e.g. with `arbiter_node()` and `archive_node()`) actually correct?

`build_graph()` has 25 INFERRED edges - model-reasoned connections that need verification.

Are the 10 inferred relationships involving `BaseAgent` (e.g. with `ArbiterAgent` and `CompletenessJudge`) actually correct?

`BaseAgent` has 10 INFERRED edges - model-reasoned connections that need verification.

What connects `mailroom` to the rest of the system?

1 weakly-connected nodes found - possible documentation gaps or missing edges.