🦉 llm-mailroom · graph report

Knowledge Graph Report — llm-mailroom

built 2026-09-14 commit 2a212e76a62b code-only AST production src/ · tests/notebooks/.opencode excluded graphify 0.9.53
0
nodes
0
edges
0
communities
74 shown · 82 thin omitted
0
source files
production Python only

↗ Open the architecture map 🌳 Module tree

What changed

Rebuilt from 2a212e76 (mailroom v0.7.1 + mailroom-dataset v9 corpus, GT-closure revision 46a4d3c2 + ground truth: pared LLM load, 13-node layered-state pipeline, Gmail triage + relations clerk auxiliary flows, review-resolve tray, judge/arbiter lanes, deterministic field scoring, and the Digital-Mailroom monorepo docs alignment).

This build indexes production src/ only (142 files) so the map follows the live architecture: review_resolve.py, posthoc_gt.py, agent_eval.py, specialist suites, honesty-gap metadata, and the 13-node conveyor (human_review pauses with interrupt()). Tests, notebooks, and .opencode/skills are excluded on purpose.

Navigate the architecture map as layers → modules → symbols. The 13-node conveyor strip at the top jumps to intake_node, classify_node, Lane A/B, Boss, catalog, and archive. Comments (rationale nodes) are hidden until you uncheck them. Double-click a symbol to isolate its neighborhood.

Community Hubs (click to filter)

God Nodes — most connected core abstractions

ensure_schema()
55
load_config()
50
async_session()
50
load_env()
46
BaseAgent
41
BaseAgent
39
get_managed_prompt()
34
build_graph()
33
setup_logging()
32
inbox_dir()
31

Surprising Connections (cross-community bridges)

Import Cycles

✓ None detected — the dependency surface is acyclic.

Communities (102 total · 28 thin omitted)

#000

Hub subclass inventories

83 nodes
cohesion
0.05

_resolved_models(), as_clause_lines(), clause_handoff(), enrich_contract_extraction(), flatten_cuad_clause_labels(), flatten_maud_clause_labels(), infer_merger_consideration(), normalize_consideration() (+75 more)

#001

Graph nodes & state

55 nodes
cohesion
0.09

DocumentState, arbiter_node(), boss_escalation_node(), _build_checkpointer(), build_graph(), _build_handoff_context(), _build_specialist_dispatch(), catalog_write_node() (+47 more)

#002

Managed prompts

53 nodes
cohesion
0.08

BaseAgent, BossAgent, ComplianceSpecialist, CorporateRecordsSpecialist, CorrespondenceSpecialist, InsuranceClaimsSpecialist, build_structured_schema(), .adjudicate() (+45 more)

#003

Post-hoc schema GT

53 nodes
cohesion
0.07

configure(), instrument_openai_client(), is_configured(), _apply_taxonomy_settings(), field_is_ambiguous(), get_type_bands(), warm_embedding_model(), _date_pair_days() (+45 more)

#004

Watcher ingest

51 nodes
cohesion
0.07

InboxHandler, Watcher, _WatcherLock, WatcherLockHeld, lifespan(), run_pipeline(), claim_file(), is_ingestion_paused() (+43 more)

#005

FastAPI intake

50 nodes
cohesion
0.07

OpsMonitor, analyze_audit_database(), _check_database(), _embed_watcher_running(), get_audit_trail(), get_document_status(), get_matter(), health() (+42 more)

#006

FastAPI intake (6)

48 nodes
cohesion
0.06

DocumentManifest, PipelineStage, _document_payload_from_manifest(), lookup_document_endpoint(), _move_rejected_to_failed(), _parse_resolve_payload(), _rate_limit_upload(), _require_token() (+40 more)

#007

Inbox bins

48 nodes
cohesion
0.11

get_queue(), _ensure_dirs(), accepted_extensions(), archive_dir(), classified_dir(), failed_dir(), get_base_dir(), _get_config() (+40 more)

#008

Catalog & audit trail

44 nodes
cohesion
0.11

AuditLogRecord, DocumentRecord, MatterRecord, Base, _write_review_audit_entry(), _touch_heartbeat(), ._query_catalog(), main() (+36 more)

#009

HF pilot & honesty gaps (9)

43 nodes
cohesion
0.10

_catalog_by_trace(), completed_filenames(), enrich_sample_row(), expected_fields_for_sample(), expected_fields_meta(), finalize_report(), find_sample_text(), hf_corpus_honesty() (+35 more)

#010

logging

41 nodes
cohesion
0.09

_RotatingFileSink, _langchain_prompt(), prompt_templates(), default_environment(), load_env(), .__call__(), .__init__(), setup_logging() (+33 more)

#011

LangChain specialists

38 nodes
cohesion
0.08

ComplianceFilingSpecialist, ContractsSpecialist, CorporateRecordsSpecialist, CorrespondenceSpecialist, _SpecialistBase, get_prompt(), .system_prompt(), .__init__() (+30 more)

#012

classifier

34 nodes
cohesion
0.08

classify_image(), clean_prediction(), extract_confidence(), extract_reasoning(), extract_runner_up(), _valid_classes(), build_text_messages(), build_vision_messages() (+26 more)

#013

Sorter classification

33 nodes
cohesion
0.09

SorterAgent, .review(), build_structured_schema(), format_sorter_subclass_catalogs(), sorter_subclass_catalog(), valid_sorter_subclasses(), _classification_user_message(), _doc_classes_for_prompt() (+25 more)

#014

CUAD corpus loaders

33 nodes
cohesion
0.11

CorpusUnavailable, Sample, load_cuad_qa(), load_family_rows(), _contracts_from_annotations(), _contracts_from_txt(), _download(), download_all() (+25 more)

#015

Vision rendering

32 nodes
cohesion
0.10

ContractsSpecialist, ._configured_max_input_chars(), ._configured_max_tokens(), ._truncate_input(), .__init__(), agent_uses_vision(), _any_specialist_uses_vision(), is_vision_capable() (+24 more)

#016

Routing & reconsideration

32 nodes
cohesion
0.11

after_arbiter(), after_boss(), after_classify(), after_extraction(), after_extraction_gated(), after_human_review(), after_judge(), after_retry_classify() (+24 more)

#017

PDF & image transcription

31 nodes
cohesion
0.11

PDFTranscriber, ._call_llm(), ._call_structured(), ._configured_reasoning_effort(), ._skill_appendix(), .system_prompt(), .system_prompt_with_skills(), ._extract_with_vision() (+23 more)

#018

Quality scores

30 nodes
cohesion
0.12

ensure_field_score_configs(), score_and_log_extraction(), _client(), create_trace_score(), deterministic_verdict_label(), emit_in_pipeline_judge_scores(), emit_pipeline_scores(), ensure_score_configs() (+22 more)

#019

Tracing backends (19)

30 nodes
cohesion
0.12

client_kwargs(), flush_langfuse(), get_langfuse_client(), get_trace_id(), install_on_dropped(), instrument_openai_client(), observation(), _optional_float() (+22 more)

#020

tasks

29 nodes
cohesion
0.13

RunResult, LegalBenchTask, build_parser(), main(), log_run(), _model_name(), print_summary(), run_task() (+21 more)

#021

LangChain BaseAgent

29 nodes
cohesion
0.13

BaseAgent, .augmented_system_prompt(), ._build_user_content(), ._call_llm(), ._call_structured(), ._call_vision(), ._call_vision_multi(), ._check_deadline() (+21 more)

#022

Run limits & budgets

27 nodes
cohesion
0.10

RunBudgetExceeded, RunDeadlineExceeded, _bounded(), compute_run_metrics(), check_token_budget(), estimate_cost(), get_deadline_seconds(), get_max_total_output_tokens() (+19 more)

#023

Agent toolkit & memory

27 nodes
cohesion
0.13

AgentTool, ._tool_context(), _memory_dir(), _memory_path(), recent_context(), record_outcome(), stats(), .__init__() (+19 more)

#024

Langfuse log sync

23 nodes
cohesion
0.15

main(), _slug(), _client(), main(), _parse_since(), sync_logs(), _trace_basics(), _trace_stage() (+15 more)

#025

Graph nodes & state (25)

23 nodes
cohesion
0.13

apply_intake(), _extract_text_from_docx(), _extract_text_from_image(), _extract_text_from_pdf(), _file_sha256(), _file_size(), intake_node(), _read_file_text() (+15 more)

#026

Graph nodes & state (26)

23 nodes
cohesion
0.11

archive_node(), _catalog_upsert(), _emit_stage_audit(), _existing_processing_doc_id(), _finalize_aborted(), human_review_node(), _latest_audit_hash(), _normalize_review_decision() (+15 more)

#027

Routing & reconsideration (27)

23 nodes
cohesion
0.16

after_report(), align_class(), _as_float(), class_misses_ground_truth(), collect_review_causes(), expected_class(), expected_field_coverage(), format_causes() (+15 more)

#028

Grounded pilot

23 nodes
cohesion
0.15

flush_braintrust(), flush(), _attach_field_scoring(), diff_report(), filter_real_samples(), _ground_truth_scores(), _ingest_scores(), main() (+15 more)

#029

LLM providers

22 nodes
cohesion
0.16

ProviderConfig, .__init__(), _check_llm_provider(), get_llm(), get_llm_client(), get_llm_model(), instrument_client(), _build_providers() (+14 more)

#030

HF corpora

22 nodes
cohesion
0.16

active_corpus(), adapt_hub_row(), example_for_class(), example_rows(), examples_by_class(), hub_sample(), load_example_pack(), pipeline_corpora() (+14 more)

#031

Langfuse evaluator sync

22 nodes
cohesion
0.16

_build_evaluator_request(), _build_output_definition(), _build_rule_request(), _client(), _current_evaluator_prompt(), _ensure_llm_connection(), _existing_rule_ids(), main() (+14 more)

#032

Experiment log

21 nodes
cohesion
0.19

append_record(), build_record(), default_log_path(), default_sibling_root(), git_snapshot(), _inside(), regenerate(), _run_python() (+13 more)

#033

HF pilot & honesty gaps

21 nodes
cohesion
0.16

_denial_reasons(), determination_consistency_is_quality(), honesty_trace_metadata(), insurance_determination_consistent(), insurance_determination_issues(), insurance_expected_set_is_homogeneous(), insurance_gt_is_homogeneous(), _norm_determination() (+13 more)

#034

Arbiter & Boss

20 nodes
cohesion
0.10

ArbiterAgent, ImageExtractor, SorterReviewerAgent, .arbitrate(), .system_prompt(), .extract(), ._fallback_extract(), .system_prompt() (+12 more)

#035

audit

20 nodes
cohesion
0.19

AuditLogEntry, archive_document(), _file_sha256(), build_audit_entry(), compute_audit_hash(), compute_audit_hash_v1(), verify_chain(), _all_chains() (+12 more)

#036

LLM retry

20 nodes
cohesion
0.19

_is_retryable_error(), _is_json_mode_400(), _is_retryable(), _retry_after_seconds(), retry_chat_completion(), _retry_config(), retry_sleep_seconds(), _status_code() (+12 more)

#037

Agent eval

19 nodes
cohesion
0.22

cases_for_agent(), evaluate_agent(), load_fixture_cases(), load_local_pack_cases(), load_manifest_cases(), _mean(), _read_text(), score_case() (+11 more)

#038

Local eval packs

19 nodes
cohesion
0.24

get_field_types(), all_local_pack_samples(), compliance_local_samples(), corporate_extraction_samples(), _hydrate(), insurance_contrast_samples(), local_pack_status(), _mean() (+11 more)

#039

Eval mocks & validation

18 nodes
cohesion
0.14

FakeLangChainLLM, _FakeStructuredRunner, .bind(), .__init__(), .invoke(), ._make_message(), ._run(), .with_structured_output() (+10 more)

#040

prepare samples

18 nodes
cohesion
0.20

ensure_process_tracing(), _escape(), generate_pdf_from_text(), is_real_sample(), _load_manifest(), prepare_samples(), _dim_summary(), judge_one() (+10 more)

#041

Pipeline guards

17 nodes
cohesion
0.18

validate_extraction(), apply_classification_guard(), apply_extraction_guard(), guard_classification(), guard_extraction(), _has_substantive_content(), _is_valid_confidence(), _valid_subtypes() (+9 more)

#042

LegalBench agent

16 nodes
cohesion
0.14

LegalBenchAgent, .answer_binary(), .augmented_system_prompt(), .classify_family(), .__init__(), .system_prompt(), .usage(), agent.py (+8 more)

#043

Langfuse tracing

16 nodes
cohesion
0.16

attach_run_scores(), ensure_score_configs_if_enabled(), _environment(), legalbench_trace(), question_observation(), is_enabled(), pipeline_trace(), langfuse_tracing.py (+8 more)

#046

Tracing backends

15 nodes
cohesion
0.15

_NoopLangfuse, _NoopSpan, .create_trace_id(), .flush(), .get_current_trace_id(), .set_current_trace_io(), .shutdown(), .start_as_current_observation() (+7 more)

#044

Catalog & audit trail (44)

15 nodes
cohesion
0.21

_apply_sqlite_pragmas(), close_db(), _engine_kwargs(), _ensure_models_imported(), get_engine(), get_session(), _get_sessionmaker(), init_db() (+7 more)

#045

Graph nodes & state (45)

15 nodes
cohesion
0.17

_chunk_config(), _extract_compliance(), _extract_contracts(), _extract_corporate_records(), _extract_correspondence(), _extract_insurance_claims(), _instantiate_specialist(), _run_chunked_extraction() (+7 more)

#047

Specialist scoring suites

15 nodes
cohesion
0.22

attach_single_doc_extras(), _numeric_extra(), score_and_log_intake(), score_intake_suite(), score_with_suite(), unwrap_suite_result(), suite_scoring.py, Any (+7 more)

#048

Field scoring calibration

15 nodes
cohesion
0.20

main(), _perturb_date(), _perturb_entity_list(), _perturb_free_text(), _perturb_money(), _perturb_name(), _predictions_for(), calibrate_field_scoring.py (+7 more)

#049

Eval mocks & validation (49)

14 nodes
cohesion
0.22

_EvalLangChainLLM, is_classify_call(), user_text_from_messages(), ._classify(), ._evidence_classify(), ._extract(), .__init__(), ._run() (+6 more)

#051

Dashboard sync

14 nodes
cohesion
0.29

WidgetSpec, _client(), _existing_placements(), json_dumps(), main(), _placement_kwargs(), _score_widget(), _spec_to_request() (+6 more)

#050

LegalBench scoring

14 nodes
cohesion
0.25

equivalent_subtypes(), _binary_f1(), _ece(), _mean(), _safe_div(), score_binary(), score_multiclass(), scoring.py (+6 more)

#053

Eval mocks & validation (53)

13 nodes
cohesion
0.24

_MockClient, ensure_dirs(), _collect_documents(), _expectation_for(), _load_manifest_expectations(), main(), _mock_get_llm(), validate_pipeline.py (+5 more)

#052

Tracing backends (52)

13 nodes
cohesion
0.21

flush_phoenix(), _init_opentelemetry(), _instrument_openai(), instrument_openai_client(), is_configured(), phoenix_enabled(), phoenix_setup.py, Arize Phoenix tracing backend — local, cost-free default for llm-mailroom.… (+5 more)

#054

Quality judges

12 nodes
cohesion
0.23

CompletenessJudge, ._field_list(), .judge_classification(), .judge_completeness(), .judge_extraction_correctness(), .system_prompt(), ._taxonomy_spec(), ._truncate() (+4 more)

#056

Extraction schemas

12 nodes
cohesion
0.35

ComplianceFilingExtraction, ContractExtraction, CorporateRecordExtraction, CorrespondenceExtraction, InsuranceClaimExtraction, Matter, get_extraction_schema(), documents.py (+4 more)

#055

Docclass prompt arm

12 nodes
cohesion
0.20

list_prompts(), PROMPT_TEMPLATES(), docclass_prompts_enabled(), langchain_prompt_version(), managed_prompt_lookup(), langchain_agents/prompts.py, docclass_mode.py, List all available prompt versions. (+4 more)

#057

mock

11 nodes
cohesion
0.27

MockLegalBenchModel, _hash(), .answer_binary(), .classify_family(), .__init__(), .last_usage(), .usage(), legalbench/mock.py (+3 more)

#058

bootstrap

11 nodes
cohesion
0.29

bootstrap_ci(), _clean(), delta_significance(), _resample_means(), bootstrap.py, Any, Random, Bootstrap confidence intervals and small-sample delta testing. Ported verbatim… (+3 more)

#059

External samples

11 nodes
cohesion
0.33

_caption_from_text(), _download(), fetch_atticus(), fetch_legalbench(), fetch_pileoflaw(), main(), _stream_pol_records(), fetch_external_samples.py (+3 more)

#060

Classification scoring

10 nodes
cohesion
0.31

classes_match(), normalize_class(), score_exact_classification(), check_contract(), pipeline_class(), classification_scoring.py, Any, Classification KPIs after ``merger_agreement`` became a live MAUD class. Dojo… (+2 more)

#061

Prompt cutover

10 nodes
cohesion
0.40

cutover_agent(), cutover_all(), list_agents(), list_local_models(), load_config(), main(), recommend_cutover_order(), save_config() (+2 more)

#062

Sorter classification (62)

9 nodes
cohesion
0.25

SorterAgent, .classify(), .classify_json(), .__init__(), _invoke_sorter(), _LangChainSorterAgent, Mailroom-configured sorter. - Model/budget defaults come from ``taxonomy.yaml``…, Classify a document, optionally with page images attached. Returns ``(doc_type,… (+1 more)

#063

env utils

9 nodes
cohesion
0.28

bool_env(), get_env(), load_env(), require_env(), env_utils.py, Load ``braintrust.env`` then ``.env`` into the environment (idempotent).…, Validate that all given environment variables are set and non-empty. Returns…, Get an environment variable with a default fallback. (+1 more)

#064

prompts docclass

9 nodes
cohesion
0.28

_append(), _build_versions(), _rules(), _specialist_rules(), prompts_docclass.py, Docclass prompt variants for every mailroom classification-chain role.…, Pure-appended docclass variant: base is a STRICT PREFIX of the result., Derive every variant from the live production template of that role. (+1 more)

#065

HF pilot & honesty gaps (65)

9 nodes
cohesion
0.31

load_ground_truth_labels(), load_hf_rows(), _paginate_viewer(), _scan_cap(), _take_rows(), _viewer_rows(), ``max_scan <= 0`` means unlimited (do not use on the 247k Enron set)., Map filename → {expected, expected_subclass} from config=ground_truth. These… (+1 more)

#066

Langfuse model sync

9 nodes
cohesion
0.39

_client(), _cost_models(), _existing_by_name(), main(), _match_pattern(), _prices_match(), sync_models(), sync_models.py (+1 more)

#067

Grounded pilot (67)

7 nodes
cohesion
0.29

_check_cost_watchdog(), _fetch_openrouter_prices(), _price_for(), _record_langchain_response(), Warn at $0.15, abort the run at $0.20 (cumulative across all samples)., Mirror _wrap_client's usage/cost accounting for a LangChain response., Fetch live OpenRouter pricing (per-token), normalized to $/M tokens. The…

#070

Eval mocks & validation (70)

6 nodes
cohesion
0.40

_Choices, _HintedEvalLangChainLLM, .__init__(), ._hint_from_filename(), .__init__(), Evidence classifier with a filename-hint override. The repository's real sample…

#068

LegalBench data

6 nodes
cohesion
0.33

_fingerprint(), _normalize_prediction(), _extract_binary(), Any, yes'/'no' normalization for binary answers (lenient)., Deterministic corpus fingerprint for the sampled rows.

#069

prompts

6 nodes
cohesion
0.40

family_classification_prompt_v1(), get_prompt(), legalbench/prompts.py, Versioned LegalBench task prompts. Prompt version = experiment identity in the…, Fill the 25-family list into the multiclass prompt (called per run so the…, Resolve a prompt version to its system-prompt text.

#071

Eval mocks & validation (71)

5 nodes
cohesion
0.40

chat, completions, _fake_client(), _fake_judge_client(), .create()

#072

Mailroom BaseAgent

4 nodes
cohesion
0.50

._build_multimodal(), ._uses_vision(), True when this agent's model accepts image input and (optionally) page images…, Build the user-message content for a document input. Vision-capable models get…

#073

Graph nodes & state (73)

4 nodes
cohesion
0.50

_prompt_versions(), _bound_prompt_versions(), Prompt versions bound during the run (best-effort; Langfuse-managed prompts…, Version keys currently wired into production / agent defaults. Used for catalog…

No communities match.

Knowledge Gaps

Rationale comments and a handful of stdlib/unattributed symbols sit in thin communities. They are hidden on the map by default. Isolated __init__.py package markers are expected.

Suggested Questions

Why does `BaseAgent` connect `LangChain BaseAgent` to `LLM retry`, `Watcher ingest`, `Eval mocks & validation`, `HF pilot & honesty gaps (9)`, `LegalBench agent`, `LangChain specialists`, `Sorter classification`, `LangChain BaseAgent (78)`, `Eval mocks & validation (53)`, `LangChain BaseAgent (85)`, `Agent toolkit & memory`, `LangChain BaseAgent (86)`, `Grounded pilot`?

High betweenness centrality (0.088) - this node is a cross-community bridge.

Why does `load_env()` connect `logging` to `Langfuse model sync`, `audit`, `Watcher ingest`, `FastAPI intake`, `Inbox bins`, `Catalog & audit trail`, `HF pilot & honesty gaps (9)`, `Eval mocks & validation`, `prepare samples`, `CUAD corpus loaders`, `Field scoring calibration`, `Dashboard sync`, `Eval mocks & validation (53)`, `Langfuse log sync`, `Grounded pilot`, `LLM providers`, `Langfuse evaluator sync`?

High betweenness centrality (0.051) - this node is a cross-community bridge.

Why does `load_config()` connect `Vision rendering` to `Hub subclass inventories`, `Graph nodes & state`, `Managed prompts`, `Post-hoc schema GT`, `FastAPI intake`, `FastAPI intake (6)`, `Inbox bins`, `classifier`, `Sorter classification`, `Routing & reconsideration`, `PDF & image transcription`, `Quality scores`, `Run limits & budgets`, `Agent toolkit & memory`, `Langfuse evaluator sync`, `LLM retry`, `Local eval packs`, `Graph nodes & state (45)`, `Quality judges`, `Langfuse model sync`?

High betweenness centrality (0.046) - this node is a cross-community bridge.

Are the 9 inferred relationships involving `BaseAgent` (e.g. with `SorterAgent` and `get_specialist()`) actually correct?

`BaseAgent` has 9 INFERRED edges - model-reasoned connections that need verification.

Are the 10 inferred relationships involving `BaseAgent` (e.g. with `ArbiterAgent` and `BossAgent`) actually correct?

`BaseAgent` has 10 INFERRED edges - model-reasoned connections that need verification.

Are the 25 inferred relationships involving `build_graph()` (e.g. with `intake_node()` and `classify_node()`) actually correct?

`build_graph()` has 25 INFERRED edges - model-reasoned connections that need verification.

What connects `mailroom` to the rest of the system?

1 weakly-connected nodes found - possible documentation gaps or missing edges.