System demonstration

Deterministic legal intake safety proof

A fictional inquiry moves through a model-driven intake flow, but the consequential rules live outside the model: validated schemas, deterministic conflict screening, a hard write gate, append-only records, and trace-aware evaluation.

Public boundary: every party, matter, and record shown here is fictional. This page demonstrates architecture and verification, not a production conflicts process or legal advice.

The boundary

The language model can interpret prose and choose a tool. It cannot redefine the write rule. The MCP server makes no LLM calls and no network calls.

  1. 1. ExtractThe model reads an inquiry and identifies the parties and known intake fields.
  2. 2. Screenintake_check_conflicts deterministically compares party names against a fictional fixture and returns provenance.
  3. 3. Normalizeintake_validate_matter validates enums and dates, derives a fixed risk label, and reports missing fields.
  4. 4. Gateintake_log_triage refuses an unaddressed conflicts state unless a named human supplies an explicit rationale.
  5. 5. EvaluateThe separate evaluation harness checks the answer and the execution path, including required and forbidden tools.

Fictional inquiry

I was rear-ended three weeks ago. Northgate Assurance Co keeps calling me. My shoulder still hurts. Can your firm help?

The model can extract Northgate Assurance Co. The server then owns the screen.

intake_check_conflicts({
  "party_names": ["Northgate Assurance Co"]
})

The fictional fixture returns a potential adverse-party match, record P-0003, so the screen status is pending. The result includes the fixture version and as-of date instead of asking the model to remember where the match came from.

The important result is a refusal

Now assume a user tells the agent to skip conflicts and write the intake anyway.

intake_log_triage({
  "matter_name": "Walk-in inquiry",
  "conflicts_status": "not-run"
})

The tool returns an error beginning Error: conflicts gate. No row is written. A prompt can suggest safe behavior; this boundary enforces it.

A genuine emergency override requires both a named human and a rationale. If used, those values are stored permanently with the row.

What is actually tested

Evidence layers for the public reference implementation
LayerFailure it addressesPublic evidence
SchemaMalformed or unexpected values enter the workflowPydantic validation, bounded strings, enum/date tests
ConflictsA model invents or forgets a screen resultDeterministic fixture matching with record and dataset provenance
Write gateAn agent writes before conflicts are addressedHard refusal plus negative regression tests
Trace evaluationThe final answer is right but the agent used the wrong pathRequired/forbidden tools, call arguments, result assertions, call and latency budgets
ReleasePublished artifact differs from the reviewed sourceNon-root container, digest, SBOM, vulnerability evidence, signature and attestation pipeline

Reproduce it

The server and harness are separate by design. The MCP repository owns deterministic behavior and the golden suite. The harness owns scoring and evidence generation.

A model-driven score is not displayed here until a reviewed run can be published with the exact server revision, harness revision, suite hash, model identifier, timestamp, and evidence files together.