Research artifact

Session Benchmark v0

A benchmark shell built from aggregate patterns across 1,299 agent sessions. It tests whether an agent can keep evidence, constraints, privacy boundaries, and useful next actions intact when the source context is messy.

Public boundary: this project publishes scrubbed aggregates, benchmark definitions, mocked candidate fixtures, subjective pilot scores, and synthetic registry data. The private source archive remains local. It does not publish raw source excerpts, source identifiers, client details, credentials, screenshots, or production routing.

From deployment artifact to session benchmark

The project began as a synthetic registry-first deployment example. The more useful anchor was the archive itself: a long trail of sessions where recurring work becomes visible in aggregate.

The benchmark asks a narrower question than whether an agent can write a polished answer. It asks whether the agent can preserve constraints while crossing between rough notes, tools, private context, and public artifacts. The current release contains 26 cases, three mocked public-safe candidate profiles, and an eight-case observed-archive pilot used as calibration.

This is not a live leaderboard. The fixture proves the scoring harness and public artifact shape before any provider-backed candidate run is claimed.

Corpus shape

Scrubbed task-family counts derived from the local archive
Task familySessionsPublic note
Core operating system674Identity, operating principles, reusable delivery context, and routing shells.
Tool and skill infrastructure298Repeatable workflows, execution surfaces, research helpers, and packaging.
Knowledge systems279Project memory, long-term knowledge bases, file organization, and context layers.
Client delivery258Scrubbed legal-operations delivery work and implementation planning.
Product bets179Product strategy, UX, public positioning, and experimental applications.
Agent infrastructure149Agent harnesses, runtime patterns, orchestration, and comparison work.

What the qualitative review added

The 18-session proportional review added cases where generic benchmark prompts usually miss operational detail.

  1. Artifact choice before agent choiceRepeated work should become the smallest useful artifact: prompt, context pack, skill, workflow, hook, specialist agent, or full operating kit.
  2. Autonomy needs a ledgerLong-running agents need explicit state: discovered, classified, moved, skipped, ambiguous, verified, remaining, and blocked.
  3. Public-safe by defaultIf a repo or page may become public, the data model must be synthetic from the start. Git history is part of the privacy surface.
  4. Evidence before synthesisComplex answers improve when source inventory, filesystem review, outside research, and final synthesis are separate lanes.
  5. Settled decisions must bind future workThe agent should not reintroduce deprecated architecture after the user has corrected it.
  6. Runtime selection follows riskDifferent jobs need different harnesses: read-only monitoring, scoped batch writing, broad-write automation, drafting, and parallel research are not the same agent problem.
  7. Design is a governance problemAvoiding generic AI design requires a pre-build ritual: diagnose convergence, choose a direction, bind the design spec, and audit the result.
  8. Scope discipline is a quality signalA strong agent preserves ambition while shipping the smallest useful version first.
  9. Communication work is operational workUseful agents turn rough notes into concise client-safe recaps, decision requests, and next actions without inflating the tone.
  10. Billing requires reconciliationChat/session history alone is not enough for client-facing billing; calendars, emails, drafts, deliverables, and unrelated noise must be reconciled.

Publication boundary

Published: aggregate counts, scrubbed task-family labels, benchmark case definitions, mocked candidate fixture results, observed-archive pilot scores, failure-mode labels, synthetic registry fixture.

Withheld: client names, staff names, matters, emails, screenshots, credentials, private field names, raw source excerpts, production automations, exact routing logic, local source UUIDs.

Data and eval downloads