Field note

The Unforgetter pilot: what we measure and why

The pilot runs 25 August to 8 September 2026. This note publishes the scorecard, the baseline problem and the go/no-go criteria before any result, so the numbers cannot be chosen afterwards.

UnforgetterWritten 27 August 20265 min readMarkdown version

Unforgetter’s website will not show a metric until it has been measured in a way that can be reproduced. This note is the other half of that promise: the measurements are published before the results, so the results cannot be selected to look good.

The setup

  • One workspace, one owner.
  • Nine explicitly allowlisted internal projects. No raw dump of anything.
  • Three agent zones: Codex, Claude and a Grok Bot team treated as one shared trust zone with a shorter-lived token.
  • A gateway hosted in an EU West region; a canonical writer on a controlled machine that publishes every 15 minutes and after each consolidation.
  • One synthetic restricted fixture that exists only to be refused.
  • Window: 25 August 21:40 to 8 September 21:40 CEST, counted from the first real end-to-end smoke test, not from the plan date.

What is measured per useful run

Measure What counts
Context correctness Stale or irrelevant facts served; extra retrieval calls needed to finish
Payload size Items in the bundle; whether the agent needed to search beyond it
Completed deliveries Work verified as finished without a new prompt from the owner
Proposal quality Accepted, corrected and duplicate proposals, traceable to actor and evidence
Handoffs Tasks continued by a second agent without re-explanation; false progress caught
Scope and injection events Reads outside scope, canary appearances in proposals or tool actions. Target: zero
Supervision time Owner minutes on context setup, handoff, approval and correction
Cost Model and cloud spend per active day
Connector setup time Minutes to connect each agent platform

The baseline problem

The manual routing that preceded the pilot was never timed systematically. Rather than estimate it, the pilot records a parallel baseline on the first three real tasks: how many minutes the owner spends on context, handoff, approval and correction, alongside the same tasks’ connector setup and independent retrieval. The comparison is between measured and measured, not measured and remembered.

Two evaluation tracks

Parity. The same bounded, read-only task is given to all three zones with the same scope. Each returns its result, the context ids it used, proposals, tool actions, elapsed time, supervision requests and errors as structured data. No comparison is made until all three outputs exist. One early observation is already recorded: an A/B on an outreach draft was not fair because the shared record lacked the owner’s writing profile. That gap is a finding, not a footnote.

Proactivity. A standing, non-sensitive goal runs over several days. Useful work counts only when the delivery can be verified without relying on the agent’s own claim of completion.

Go/no-go for a real alpha

The pilot advances to a managed alpha only if all of these hold:

  • all three zones read correctly scoped context and deliver proposals
  • no agent can write canonical memory directly or read outside scope
  • retries create no duplicates
  • proposals trace back to actor and evidence
  • at least one standing agent delivers verified useful work without a new prompt
  • supervision time and cost measure better than the recorded baseline
  • the negative contract tests and the injection canary have run, and no breach succeeded

If any fail, the contract, the context composition or the write model is fixed before signup, billing, sync or multi-tenant infrastructure is built.

What will be published

When the window closes, the scorecard is summarised on this site with the method above, the measured numbers, and the limitations. Where a measurement was not taken, the summary will say so. Any result will be published as a dated follow-up in Product updates, not turned into an undated homepage claim.