All docs · North OS2026-09-16

Mail pipeline off the Anthropic API — PARTIAL design spec

Status: in-progress · Updated: 2026-09-17 · Related: implementation plan. Supersedes the mail half of cheaper classification.

PARTIAL delivery — user decision, 2026-09-17

This PR delivers gate/classifier/matter changes only. Kwiss explicitly selected Separate the rollup change: restore the existing rollup contract, defer rollup replacement and exclude rollup reliability from this PR's acceptance. This is not a claim that the original rollup is safe. No product feature is withdrawn and no new fallback is introduced. Status remains in-progress, not whole-feature implemented; product-wide Anthropic removal is not delivered.

Current routing and preserved contracts

Only the three OpenRouter stages use the shared registry policy zdr: true, data_collection: "deny", allow_fallbacks: false. Kwiss explicitly authorized the gate's reuse of the existing DeepSeek/OpenInference route in the implementation conversation; no new provider approval is implied. The unused Ling production registry entry is removed, not replaced by another cheap rollup route.

Normalized available body text precedes snippet fallback at the unchanged 2000-unit gate cap. The privacy/capacity decision precedes terse-work actionability. Memo namespace v3-deepseek-v4-flash-relevance-json1 invalidates the obsolete gate contract; private exclusions are never memoized. Thin classification precedes the judge so a thin failure cannot repay a completed judge. Only the gate is memoized: fresh retrieval, authoritative drive pins, membership/confidence guards, CAS and opaque legacy sonnet identifiers remain intact. Resolver model failures abstain; later job failures can still repay completed thin/judge work.

Noise skips the judge only without a valid hard drive pin or explicit matter number. Thread continuity applies only after higher-priority unique matches fail; ambiguous earlier evidence abstains rather than falling through to thread evidence. Unterminated script/style bodies retain their text after tag removal: normalization is not sanitization.

Measured evidence and limits

The frozen current-source run reports gate 154/156, with 40/40 private exclusions, one actionable miss (ga-fa-bg-acris-info) and one false positive (gn-bare-matter-invite). Matter is 26/27, with no false links; thread-bg-tokenless abstains despite an inherited thread matter. No invalid/fallback outputs in that run. Separate targeted judge-precedence validation passed 15/15 full frozen contracts, using 45 generation attempts; this is not 15 new corpus cases or a complete corpus pass.

These measurements exercise real helpers/guards with frozen retrieval, not live ingestion, persistence, Redis or publication. Source manifests, precise costs, historical review provenance and post-restoration checks are recorded in the durable validation report. Prior root checks predate scope cleanup and must be rerun; the report distinguishes the new checks. Gate prompt executable text is unchanged by the standalone comment correction.

Rollup benchmarks are a no-go, not acceptance evidence for this partial PR. Historical Ling evaluation had one sensitive leak and 11 incomplete attempts in 34×3. Later GLM C retained actor/date/status errors; GLM D produced 4/16 invalid JSON responses, and DeepSeek D had 13 material findings on 70 plus one on eight challenges. The external operations prototype stopped at an invalid reference on call eight, with three material findings among seven valid outputs. None is shipped, and none establishes incumbent safety.

Cost accounting retains actual-cost subtotals and explicit unknown-cost attempts, never token-price estimates. The retained rollup evaluator now executes the restored Anthropic SDK path, records every HTTP attempt including SDK retries, and treats unavailable direct-API charges as unknown (therefore failing its unchanged cost-integrity gate). Its optional pairwise judge alone uses OpenRouter. This utility is not a rollout or a way to re-enable Ling.

Historical benchmark context (not current-source clearance)

Earlier alternative gate/judge/rollup proposals are historical, not current routing instructions; rollup now retains the baseline contract by explicit user decision. Historical evidence remains unchanged at /home/kwiss/north-data/investigations/2026-09-16-mail-model-benchmark.md, with raw artifacts in ~/north-data/investigations/2026-09-16-mail-model-bench/. Its successive benchmark totals were $0.66, $1.07, $2.34 and $2.78; these are historical reported amounts, not current actual-cost accounting or a savings claim.

On release 10a5bace, the historical seven-day volumes were 25,240 gate calls, 20,062 judge calls and 2,826 rollups. The retry investigation reported 6,516 retried attempts across 2,245 emails, 5,179 HTTP 429s, 1,328 timeouts and 784 exhausted into the DLQ. The old gate → judge → chip → commit ordering explained why a late chip failure re-paid upstream calls; it is not a description of the current Anthropic routing.

The initial 62-case × 3 gate benchmark reported Haiku private/actionable recall 100% at $0.8649/1k, versus Nemo private recall 90% at $0.0113/1k. The capacity-rule rerun reported Nemo private recall 100%, actionable 92% and $0.0162/1k versus Haiku $1.1246/1k. Those small-corpus results did not override the later historical 38/40 private recall; current-source measurements are separated above.

The historical judge payload was 79.2% drive evidence, with up to 421 candidates. The proof-of-concept top-10 cap reduced payload characters by 69.7%; it was not the production ranking specification. The initial replay reported Sonnet 16/27 at $21.165/1k and capped Haiku 16/27 at $4.591/1k, both with zero false links. The later 81-case-run replay reported Haiku 48/81 with zero false links, Sonnet 48/81 with three, and DeepSeek 44/81 with three ($4.591, $9.461 and $0.050/1k respectively).

The real full-guard battery then reported Haiku 36/81 with six false links at $8.753/1k and DeepSeek 35/81 with three at $0.168/1k. This withdrew the replay-based safety conclusion. DeepSeek's three links traced to the spurious noise-clio-eu drive pin. Historical status-fail counts were six for DeepSeek versus zero for Haiku; all-case correctness is not the shipped expect-matter link-recall denominator.

The historical rollup probe covered only 8 cases × 3: Sonnet, Haiku and Ling each scored 24/24 on firm_shareable, at $2.9774, $0.8769 and $0.0605/1k respectively. Ling won 7 of 8 blind pairwise comparisons against Sonnet with one cross-family judge. This small probe cannot override the later sensitive leak and incomplete attempts.

The old ~$447 → ~$50/week and ~$20,600/year projections multiplied historical sample costs by weekly volume. They remain historical estimates, not delivered savings or current cost totals. No current savings claim is established here.

Acceptance and uncompleted scope

Gate/classifier/matter acceptance retains private recall 1.0, actionable recall at least 0.9, zero false matter links, strict parser/privacy guards and complete attempt/cost denominators. The observed three current-source quality misses remain visible; no label or threshold is changed. Rollup reliability and migration are explicitly deferred, not quietly marked passed. Broader nonmail direct-Anthropic inventory/removal also remains open.

Root lint/typecheck and affected package tests must cover the restored scope. The plan's broader three-repeat exclusion/resolver evidence is not replaced by a single frozen-helper run; this cleanup performs no paid or database evaluation. Two independent adversarial reviews precede any substantial merge, including a deletion audit. Prior reviews are not automatically reviews of this restoration. Merge and deployment remain kwiss's decisions.

No schema migration, reclassification, backfill, golden relabeling, cascade, shadow-run framework, UI change or new fallback. Generic routed transport/accounting and reviewed gate/classifier/matter fixes remain. Rollback uses the previous deployment and configuration.