Mail pipeline off the Anthropic API — PARTIAL implementation plan
Status: in-progress · Updated: 2026-09-17 · Related: design spec.
PARTIAL delivery — user decision, 2026-09-17
This PR delivers gate/classifier/matter changes only. Kwiss explicitly selected Separate the rollup change: restore the existing rollup contract, defer rollup replacement and exclude rollup reliability from this PR's acceptance. This is not a claim that the original rollup is safe. No product feature is withdrawn and no new fallback is introduced. Status remains in-progress, not whole-feature implemented; product-wide Anthropic removal is not delivered.
Current routing and preserved contracts
- Gate and matter judge:
openrouter:deepseek/deepseek-v4-flash, pinnedopen-inference/fp8. - Chip classifier:
openrouter:z-ai/glm-5.3-flash, approved hostsrelaceandfireworks(not ordered failover). - Rollup, unchanged from merge-base
7b017a97d3: direct Anthropicclaude-sonnet-4-6, the original prompt, 6000 UTF-16 units per message, 800 output tokens, ephemeral system-prompt caching and existing failure/quarantine behavior. Both inbound and outbound callers use the worker chassis client configured withANTHROPIC_API_KEY.
Only the three OpenRouter stages use the shared registry policy zdr: true, data_collection: "deny", allow_fallbacks: false. Kwiss explicitly authorized the gate's reuse of the existing DeepSeek/OpenInference route in the implementation conversation; no new provider approval is implied. The unused Ling production registry entry is removed, not replaced by another cheap rollup route.
Normalized available body text precedes snippet fallback at the unchanged 2000-unit gate cap. The privacy/capacity decision precedes terse-work actionability. Memo namespace v3-deepseek-v4-flash-relevance-json1 invalidates the obsolete gate contract; private exclusions are never memoized. Thin classification precedes the judge so a thin failure cannot repay a completed judge. Only the gate is memoized: fresh retrieval, authoritative drive pins, membership/confidence guards, CAS and opaque legacy sonnet identifiers remain intact. Resolver model failures abstain; later job failures can still repay completed thin/judge work.
Noise skips the judge only without a valid hard drive pin or explicit matter number. Thread continuity applies only after higher-priority unique matches fail; ambiguous earlier evidence abstains rather than falling through to thread evidence. Unterminated script/style bodies retain their text after tag removal: normalization is not sanitization.
Measured evidence and limits
The frozen current-source run reports gate 154/156, with 40/40 private exclusions, one actionable miss (ga-fa-bg-acris-info) and one false positive (gn-bare-matter-invite). Matter is 26/27, with no false links; thread-bg-tokenless abstains despite an inherited thread matter. No invalid/fallback outputs in that run. Separate targeted judge-precedence validation passed 15/15 full frozen contracts, using 45 generation attempts; this is not 15 new corpus cases or a complete corpus pass.
These measurements exercise real helpers/guards with frozen retrieval, not live ingestion, persistence, Redis or publication. Source manifests, precise costs, historical review provenance and post-restoration checks are recorded in the durable validation report. Prior root checks predate scope cleanup and must be rerun; the report distinguishes the new checks. Gate prompt executable text is unchanged by the standalone comment correction.
Rollup benchmarks are a no-go, not acceptance evidence for this partial PR. Historical Ling evaluation had one sensitive leak and 11 incomplete attempts in 34×3. Later GLM C retained actor/date/status errors; GLM D produced 4/16 invalid JSON responses, and DeepSeek D had 13 material findings on 70 plus one on eight challenges. The external operations prototype stopped at an invalid reference on call eight, with three material findings among seven valid outputs. None is shipped, and none establishes incumbent safety.
Cost accounting retains actual-cost subtotals and explicit unknown-cost attempts, never token-price estimates. The retained rollup evaluator now executes the restored Anthropic SDK path, records every HTTP attempt including SDK retries, and treats unavailable direct-API charges as unknown (therefore failing its unchanged cost-integrity gate). Its optional pairwise judge alone uses OpenRouter. This utility is not a rollout or a way to re-enable Ling.
Implementation disposition
Delivered: gate input normalization and memo invalidation; approved OpenRouter gate/classifier/matter routing; classifier schema reminder and judge precedence; bounded parsing and generic attempt accounting. Retained baseline: rollup helper, original prompt and both worker callers. Deferred by user: any Ling/C/D/operations rollup replacement and its reliability acceptance. The plan is not whole-feature implemented.
Historical evidence
The design spec retains the superseded benchmark sequence, including the withdrawn replay-based safety conclusion. Historical accounting and compilation repairs at 36f2539b9 and judge-verdict correction at 65375e213 are not proof for today's working tree. No measured product-wide savings or complete Anthropic removal is claimed.
Honesty rules for this lane
These are not decoration; the episode this plan corrects was caused by their absence.
- A golden case that fails stays red until it is proven a label bug. Labels are never softened to make a run pass.
- A hard gate is never relaxed to ship. If a gate blocks the change, the change is not ready.
- Report the whole surface, not the part that was changed. "Mail spend" means all four calls, not one.
- Never quote a saving against an unmeasured baseline. If the incumbent was not measured, say so.
- Every claim of a saving cites a measured before and after, with its method and its sample size.
- A fallback that silently returns a safe default is a failure in the report, never a pass.
- Name every shortcut in the PR body.
Acceptance and uncompleted scope
Gate/classifier/matter acceptance retains private recall 1.0, actionable recall at least 0.9, zero false matter links, strict parser/privacy guards and complete attempt/cost denominators. The observed three current-source quality misses remain visible; no label or threshold is changed. Rollup reliability and migration are explicitly deferred, not quietly marked passed. Broader nonmail direct-Anthropic inventory/removal also remains open.
Root lint/typecheck and affected package tests must cover the restored scope. The plan's broader three-repeat exclusion/resolver evidence is not replaced by a single frozen-helper run; this cleanup performs no paid or database evaluation. Two independent adversarial reviews precede any substantial merge, including a deletion audit. Prior reviews are not automatically reviews of this restoration. Merge and deployment remain kwiss's decisions.
No schema migration, reclassification, backfill, golden relabeling, cascade, shadow-run framework, UI change or new fallback. Generic routed transport/accounting and reviewed gate/classifier/matter fixes remain. Rollback uses the previous deployment and configuration.