All specs and plansNetDocuments

NetDocuments investigation — implementation plan

2026-09-17 addendum — v5 grounding and continuation (implemented in the lane worktree)

Opened in progress and completed on 2026-09-17; final verification and paid synthetic adjudication are recorded below. Historical implementation metadata and enablement status are unchanged.

  1. Implemented prompt/finalizer evidentiary rules without a runtime semantic judge: acts/status evidence is distinct from definitions, obligations and conditional scenarios, with explicit positive/pending and sufficient factual-answer cases.
  2. Implemented durable receipt origin/completed-work provenance, conservative legacy handling, retained/superseded partial publication, server-derived recorded coverage and handle-aware guidance. Added explicit historical work scope on replay without changing per-invocation accounting.
  3. Implemented monotonic terminal lineage and Lua-derived counter saturation without new generations, TTL refresh, money resets or limit changes. Research get suppresses terminal execution retries; revocation retains terminal authority after source-context withdrawal. Uncertain unpersisted terminal observations retain the existing bounded recovery limitation.
  4. Implemented four Opus 5 low synthetic fixtures and exact stored-window reporting with 65,536-byte evidence / 262,144-byte report bounds, omitted refs, truncation withholding, separate answer/finding review and non-semantic mechanical exit. The eval mirror follows per-session ownership and retained 3/30/300 capacity rather than an actor execution mutex.
  5. Final verification: runtime 5,194 Vitest + 139 and 77 Bun tests passed, 171 skipped; five scoped suites 369 passed with one opt-in PostgreSQL proof unrun; connector limiter eight passed; eval 72 passed. Fourteen temporary mutants failed at intended assertions and were restored. Concurrent retry survives refundable read saturation; definitive caps, revocation and unknown charged reads remain enforced. bun run lint && bun run typecheck && bun run eval:gate:cheap exited 0. Cheap router gated cases passed 4/4; four nongated boundary cases scored 0/5 versus 5/5 in an earlier run: variability, not a claimed regression fix or all-router-case pass. Memory passed 9/9 with a 48-active versus 20-cap warning.

Final pressure proof after the three-line coverage prompt correction: context 120,992; messages 92,724; system 4,974; tools 14,351; framing 8,192; rejected preflight 120,241; largest accepted preflight 101,952; largest actual model request 85,689 bytes. This run recorded 16 useful reads + 7 refreshes = 23 attempts, 0 denials, 35 HTTP admissions. Request limit remains 120,000 bytes; no budget or runtime policy changed.

Paid synthetic rerun c74b2c690f0018c353de174b53555e6e12100189c3af6359fb27b9a6477b2e72: four fixtures, six reports, six answers and 21 findings; independent Astra LOW and Fable MEDIUM semantic PASS. New corrective generation 1 incorporated one research read and two model calls; replay used zero models, zero research reads and $0 model cost, with four revalidations accounted. All displayed evidence was report-complete.

The first paid run failed whole-report acceptance only on unread-content characterization; the generic three-line prompt correction addressed it, and the final run did not repeat it. An immediately self-corrected lead ID is a minor presentation defect, not a material false claim. The factual zero-hit find result is not independently checkable from the report because query/hits are not exposed; only the operation is corroborated. Exact source-length assertions are not uniformly independently verified. This is semantic acceptance, not blanket metadata verification. Final gates passed with the nongated diagnostics noted above; no CI/PR, merge, deploy or enablement claim. This documentation worker made no paid calls or ledger writes.

2026-09-14 top addendum — direct-first research record, implementation plan (implemented in the lane worktree 2026-09-14; default-off)

Approved, amended after review. Concrete plan for the direct-first design addendum and its seven review amendments, which override any contradictory step below. kwiss approved implementation and the fixed 30-minute retention; the steps below were implemented on 2026-09-14 by the Fable HIGH worker (14 production files: 4 new, 10 modified; tests untouched, owned by the omp @write Astra HIGH seat), with the as-built deviations listed at the end of this addendum. Verification: independent focused suites 216 passed, 2 skipped, 0 failed (research 45 passed; investigation 124 passed, 1 skipped; state 47 passed, 1 skipped); full runtime 5554 passed; cheap eval coverage 11/11, retrieval 27/28, router 4/4, memory 9/9; actual MCP smoke 2 direct reads, save/get, Opus 5 low escalation with 3 findings, 3 citations, 2 documents in 109.7s; exact-quote save/get 4 calls, 2.8s on a 22911-character window; web production build passed. The feature is implemented and stays default-off behind APP_NETDOCS_INVESTIGATION_ENABLED; draft PR #423 is updated, never duplicated. Seat deviation (Devin Fusion Fable 5.1 high + SWE-2 medium, at kwiss's request) is recorded in the spec. Scope estimate: about 12 files; no schema, migration, grant, dependency or new env key (the research tool runs under the existing APP_MCP_READ_TIMEOUT_MS policy, so every save/get verification runs under a bounded deadline: an unfinished check reports unavailable, preserves state and never implies verified. Eight windows per save is the maximum, not a promise that all eight finish within the 45 s budget).

1. Provenance seam (packages/agent-runtime/src/tools)

2. Record store, tool and escalation (packages/agent-runtime/src/tools)

3. MCP surface (apps/mcp-server/src)

4. Tests: real behaviour through existing suites, no wording assertions

As built, 2026-09-14 — deviations from the steps above

Acceptance and release gates (implementation approved)

Acceptance on the owned read-only MCP server with synthetic identity: the original private question, unhinted, through a fresh direct Fable host and a fresh direct Opus host given the procedure; save, then get from another fresh host; a forced unfinished direct search escalated to Opus low carrying provenance and pending work, with no repeated-read claim in the import block; exact-citation rejection; revoked, changed, expired and cross-actor records; two concurrent escalations of one record; a lost escalation response retried from a fresh host recovering the same session; the existing deep-search regression. Gates: the package tests above; root bun run lint and bun run typecheck; bun run eval:gate:cheap (routing changed); independent fresh Fable and omp @review Astra reviews to convergence with the deletions pass; rebase onto current main without force-push; update draft PR #423; CI green; MERGE-READY, never merge or deploy.

2026-09-14 top addendum — pre-R2 actual run and verified repairs (still in progress, default-off)

Actual run (09:28 UTC, pre-R2 server, Opus 5 low on the approved OpenRouter Bedrock pin; Astra still rate-limited upstream): one native context of 26 entries across two generations; 12 model calls; 11 useful reads (8 search, 2 find-in-document, 1 fetch); 154 refresh reads; 24 limiter denials; 301 HTTP; ledger $0.850860 actual, $0 unknown. First receipt inconclusive (9 models, 9 reads, 81 refreshes, 276,681 ms, no findings, insufficient_walltime_for_model_call). The generation-1 finalization was a negative/limited-coverage statement, not a budget answer, and was refused because its coverage_statement (1,035 characters) exceeded the maxLength: 800 disclosed in the tool JSON schema; about 20 s remained. Only one source (about 39k characters) was partially fetched; neither budget workbook was reached. The generation finished server-side but its receipt never reached the client within 360 s; the exact delivery cause is unobserved. No semantic pass; model and route differ from the successful Fable CLI experiment, so context shape alone is not isolated.

Since the live run: R2 (coverage, double-resume) and R3 (text-only ACL) fixes landed in the lane tree and are verified by suites and converged Astra/Fable reviews (268 pass / 2 skip; Lua 1 pass; MCP 40 pass; root lint and typecheck exit 0), not by a rerun of the original question. Kept: source ACL, counters, ids, timestamp/lifecycle/test repairs. Rejected: ignoring file size; the hidden-limit inference. Release stays blocked by the failed acceptance and continuation-delivery uncertainty; PR #423 is a draft with a known conflict; no git action this pass.

Proposed next decision (not approved, not implemented): stop re-reading the entire history before and after each model thought; keep fresh access on new reads, on resume and at citation publication, same actor-bound history. Tradeoff: a source edit or revocation during an active generation would not retroactively remove already-read context before the next model call. The strict policy remains in force until kwiss decides; the recommendation rests on 154 refreshes for 11 useful reads, not on proof it finds the budgets.

2026-09-14 top addendum — native research agent cutover (in progress, default-off)

In progress / default-off; implemented in the lane worktree, not committed, not reviewed, not accepted. kwiss approved the fast-track cutover to a real internal research agent with cumulative context after a 67-second fresh Fable CLI experiment (20 read-only NetDocs calls, 4 cited documents, 7 of 8 exact quotes). Two independent plan reviews (Main and this worker) converged before implementation. Effort deviation, recorded at the start: this cutover was written by a Fable LOW worker seat, below the seat table's Fable high; every other routing rule is unchanged. Product model selectors (astra, opus, grok, muse, efforts) are unchanged.

What replaces what

Verification in this pass

Targeted package Vitest (synthetic DATABASE_URL/APP_DATABASE_URL): investigator 114 pass / 1 skip, state 47 pass / 1 skip plus the opt-in actual Lua case against the disposable Redis on 16383 (1 pass), eval 60 pass, capability adapter 36 pass, MCP tool contract 40 pass. Package typecheck clean; scoped lint 0 errors. The eval suite drives the real investigator over the synthetic JSON-RPC corpus with Anthropic-shaped wire messages, including the forced continuation and replay. No real provider or model call was made; no live acceptance is claimed.

2026-09-11 top addendum — approved deterministic-loop production integration

In-progress / default-off; approved for implementation after plan-review convergence, not implemented or semantically accepted. The user explicitly approved integration and continuation without interruptions unless post-implementation tests fail; no repeated user-approval gate. The top design addendum governs. Together these supersede prior unproven replacement acceptance and stale limits, including prompt-only/per-read scheduling, first-call-only proof and old 34-path/5,600-line ceilings. Preserve historical evidence and safety failures; none is current acceptance. This correction pass touches only the two top addenda and NetDocs STATE row, at most 120 added lines.

Reference and proof limitations

Private reference SHA256 7948acc3ba06dfe0301041103fbd3932c55d88a4529aacb46e80203ac7090c40 was checked. The clean unhinted run reports 3 model calls / 39 research / 38 refresh / 20 pages / 257 unique candidates / 12 distinct documents / 396,866ms / 418,225 microUSD, zero failures or throttles and fresh validation of every accepted citation. Six pages and 249 reads remained pending; cached continuation windows were withheld. This is one successful scenario, not exhaustive coverage, general reliability, a controlled speed comparison or production-tool acceptance. Keep private source, run artifacts and confidential names/IDs/quotes outside the repository; port behavior, not the experiment's rendered-text parsers or private corpus.

1. Review the contract before implementation

Main obtains two independent plan reviews before assigning code: fresh Fable high in Claude Code and omp reviewer @review Astra high. Planning is @plan Astra medium. Implementation is Fable high via Claude when available; the latest lane routing sends Claude exhaustion to Devin Fable 5.1 low before any different-model fallback, including the fresh Fable review leg. This supersedes the older Astra-writer degraded route below. Record only actual routing deviations with their reason; planned normal routes are not deviations and no new exhaustion probe or worker is launched here.

Reviewers must resolve budget classification, storage fit, interruption-safe invalidation, frozen dependencies and actual start/resume acceptance. Use current source and the unfinished branch's verified findings, not dated green counts. Substantial code review adds Grok xhigh for deletions; fixes use @fix Luna max, gates @gates Sol low, git Main. Both independent code reviews must converge after fixes before merge readiness. No merge or deploy authorization.

Review adjudication: Main verified and kept all six findings: mechanical drainage after model exhaustion, retention versus revocation, durable exposed-slice citations, three-slot money semantics, pre-model refresh admission and cooldown for every typed read. Keep byte-budget planning; reject the overclaim that independent field maxima must fit simultaneously or that contextBytes must automatically increase. The aggregate guard is intentional.

2. Extend existing durable state and accounting together

Owner file: packages/agent-runtime/src/tools/netdocs-investigation-state.ts. Extend investigationContextSchema, emptyInvestigationContext, invalidateInvestigationMaterial, command/reply schemas, createInvestigationState and investigationStateLua in one coherent contract. Queue records need candidate deduplication and originating-page membership, query/window pending positions, unprocessed page remainder, dependency fingerprints, fairness position, priorities/revisits and cooldown. Cumulative counters belong to the fenced whole-investigation ledger and are echoed as trusted progress, not model arithmetic.

3. Replace per-read model control with the fair scheduler

Owner files: packages/agent-runtime/src/tools/netdocs-investigation.ts and netdocs-investigation-prompt.ts. Rework createNetDocsInvestigator's step, execute, modelDecision, read, refresh, checkpoint, decision/final schemas and result construction. Keep the owned model transport, selected model/effort, reservations/settlement, request guard, zero retries and cancellation/drainage. Remove obsolete direct model-read tool dispatch and its prompt/callers, not a second optional path.

  1. Model decisions propose searches, prioritize exact returned IDs and request revisits. Code alternates page/document work and round-robins query pagination. Every candidate remains eligible, including zero-score hits; deduplicate memberships and reads. Model choices reorder work but cannot invent IDs, titles-to-ID aliases, scope or permission.
  2. Process bounded research batches between decisions, preserving the reference's fair ordering and explicit priorities. Use remaining global budgets and pending work to schedule the next decision/finalizer, not one model call per read. Do not exhaust all five decisions before evidence can be interpreted; finalization/correction spend the same ceiling.
  3. Before each model call, reserve typed-attempt capacity for the exact frozen exposed post-decision refresh set, including its dependency traversal, and wall time for bounded model plus refresh. This ports reference loop.ts:326–329, not an extra strategy. If admission cannot fit, disable further model calls and continue already-authorized mechanical work; do not spend the protected refresh capacity on unrelated dispatch.
  4. Checkpoint the selected pending item before dispatch; after success atomically persist its result, newly queued work, fairness position and cursor advancement. A forced yield resumes that exact frontier after fresh authorization. Empty/duplicate advancing pages continue; cycles or unavailable process-local cursors stay explicitly incomplete with bounded restart/deduplication, no counter reset.
  5. Apply the same persisted cooldown to EVERY typed read: research, pre-exposure, post-decision and final/replay refresh. Replace immediate refresh-throttle budget stops; retain the same pending item and phase, wait/retry only when retry-after and dispatch fit remaining guards, otherwise yield partial. Missing retry-after or deadline-crossing wait yields partial. Throttle is never permission withdrawal; actual permission denial invalidates dependent work. Count every typed attempt, including denials, separately from HTTP.
  6. Model-only exhaustion/admission failure disables further model calls, not already-authorized mechanical traversal. Drain pending reads under remaining read/HTTP/time/access/state guards and applicable financial guards; no sixth model call after decision five. Only a guard blocking that work causes a partial stop with the recoverable frontier and precise reason. No hidden page/window completion cap; an answer with pending pages cannot claim project completeness or absence.

4. Enforce live dependency and exact-exposure evidence

Owner files: runtime investigator/state/prompt plus packages/agent-runtime/src/tools/netdocs.ts. Reuse createNetDocsReadClient.read, never parse rendered MCP text. Source confirms readScope.run, readDocumentSession's scoped cache bypass and fetchNetDocs's requested-ID check. Custody-only authorize is not a fresh document ACL; session timestamps are not permission evidence. Preserve ordinary direct cache behavior and provider-confirmed offset/source fields.

5. Migrate the existing entry surfaces, not model semantics

netdocs.ts:createNetDocsDeepSearchTool renders the same authorized findings and partial-coverage contract on product/MCP while keeping audit data redacted. Change packages/agent-runtime/src/index.ts only if exported types change. apps/mcp-server/src/tools/netdocs-deep-search.ts:netDocsDeepSearchDef and apps/mcp-server/src/tool-deadline.ts:invokeWithinMcpToolBudget change only where the state/result/budget contract requires it. Preserve the 330,000ms default deep-search envelope (activeMs + 2 * cleanupMs), 360,000ms clamp, shorter outer deadlines and ownedSettlement; 600,000ms is lifetime, not a larger RPC. Preserve default-off catalogue/call-time guards and public start-selected model/effort pinning through continuation; no REST or auth/connector-storage edits.

6. Verification that earns acceptance

Existing files: packages/agent-runtime/scripts/eval-netdocs-investigation.ts (createEvaluationState, runSyntheticEvaluation, assessSyntheticResult, evaluationExitStatus), its existing test, runtime netdocs-investigation.test.ts/netdocs-investigation-state.test.ts, affected capability/direct/window tests and apps/mcp-server/test/netdocs-tools.test.ts. Use existing typed-read/provider/state seams. Port only meaningful behavior from the nine private synthetic groups; do not add the private receipt parsers, a fixture corpus, a huge DB suite or tests that merely pin implementation wiring/wording.

Behavioral proofObservable failure it must catch
Fair pagination beyond old limits; zero-score retention; duplicate candidates; multiple queriesPage work starves fetches or vice versa; a low-ranked hit is discarded; duplicate/empty advancing pages strand a later unique result.
Forced continuation during page processing, document progress and cooldownUnprocessed page entries disappear when cursor advances; pending offsets, fairness, counters or cooldown reset; a resume repeats completed work as fresh research.
Typed cooldown vs permission denial, missing retry-after, deadline-crossing waitA throttle revokes evidence, permission denial retries as a throttle, or a denial escapes total-attempt/time accounting.
Origin dependency refresh, visited-set traversal, title/metadata-only changes, interruptionStale dependencies authorize derived work, repeated traversal wastes unbounded reads, or an interrupted invalidation resurrects an old query/receipt.
Frozen exposure, explicit revisit, hidden retained text, exact exposed IDs/quotes and live final/replay refreshUnrefreshed replacement slips into a prompt; session material is accepted as fresh; title alias or unseen quote passes; unconfirmed location is emitted; revoked evidence is replayed.
Reference fetch/exposure parity and Unicode boundariesA 6,000-code-point retained cap silently halves the 12,000-character fetch; exposure exceeds 3,500 characters per slice, 8 windows or 12 pages; non-BMP clipping breaks quotes/offsets; retained/replay inventory escapes total byte guards.
Admission, settlement and every independent bound across resume/replayUseful reads are confused with attempts; 5/80/200 or lifetime resets; HTTP/reservations bypass admission; failed output loses actual cost; serialized frontier/replay overflows silently.
Shortened outer deadlines and cancellation during model/read/refresh/checkpoint/closePost-cancel dispatch, premature output, lost useful resumable progress, or cleanup used as extra research time.
Fifth decision followed by further authorized reads; model admission lacks refresh attempts or wall timeMechanical traversal stops with model-only exhaustion, a sixth model call occurs, or a model is admitted without capacity for its exact frozen post-decision refresh set.
Pre-exposure, post-decision, final and replay refresh throttles across yield/resumeImmediate budget stop replaces a feasible persisted cooldown, pending refresh phase is lost, or throttle is treated as permission withdrawal.
Routine retention eviction versus explicit revocation/change, including replayOrdinary eviction empties valid snapshots; revocation leaves derivations reusable; any cited source bypasses independent live validation or stale replay is returned on failure.
Quote exists only after the 3,500-code-point exposed boundaryTS, durable Lua resolve or eval accepts the hidden quote; an exact exposed quote must still succeed after fresh validation.
Unicode/state pressure with pending page remainder and replay snapshotsActual serialized bytes exceed context/envelope/120,000-byte wire guards, stop space is unavailable, candidates disappear, or resume skips unprocessed cursor work. Exercise feasible allocations above, byte-heavy escaping and smaller dynamic selections with honest omissions, not fixed quotas.
Three money-turn slots and old late settlement after content-version cutoverGeneration count changes, attempt.turn is reassigned, money keys/liabilities disappear, or settlement charges a newer turn instead of the original one.

Do not promote unfinished branch repairs to a green baseline. Current source's text-only window-change check, end-of-walk invalidation and lack of an exact frozen post-model dependency refresh need explicit regression proof. Retain the verified historical causes as acceptance checks even where a repair exists: cold authorization outside owned timing; connector SSE waiting for EOF; inherited implicit model streaming; stale persisted candidate exposure; sibling work deleted on replacement; advancing-page stagnation; MCP count-only answer; missing supported search fields; pre-dispatch body-guard unknown reservations. Review actual causes and deletion impact, not just final happy-path output. Existing narrow connector/MCP tests may cover these; no broad new infrastructure suite.

Offline/package gates: relevant agent-runtime and MCP package tests, connector tests where changed behavior reaches that seam, then root bun run lint and bun run typecheck. Run the targeted cheap routing gate and bun run eval:gate:cheap where required by retrieval/routing changes within established lane authorization and accounting; record any exact waiver, never relabel paid subgates offline. The eval reference mirror does not prove Lua atomicity: exercise the affected existing isolated Lua cases on the permitted disposable target, never the shared database. Main resolves tool/repo-provided configuration without renewed user permission. No large database matrix.

Real acceptance, after fixes and reviews: run the original unhinted question through actual north__netdocs_deep_search using Opus medium, the proven route, following returned continuation handles rather than manufacturing state or bypassing the tool. Also run the existing Astra acceptance flow without changing its model/effort semantics. Recover and cite the relevant identity/evidence chain and supported budget contents; distinguish existence, exact filing scope, latest and approval, with no invented flattened-workbook associations. Report partial coverage and pending/withheld work. Transport success, private bypass and direct fallback are not acceptance. Record per-call/end-to-end time, models, useful reads, attempts, refresh, HTTP, limits, actual/unknown/reserved cost, fresh citation validation and unsupported claims; keep confidential evidence private. Refresh this evidence after fixes, not before them.

Approved continuation, engineering decisions and ownership

The approved integration uses 600,000 active ms lifetime, 300,000ms per call and 5/80/200 ceilings, not a latency SLA. Read-only acceptance under the established lane identity and $50 lane allowance is already authorized. Continue after plan-review convergence without repeated approval or interruptions unless post-implementation tests fail. Main resolves configuration/accounting from available tools/repo information without increasing caps. Dynamic retention requires proof of byte fit; the existing three money-turn slots are unchanged. Name and justify reference differences; no extra search strategy, retry platform, default-model change or broad reliability claim follows from the private run.

Implementation completes only when scheduler, state/Lua/eval contract, consuming types/renderers and affected tests agree, behavioral checks and real-tool acceptance have fresh evidence, and independent reviews converge. Keep in-progress/default-off until accepted; enablement remains separate. Remove obsolete per-read code and temporary proof artifacts after verification, preserve historical documentation and ordinary direct tools. Main owns git and later merge/deploy decisions; none is authorized here. This documentation-only seat stops at PLAN-CORRECTIONS-COMPLETE.

Latest superseding plan — completion-driven investigation, design only

In-progress / default-off. No implementation assignment or acceptance. The user's ok let's try that, then we don't care about the limits to be honest it's a critical feature so we need that, supersede the 34-path/5,600-LOC and 6-model/8-research/35s/3-turn/105s acceptance ceilings. Implement for reliable completion through normal continuation, not a larger arbitrary search quota. The completion-driven spec governs; prior claims below remain historical evidence, not competing instructions.

Routing: Main reports the current Fable probe failed out-of-usage: degraded: fable, no repeat probe. Author @plan Astra medium matches config, not a deviation. Under WORKFLOW's prose/degraded rules, implementation when assigned goes to @write Astra high; plan review requires fresh Astra medium plus Grok xhigh as the different-model/family leg. This is the documented degraded route, not completed reviews. Gates remain @gates Sol low; substantial implementation review requires Astra high and Grok xhigh deletion review plus a fresh same-family session under the degraded rule. Old Astra-only/medium overrides below are history, not current routing. No workers launched here.

Transport addendum — 2026-09-11 multi-model bounded transport (in progress, default-off)

Scope. Owners: packages/agent-runtime/src/model.ts (registry: supported efforts, default effort, temperature and routing policy), model-pricing.ts, tools/netdocs-investigation-selection.ts (pure selector → canonical pair, leaf module), tools/netdocs-investigation.ts (refusal order, pinning, accounting), tools/netdocs-investigation-state.ts (pair stored on the envelope, echoed by every claim) and apps/mcp-server/src/tools/netdocs-deep-search.ts (advertised V3 choices, pre-access refusal). No cursor, DB, cap or continuation-format change; the completion-driven redesign above remains design-only and open. Chat default, chat picker and every Anthropic construction path are unchanged.

Selection contract. A start accepts optional modelastra | opus | grok | muse and optional effort. Supported efforts and defaults come from the registry, never from the tool: Astra and Opus 5 low, medium, high, xhigh, max (defaults low / medium); Grok 4.6 low, medium, high, xhigh (default low); Muse Spark 1.2 minimal, low, medium, high, xhigh (default low). Omitted model keeps the env/default direct-Anthropic policy (AGENT_NETDOCS_INVESTIGATION_MODEL, else claude-sonnet-4-6); a direct-Anthropic model runs with reasoning disabled, internal effort none, which is not a public selector. Fable and the rest of the catalogue are not selectable. An unsupported pair is refused before authorization, state or budget on both surfaces. Resume takes only the existing handle and optional guidance and rejects model/effort. The V3 MCP schema advertises the choices; the canonical strict union still validates first.

Pinning. Every new production or eval start stores the canonical model id and effort; every claim echoes them and the investigator re-validates the stored pair before replay or research, so a later env change cannot switch a running investigation. Old envelopes with both fields absent resolve today's default; a partial or unsupported stored pair stops configuration. accounting.modelId / accounting.effort report the effective pair.

Routing. Astra Azure-only ZDR pin; Opus 5 Amazon Bedrock-only ZDR pin; Grok 4.6 xAI ZDR endpoint under the shared policy; Muse Spark 1.2 the sole approved non-ZDR exception (only: [meta], zdr: false) with data_collection: deny and no fallback retained. Every other unpinned route stays refused; an unavailable route stops configuration, never heals to a default. Chat posture is per model: Astra and Opus 5 are utility/eval-only; Grok 4.6 and Muse Spark 1.2 keep their existing, intentionally preserved ordinary-chat support, neither added nor withdrawn here. The investigation's bounded effort stays utility/eval on every route, with custody and no-fallback unchanged.

Bounded request. Complete serialized body at most 120,000 UTF-8 bytes (the 49,152-byte figure first written here is the historical cap; the canonical contract §6 carries the current 120,000); 4,000 total output tokens including reasoning for every OpenRouter route (direct Anthropic keeps 2,000); zero retries at both the LangChain and SDK layers; non-streaming; the pinned effort sent as OpenRouter's native reasoning: {effort} because the SDK heuristic drops effort for vendor-prefixed slugs; effort raises no limit. Temperature follows the registry: Astra and Opus omit it, Grok and Muse keep the supported setting. A final guard on the SDK-serialized body enforces the exact stripped slug, the model's provider block (data_collection: deny, allow_fallbacks: false, the pin, zdr: true except for the Muse exception), usage.include, tool_choice: auto, parallel_tool_calls: false, and rejects caching directives, model aliases, routing keys and output overrides before dispatch; the existing 404-marking OpenRouter fetch delegates to that guard, then to the owned fetch.

Accounting. The state reservation stays 3,030,000 microUSD per attempt. A bounded-price capability (wire bytes as the input-token ceiling) must fit inside it. At the current 120,000-byte cap with 4,000 output tokens: Astra 1,540,000, Opus 5 770,000, Grok 4.6 264,000, Muse 167,000 microUSD; direct Sonnet with 2,000 output tokens 390,000 microUSD; all inside 3,030,000. Conservative regional ceilings 11/55, 5.5/27.5, 2/6 and 1.25/4.25 USD per million, verified 2026-09-11 separately from the unchanged global price date. The historical 760,672 / 380,336 figures were computed against the 49,152-byte cap and are retained as history only; the existing pricing regression is updated by the core worker, not by this document. OpenRouter actuals come only from the raw owned payload's finite non-negative usage.cost (cache and reasoning included), never from token estimates; missing or invalid cost leaves the attempt unknown. Normalized and native finish reasons plus message.refusal map to refused / provider; the owned fetch and body drainage are awaited before settlement.

Proof, 2026-09-11 real MCP smoke on frozen source (Main-owned, paid, env-selected Astra/Opus). Astra low: 3 receipts, 11 model calls, 18 research, 28 HTTP, $0.327326. Opus 5 medium: 3 receipts, 14 model calls, 13 research, 23 HTTP, $1.120465. Both: zero unknown attempts, zero unresolved reservations, correct accounting.modelId. This proves transport, start/resume and accounting on both routes and stays valid for the same pinned routes. Per-start selection smoke (server default Sonnet, actual tool calls): grok/high and muse/minimal, three receipts each, canonical model and effort correct throughout, both inconclusive with no findings. Those runs exposed a pre-dispatch accounting defect (the local body guard can retain an unknown reservation) now being fixed by a separate worker; no financial proof or final test counts are claimed for these routes until Main re-runs. Known semantic failure: no run produced a finding; semantic retrieval acceptance remains failed and open. Feature remains in progress and default-off.

1. Replace fixed generation/depth admission in the existing state module

packages/agent-runtime/src/tools/netdocs-investigation-state.ts: update createInvestigationState, command/reply schemas and investigationStateLua together. Replace three precomputed hashes/turn arrays and fixed model/research/HTTP/lifetime-active ceilings with demand-minted fenced generations, dynamic financial turn entries and measured counters. Continue within fixed 30-minute expiry, existing limiter, monetary caps and per-call deadline. Keep three replay receipts as bounded retention, not a continuation ceiling. Preserve HMAC identity binding, replay-only claims, CAS, stale-owner exclusion and crash takeover.

Add an authoritative allowance snapshot to claim/admission/checkpoint replies, validated at the state seam: deadlines, expiry, consumed counters, monetary headroom including unresolved reservations, limiting reason and revision. Use a read-only fenced refresh only where existing replies cannot give current admission information. No model-derived budget or untrusted public cap override. Retain $5/turn, $15/lifetime, $50/org-day and full $3.03 reservations; reconcile Lua's observed $5 eval comparison with the separately approved $50 aggregate eval policy. Preserve historical ledger identity and late settlement across content-version cutover; do not reset financial keys.

2. Make the investigator pursue evidence and preserve useful progress

packages/agent-runtime/src/tools/netdocs-investigation.ts: update modelDecision, decision/final schemas, step, execute, validate, checkpoint and result construction. Track cited identity, budget contents, filing scope, latest and approval as separate obligations with provenance. Decisions identify their obligation and expected gain: read anchor → reconcile linked identity/contradiction → discover and inspect alternatives → answer supported claims with gaps. No customer-specific ID, alias or parcel rule.

Build canonical progress over validated state before the current last-24-node/14KB presentation trim. Remove the 25-line repeatedSearchSameLeads projection and its presentation-order assumptions when the replacement owns that contract. Compare equivalent searches with completeness/uncertainty intact; report stagnation to replanning, not fabricated exhaustion. Carry unresolved obligations through compaction and resume; remove obsolete wording-pinning tests instead of preserving a second feedback convention.

Keep discovery = [] single exposure, opaque internal candidates, fresh requested-ID-matched ACLs and transitive pruning. Choose a currently exposed lead for immediate fetch; if it was not fetched before yield, reacquire via the saved authorized query rather than resurrecting its ID. Promote only read-supported identity links and claims. netdocs-investigation-prompt.ts explains the new control/progress fields and evidence sequence; it already requests fetching/aliases, so prompt edits alone are not the fix.

Send trusted remaining-budget/progress controls outside UNTRUSTED_DATA after ACL costs and model admission are accounted for. State snapshots guide planning but never bypass atomic admission. Preserve Sonnet 4.6, 2,000 output tokens, zero hidden retries, 48KiB actual SDK request guard and source injection boundaries; no model upgrade is assumed.

3. Cooperatively yield before the hard stop; retain transport ownership

Replace private graph recursionLimit 9 with an explicit serial completion/progress/deadline loop inside the same module. Initially keep 35s hard-active/40s cleanup and begin cooperative yield at 30s or earlier when the planned operation threatens the five-second active receipt reserve. New research stops at yield; fresh receipt ACLs/checkpoint run within the hard-active envelope. Preserve suppression after hard abort and await withNetDocsConnection close/drain before publishing. Never use cleanup time for research.

apps/mcp-server/src/tools/netdocs-deep-search.ts:handler, runtime run and state claim deadlines must consume one shared policy, not three drifting constants. Keep tool-deadline.ts:invokeWithinMcpToolBudget's 45s default, 5–55s override clamp and ownedSettlement semantics, plus tools/netdocs-context.ts:invokeNetDocsTool identity/cancellation propagation. Shorter outer deadlines shrink all stages. Retain direct 15s/25s operation and 6s provider defaults; no transport-wide timeout lift or new background service.

Extend NetDocsInvestigationResult with semantic progress and a distinct cooperative-yield reason; migrate both existing adapters/renderers and their descriptions together. Return cited findings, outstanding obligations/actions, scope limitations, measured accounting and an opaque expiring continuation. Normal same-tool resume consumes the saved frontier after fresh authorization; duplicates revalidate/replay, not rerun. Finalized supported answers do not force pointless continuation solely because latest/approval remain unknown. Host may resume or offer the user continuation; this plan cannot promise host obedience.

4. Prove behavior through existing test and eval seams

Acceptance: the original private question (kept outside the repository) independently recovers the cited identity chain to the independently linked project identity and both budget contents (two financing alternatives) with relevant source citations and stated limitations; clearly distinguishes project existence, observed filing scope, latest and approval; 43/55 is incomplete coverage, not proof of a third-matter filing. Record per-slice/end-to-end latency, model/research/ACL/HTTP counts, actual/unknown/reserved cost and unsupported claims. The historical experiment's one-deep-call/12-tool harness restriction is removed in the future acceptance driver, not mislabeled a full product failure. Direct fallback success does not pass the investigator.

Review, evidence and authorization gates

Main owns fresh independent plan review and subsequent implementation assignment. After implementation, update the existing canonical provider contract and affected tests/callers; this documentation-only seat leaves them unchanged. Required gates remain root lint/typecheck, affected package tests and eval gates per WORKFLOW/config. The scorecard records that cheap-eval subgates may perform paid/provider/DB work: Main must preserve authorization or record an explicit waiver, not run them blindly. No prior green suite is proof of this proposed behavior.

Main reports latest pair: Opus 5 A 12 direct calls/66.211s/no leads; B 10 direct + 1 deep/106.940s, both expected proformas through DIRECT fallback. Deep: 4 Sonnet calls/8 research/11 HTTP/30.961s, inconclusive budget, zero findings/aliases; requested resume and later reads blocked by harness. Latest ledger: 91 actual attempts, 2,898,528 microUSD, zero unknown/reserved, $50 total authority unchanged. No replay/model/provider call in this design pass.

Historical record below — prior caps, routing and sequencing superseded

Current disposition — invocation reuse implemented; paid core acceptance FAILED

Invocation-scoped reuse remains implemented and offline verified. Approved four-line prompt correction: independent Astra medium nd-prompt-a p6E / nd-prompt-b p6F CLEAN, source-only; no cross-model or semantic acceptance. Main's post-prompt root lint exit0/52.25s, typecheck exit0/64 of 64/52 cached/Turbo1m53.273s/wall129.23s. Feature remains in-progress/default-off; prompt correction INEFFECTIVE on the actual MCP run.

Latest Main-personally-executed /tmp/netdocs-deepsearch-main.RBZWm93I/client.ts --paid called actual north__netdocs_deep_search, unchanged budget question: capture.jsonl proves 1 deep call/6 nested dispatches/1 receipt; deep-search-receipt.json is partial/budget, 6 models/7 research/19 ACL/57 HTTP/19,673ms. Repeated original query; budgets not fetched. Full command 21.76s/exit0 is transport ONLY, not semantic acceptance or tool-research success. Increment190839 microUSD; cumulative 2563944 microUSD/81 actual/0 unknown/0 reserved; paid retries STOPPED, original $50 TOTAL unchanged. Prior paid-core failure remains historical:189687 incremental,2373105 cumulative,75 actual,0 unknown/reserved.

Exact captured-response replay through actual SDK/core/fixture/ACL and a removed cloned ledger is RED: 6 models / 7 research / 19 ACL / 57 HTTP, exit1, 2.39s wall with fixed accounting clock. Requests3–6 retain source linkage and contradiction; SDK translation preserves the repeated search. Per-query noProgress=false does not establish global novelty. Command and source citations: scorecard, latest paid-core diagnosis.

Prior natural-host localhost synthetic-MCP direct-tool success found both budgets but did not select deep_search; it is not deep acceptance. Prior prefix counterfactual 5 models/6 research/16 ACL/35 HTTP/1,518ms with bothBudgets=true remains offline feasibility only.

User approval now authorizes exactly four generic lines after the existing Follow aliases line in netdocs-investigation-prompt.ts; this supersedes the earlier no-prompt-edit restriction only for those lines. Main's baseline 5,562 + 4 = 5,566/5,600 gross production additions, 34 paths, 34 remaining; no other product/test, fixture or runtime/financial-limit change.

Hermetic investigator 187 passed/1 skipped, 5.98s; prompt Prettier/ESLint exit0 and prior private offline TS/catalog/preflight/SDK checks remain preparation evidence. Main's paid flow exercised actual SDK client → MCP handler → factory → core/model SDK → source reads → receipt, with synthetic auth/principal/state/source, not real-user documents. All old failed captures and prior direct-host success remain distinct. Read-only per-query feedback proposal and LOC estimate are in the scorecard, awaiting Main/user approval before any contract edit; no product fix or extra calls here. Metadata in-progress/default-off, generated STATUS/index and canonical contract unchanged; no PR/commit.

Latest approved implementation: model-projection-only repeatedSearchSameLeads compares all supplied arguments and equal observed lead sets only among retained validated certain complete first-page records; excludes failed/pending/restarted/cursor-history/truncated or 200-ID records. No IDs, receipt/state schema, noProgress, pagination/retry gate or prompt change. +25 production: 5,591/5,600, 34 paths, 9 remaining; +75 existing-test lines. RED before fix; final hermetic investigator 191 passed/1 skipped (6.23s), scoped format PASS/lint0 errors12 warnings, fresh private strict TS/catalog/SDK offline PASS. Main-only paid after fresh dual review using /tmp/netdocs-repeat-main.oIeTlbwr/client.ts; not run. Ledger unchanged 81 actual/2,563,944 microUSD/0 unknown or reserved. In-progress/default-off; no semantic acceptance. Prior failed captures remain binding; scorecard records commands and capture decision.

Historical preparation below — superseded sequencing, preserved evidence

The following DESIGN READY, not-implemented, review-pending, prior-spend and next-review statements describe earlier handoffs, not the current disposition above. Their measurements and failures remain evidence; technical contracts and acceptance criteria remain binding.

Status: in-progress, default-off; connection correction DESIGN READY for fresh dual review, not implemented or accepted; paid execution remains STOPPED. Author: Astra medium under the existing explicit Astra-only/medium override; not cross-model review. Updated: 2026-09-10 (original plan 2026-09-08). Main's recorded base/HEAD is 6bd0ae57, with origin/main 11 commits ahead; integration before PR remains Main's responsibility, not current-main verification. The earlier c789e830 rebase from 3aa016ea and eight byte-identical docs remain historical evidence. Related: design; canonical contract.

Substantial despite small plan: public tool names, model instructions and resumable authorized state. Factory cleared all NetDocs surfaces; no reservations. R2 A=w7E:p3X and B=w7E:p3Y corrections and prior historical dispositions remain in the scorecard. Independent R3 A=w7E:p30 and B=w7E:p41 returned CLEAN TO IMPLEMENT; Main read both full reports, accepted plan convergence and explicitly assigned STEP1 after rebase. No implementation-review convergence or enablement is claimed. Main's workbook format gate PASSED at 3aa016ea; new-loop behavior remains unproven. Earlier draft/unassigned permission-stage text below is historical, not current implementation status.

Current scope and evidence — 2026-09-10

Approved authority: kwiss's exact Approve invocation-scoped ok let’s go / yes ok: invocation-scoped reuse, 34 implementation files / 5,600 total normally formatted gross production additions. Main's reported baseline remains 33 actual files / 5,230 additions; 370 remaining, no deletion credit or feature cut. Earlier 33/5,500 and smaller allocations are historical. This assignment changes design/docs only; no schema, grants, migration, dependency or product edits.

The owner granted isolated fixture execution. Prior full runtime verification reported 5,375 passed / 111 skipped, the trajectory script exited 0, and earlier actual Lua/PG verification reported 2 passed. These are historical results before the current cap change, not reruns or current-main acceptance. Latest cardinality repair uses native auto plus disable_parallel_tool_use, preserving one-decision, refusal and serial-plan semantics. Affected verification: 249 passed / 1 skipped; root lint and typecheck each exited 0. Two fresh independent Astra-medium delta reviews were scoped clean under the explicit Astra-only override: not cross-model review, full paid acceptance or release authority.

September 10 kwiss authority: yes tell it it’s ok to spend more like 50$. The synthetic-model EVAL ceiling is $50 TOTAL, including all historical actual cost, unknown costs and outstanding reservations; not an additional $50 or Fable authorization. Runtime limits remain $5/turn, $15/lifetime, $50/organization UTC-day. Main reconciled 64 actual calls totaling 2,025,969 microUSD before today's first case; the persistent ledger must never reset.

First positive synthetic core-gate invocation today — failed acceptance: 22.00s, 5 model calls, 64 actual HTTP attempts, stop reason budget, findings 0, bothBudgets=false. Raw SDK sequence: first search, then fetch memo plus access document, then repeat search. No later fixtures or resume ran: the wrapper intentionally stopped after the first invocation; CLI exit 2 and row unrun are not a completed paid matrix. Added cost 157,449 microUSD; cumulative 2,183,418 microUSD, 69 actual calls, 0 unknown, 0 reserved. Paid work is STOPPED for unpaid captured-replay diagnosis. No acceptance, PR/CI, merge, deploy or enablement authority; no live calls allowed.

Approved invocation-scoped repair — corrected design ready, implementation pending

Latest disposition, superseding the diagnosis-pending sequencing above: Main accepts source-context loss hypotheses as rejected. Unpaid captured-response replay exposes the memo/linked alias and conflicting access text in reconstructed SDK requests 3–5. Original outgoing requests were not captured: no historical byte-identity claim. Fetch coverage returned=0 is count ambiguity, not missing source text; no storage/ACL patch is justified.

Actual-prefix counterfactual FAIL: preserve the actual initial search and BOTH memo/access fetches, then search the passage-supported alias and fetch both discovered budgets. Real SDK plus existing synthetic reads/state: 4 models, 6 research reads, 10 revalidation attempts, 64 HTTP, 1.216s active elapsed; stop budget, grounded fallback findings 1, bothBudgets=false, exit 1. Budget fetches succeed at HTTP 49 and 61; next memo ACL initialize/notification/GET consume 62–64, preventing the fifth model/finalizer. Exactly 48 handshake requests plus one catalogue, two searches, four full fetches and nine dispatched ACL fetches; the tenth logical ACL attempt cannot dispatch its tool call. This is not acceptance or proof that real-model latency fits 35s.

Tiny unpaid native-SDK lifecycle proof: two serial RPCs share one initialize/notification/GET; PASS, exit 0, 0.47s. Native close aborts/rejects before a held local fetch settles, so close alone is insufficient: pending fetch/body work must drain. A closed transport cannot restart; reconnect needs a new Client/transport. This proves lifecycle facts, not the proposed implementation or semantic success.

Corrected interface: connector withNetDocsConnection<T>(run: (close: () => Promise<void>) => Promise<T>), root export; private ALS owner, one serial generation, no SDK/context/resource bag or global close seam. At existing investigator timing callback, wrap run(args, outer, started, close); add its required close parameter. In existing final finally, await terminal idempotent close/drain, then revoke provisional publication on signal/deadline before removing timer/listener. Findings, continuation and duration/event remain after that seam; wrapper finally awaits the same close promise as backstop. No 191-line reindent.

Successful RPC intentionally finishes/cancels and drains only operation POST bodies without waiting EOF/timeout or aborting retained transport. GET fetch/body ownership belongs to the generation. Sticky generation-bound GET/native SDK failure rejects the active same-generation operation, retires idle immediately and awaits idempotent retirement before allowed replacement; old callbacks cannot poison new work. Ignore only intentional POST finish, expected GET405 and already-retiring close callbacks. Block idle dispatch including automatic SDK ping/error replies. Preserve fresh ACLs, binding/rotation, limiter, HTTP reservations, deadlines, byte limits and direct defaults; no new retries or remote rollback claim.

Planning ceiling, not measured fit: client 320 + root export 1 + investigator 33 + canonical contract 8 = 362, plus 8 unallocated = 370; 5,230 + 370 = 5,600. Count every formatted replacement/reindent without deletion credit; stop if measured implementation exceeds authority. Root export consumes approved file 34. No prompt edit planned, no new repository file or raised runtime limit. Exact lifetime, race handling and proof cases are in the private connection-design correction supplied to reviewers; two fresh independent Astra reviews remain required.

After design/scope approval, prove the unchanged actual-prefix first-call both-budget criterion within 6 model / 8 research / 64 HTTP / 35s, including mandatory ACLs and cold protocol overhead; test serial isolation, current-operation guards, lease rotation/recovery, persistent GET and held-close drainage without weakening existing acceptance criteria. Prior root gates are historical after the cap change; no fresh root-gate claim.

Approved cap-test repair, separate from the unapproved connection design: Main's affected run reported 248 passed / 1 failed / 1 skipped, 75.30s: the unknown-provider-failure test expected one dispatch but got nine because the larger cap permits more unresolved reservations. Only the existing eval test was repaired: valid historical settlements, each at most one reservation, leave exactly EVAL_RESERVATION available under EVAL_CAP. Assertions retain one dispatch, unknown reservation, later cases unrun and all historical charges; the incidental reason-string assertion becomes ledger-state assertions. No production stop policy or expectation of nine calls. Post-fix targeted test 1 passed / 64 skipped, 2.69s; full affected investigation/eval suites 249 passed / 1 skipped, 68.80s, exit 0. These are hermetic test results, not actual-prefix feasibility or semantic acceptance; money and pending design/scope approval are unchanged.

Final cap-test form: split the seed's nested reservation call into a local attempt variable to stay within 13 added test lines; no behavior or production change. Formatter reports unchanged. Re-ran both affected suites on that final form: 249 passed / 1 skipped, 69.25s, exit 0 (formatter plus test command wall time 70.83s). The preceding 68.80s result is the earlier passing run; the targeted test ran once.

Next: fresh dual review of the corrected design, not financial reapproval or implementation. Main accepted both reviewers' four findings: pre-close publication, non-EOF successful POST, generation-bound GET/native failure and 191-line reindent cost. Historical NOT APPROVED/pending-allocation statements record their earlier dates, not current authority. All failed paid/replay/counterfactual evidence and acceptance criteria remain intact. $50 TOTAL; $2.183418 actual / $0 unknown / $0 reserved, 69 actual settlements; ledger/captures untouched. No tests, paid calls, services, git, product/prompt or unrelated telemetry changes in this correction.

Historical September 8–9 repair handoffs

All “current,” “latest,” ownership, allocation and unrun-proof statements in the next three paragraphs refer to their original handoff dates, not September 10. Their failures and bounded successes remain evidence; the current scope, executions and stop above supersede their sequencing and ceilings. Acceptance contracts remain binding.

2026-09-09 current update, superseding the following historical slice allocations: Main relays kwiss's exact Allow 5,500 formatted additions selection via functions.ask; 31 implementation files including contract, no deletion credit, no feature cut. The eval/proof slice is partially resumed: initial handle/HTTP/fake-bearer repairs and a real typed-read/graph long-window regression are written. Latest focused three-file result is 219 passed / 1 failed / 1 skipped; empty-page-long first-call grounding remains red at the 64-HTTP limit. Long omitted-length A60000/B44878 first-call grounding and bounded wire pass separately, not Lua/resume proof. Further proof expansion stops for Main's test-allocation reassessment: current eval tests project to 579 formatted lines of approximately 900 total. No PG child proof or expanded Lua tests yet; no paid/infrastructure execution here. Exact counts, authority, limitations and unchanged remaining requirements are in the scorecard. Status remains in-progress; no acceptance or review convergence.

Latest bounded writer handoff, superseding the preceding RED/allocation stop: Main reassessed test additions to 1,600. Owned hermetic suites pass 236 tests with two real-target skips (20.08s); eval script/test are normally formatted at 1,487/958 lines. Guarded PG and expanded Lua proof code are written, not executed. Long-window first-call grounding passes at 40,216 maximum wire bytes; resume makes a sixth model request but stops storage in the bounded reference adapter, not successful-resume acceptance. No core wire-cap fix was warranted or made. Main owns real proofs, paid semantic evaluation, remaining formatting/gross-scope measurement, independent reviews and gates. Full feature remains in progress and default-off; see the scorecard for exact evidence and opt-ins.

The full feature continues unchanged under the 31-file / 3,600-gross-production-addition ceiling, not a target. The scorecard preserves the exact authority relays, superseded sequencing hints, source-confirmed eight repair contracts and measured gates. This writer owns only the eight-repair slice and bounded model-test isolation, with at most 890 additional production additions from 1,744; no new eval files or infrastructure execution. Preserve existing reservation/settlement/replay behavior and the injected state interface. Local cancellation means awaited LOCAL teardown and no later dispatch, never remote rollback. Main proposes two later eval files rather than three, preserving all accounting/tests and logical contracts; no eval implementation is claimed here. Steps below retain their planned order and historical acceptance requirements; unchecked does not mean no source implementation exists.

Routing deviations only

Explicit kwiss September 8 override, Claude subscription out: write Astra high → Astra medium; review Astra high → Astra medium; review-fresh Sol max → fresh Astra medium; substantial deletions Grok xhigh → Astra medium deletion audit; fix Luna max → Astra medium; gates Sol low → Astra low. Two reviewers remain independent, fresh and not the author; the same-family-only arrangement is an explicit override, not cross-model review. Orchestrate Astra low and plan Astra medium already match .omp/config.yml and are not deviations. Production model configuration is independent.

Additional explicit deviation: read-only scout Astra medium instead of smol Luna medium under kwiss's Astra-only override. Other routing is unchanged; Main owns integration, scout only reads.

Five implementation steps (acceptance unchecked; sequential ownership)

1. Establish safe read and model seams — Astra medium

After assignment, extract typed read outcomes within packages/agent-runtime/src/tools/netdocs.ts, preserving direct renderers, prepareOperation/runWithRecovery/window semantics. Bind stable-per-bearer leaseId plus actor/org/credential-owner/connector/accessMode/scope/endpoint, never custody updated_at; retain fresh resolution/rotation invalidation. Inject access/lease acquire-refresh-reauth/read transport/limiter/clock plus narrow TIMING and EVENT SINK dependencies, with lazy production DB/custody/timing/event defaults and a lazy deep-search wrapper. audit/tool-timing.ts:6 eagerly imports DB; analytics/src/index.ts:622–646 imports/writes DB and swallows failure. Route ALL nested/read/rate-limit/recovery/ordinary/new settlement emissions through these seams; preserve direct default timing/events. Keep changes in already-budgeted netdocs/investigator files, not an extra timing file/platform. Add investigation-only fresh ACL reads bypassing document caches and requiring requested-ID match; direct cache/defaults unchanged. Format gate is passed; no Excel parser/download seam.

In packages/connector-netdocs-mcp/src/client.ts, require every outbound search argument to be explicitly declared and compatible, including bound-page calls; never strip undeclared scope. Cover omission, additionalProperties true/absent/false, incompatible declaration, preserved plain search and both scoped positive controls in packages/connector-netdocs-mcp/test/client.test.ts. This addresses surviving historical finding 4; coordinate its separately tracked ownership with Main, not a second lane.

In already-budgeted connector types.ts/client.ts, add optional signal and investigation-only atomic pre-fetch reservation hook; propagate min(local, outer deadline) through typed reads, preparation/recovery, limiter waits/retries, window/fallback and body I/O. Compose SDK/per-request/outer cancellation and recheck before dispatch; preserve direct defaults. In runtime model.ts add optional constructor retry cap (zero here); enforce COMPLETE serialized 48 KiB input/memory cap before invoke, not financial token proof. model-pricing.ts only adds verified truthful Sonnet 1M/128K metadata; retain resolveModelBudgetCapability's interface and full $3.03 reservation with constructor 2k output. DROP 200k proof/optional pricing-input-limit API per Main arbitration accepted independently by both R2 reviewers. Registry approves utility/client/tools and constructor output cap; enforce exact resolved policy/pricing, caching AND thinking off. model.test.ts covers complete request framing/tool schemas/Unicode byte boundary, oversize refusal, exact model, output/retry caps and cancellation; connector tests cover actual-fetch accounting and no post-deadline dispatch. Estimate: 350 production LOC, subject to stop rule.

2. Implement bounded authorized state — Astra medium

Add packages/agent-runtime/src/tools/netdocs-investigation-state.ts and netdocs-investigation-state.test.ts. Inject clock/random/Redis; use existing REDIS_URL server-only pattern with short timeouts/offline queue disabled. Implement TTL/caps, hashed handles, schema version, binding and claim/revision fencing. SERIAL model admission atomically checks settled actual + unresolved reservations + new full reservation against each $5 turn/$15 lifetime/$50 org UTC-day and applicable $50 TOTAL eval cap. Durable unique-attempt fenced idempotent settlement replaces reservation with VALIDATED actual once; missing/invalid usage, failure/cancel/crash retains full. Schema-invalid output can have valid billable usage; account independently. Preserve original UTC-day bucket on late settlement; usage above bound stops/records violation and actual cost, never clamps. Test duplicate/stale/late settlement, races, all failure paths, five $0.10 actuals + sixth $3.03 = $3.53 admission and unknown $3.03 leaving at most $1.97/refusing next call. Lua operations fail closed; no DB/queues.

Test cross-org/caller/credential-owner/connector/scope/mode/endpoint/lease, normal mint stability despite custody updated_at, rotation/disconnect/reconnect, expiry, concurrency, replay/guidance, lost response/crash takeover, stale writers and all limits. Gate EVERY model request/receipt, including FIRST-TURN cached fetch/find, plus derived-frontier execution. Track explicit document dependencies for all retained aliases, queries/frontier, scopes, coverage, guidance, pending questions, findings and receipts; prune transitively, drop uncertain derivations, rebuild context from authorized structured state with no freeform assistant/tool transcript replay. Fresh cache-bypassing ACL reads require requested-ID match. Keep unvalidated candidate IDs opaque/internal, omit titles/snippets/attributes; no 200-fetch requirement. Fresh authorized search results are leads, not budget proof; cached leads need fresh validation. Test direct cache populated → ACL revoked → first deep call and same-lease alias revocation blocking pending query, coverage and old assistant leakage. Keep 8 docs, 64/192 actual HTTP, fixed 30-minute TTL and replay within original HTTP/105s cumulative active allowance; exhausted validation yields authorized partials. Real Lua proof requires explicitly isolated disposable Redis; earlier authorized Lua/PG results are recorded above, not new execution authority. Historical estimate: 400 production LOC, not the current measured allocation.

3. Implement the investigator and grounded receipt — Astra medium

Add packages/agent-runtime/src/tools/netdocs-investigation.ts, netdocs-investigation-prompt.ts and netdocs-investigation.test.ts. One private bounded StateGraph uses createModel, bindTools, injected timing (lazy production invokeWithToolTiming), event sink and real typed reads; schemas live at caller/model/state trust seams. No registry-wide graph exposure or arbitrary production tools. Enforce strict start/resume union, trusted prompt, five-read allowlist, small frontier, serial execution, citation/quote validation, honest completion and deterministic authorized partial/error fallback.

Use injected model transcripts and mock provider transport through the actual typed-read seam, not mocked final findings. Assert real executed read arguments, alias evidence chain, different-project rejection, financing-option distinction, unanswered approval/version question, budget exhaustion versus completion, discovery failure with successful plain search, stale cursor restart/dedupe, false totals, empty continuation pages, long-read limits and attempted instruction injection. Preserve cached retrieval timestamps; do not claim fresh/approved merely because a call succeeded. Estimate: 550 production LOC.

4. Register both adapters, gates and accounting — Astra medium

Update packages/agent-runtime/src/tools/netdocs.ts, packages/agent-runtime/src/index.ts, packages/connector-netdocs-mcp/src/types.ts, packages/agent-runtime/src/tools/netdocs-capability.test.ts. Expose the wrapper through netDocsTools and the existing dynamic product capability list, filtered by the new default-off flag; direct reads stay available. Names remain first-party only; never resolve deep_search against the provider catalogue.

Add apps/mcp-server/src/tools/netdocs-deep-search.ts; update apps/mcp-server/src/tools.ts and apps/mcp-server/test/netdocs-tools.test.ts. Reuse invokeNetDocsTool/invokeWithinMcpToolBudget: nest 35s active + 5s settlement within the outer signal/deadline and pass both through step 1's typed/provider seam, not adapter-only abort. Annotate read-only/open-world, redact to counts/booleans, and use enabled plus call-time flag guards. Extend catalogue/audit allowlists through derived lists.

Update packages/analytics/src/index.ts with typed investigation settlement (surface, stop reason, model ID, counts, duration, reserved/actual micro-USD, continuation/replay booleans). Invoke through the injected event sink, including nested/read/recovery/ordinary settlements, not only the new event; production defaults lazy, direct timing/events preserved. Publish existing diagnostic model lifecycle channels for nested usage, one call ID per reservation, no double counting. Surface tests cover real names, disconnected/pilot/flag-off hiding, reconnect/call-time revocation, ordinary direct availability, deadline and redaction. Estimate: 200 production LOC.

5. Prove behavior, independent reviews and controlled rollout — writer/review/fixer Astra medium; gates Astra low

Add packages/agent-runtime/scripts/eval-netdocs-investigation.ts: import investigator/typed-read/model leaves, not runtime index; inject synthetic access, lease/recovery, actual provider fetch transport, state/clock and in-memory TIMING/EVENT SINKS. Lazy defaults prevent real DB/custody/timing/analytics initialization on import AND execution. Exercise real typed reads, connector parsing, recovery, windows/cache-bypass and start/resume with configured production model; no fake final findings. Export inert fully synthetic fixtures, guard CLI with import.meta.main; no test-suite import or arbitrary production tools. Existing investigation tests use absent DB configuration plus DB-import/init attempt sentinel and throwing custody/network sentinels: assert ZERO DB attempts even when analytics would swallow an error, on success/rate-limit/recovery/error. Public wrapper remains in index.ts. Append verified additive contract to docs/connections/netdocuments-mcp-contract.md. The provisional 20-file fit is superseded by the current 33-file measurement; stop for Main on further scope expansion, including timing-file edits.

The current paid synthetic-model eval ceiling is $50 TOTAL across ALL cases/repetitions, including historical actuals, unknowns and reservations; current spend and STOP are recorded above. One persistent aggregate ledger, never reset; SERIAL full $3.03 pre-call reservations settle to validated actual independently of output schema, unknowns retained. Report unrun matrix whenever next reservation will not fit. Target unchanged natural-budget request three times with shuffled results/held-out alias plus negative/no-budget, disconnected and injection scenarios; no fabricated pass to complete matrix. Report latency, model/research/revalidation/actual-HTTP counts, cost, evidence, stop reason and unsupported claims. Core case must answer from both fully synthetic budgets shaped by verified format on first call. Measure cold catalogue, retries and every exposure/transitive-pruning check within 35s/64 HTTP; if infeasible stop for reviewed contract/cap revision. Six model calls is a ceiling; unknown $3.03 prevents another default call that turn, yielding deterministic authorized partial.

Historical September 8 format probe and permission record

The following $0/$5 spend, assignment and reviewer-in-progress statements are preserved at the time of the probe. September 10 changes only the aggregate EVAL ceiling and subsequent evidence, not the spent live exception; no live calls are allowed.

Text-read feasibility PASSED by Main at 3aa016ea via exact production netDocsSearchTool/netDocsFetchTool: first run 3 reads, 8.09s, exit 0; two shape-only runs each 2 reads passed (5.14s, 6.25s), total 7 successful top-level source calls. First run A 60,000 chars truncated, B 44,878 complete; budget/numeric/financing markers readable. A's first 60,000 chars are FLATTENED single-line text: zero physical newlines/tabs, zero literal backslash-n/r/t, zero pipe Markdown rows, no HTML table/JSON-like start; headings not standalone, no Sheet/Worksheet prefix. Synthetic extracts MUST exercise flattened inline headings/numeric groups and ambiguity; no invented preserved newlines/tab structures/cell coordinates or verified numeric association. Exact numeric accuracy/column association/full coverage/new-loop behavior remain acceptance; ambiguous associations stay unanswered, not automatic parser need. Own actor via authorized firm sharing, different credential owner legitimate. Earlier 3 diagnostics stopped before reads due overstrict owner==actor probe guard, not app/provider failure. Only normal lease/rate/analytics writes authorized, no document mutation/upload/manual DB writes/DB tests/alternate actor. Evidence counts/booleans only, no real values/IDs/titles/text stored. $0/$5 model spend. STOP live checking; final metadata received, no waiting/additional probe. Separate host choice/pilot remain unproven.

Binding permission record — LIVE EXCEPTION SPENT. Dispatcher independently confirmed with kwiss, relayed in Main w7E:p3F in the current September 8 session after final shape metadata: “That word does not generalise. Any further live call under his actor, any write beyond lease/rate/analytics rows, or any spend past the $5 cap needs a fresh word through me — not a re-reading of the same approval.” Original “Allow normal reads” / “Allow up to $5” were UI selections in Main w7E:p3F shortly before the 21:23 UTC plan-fix launch, not typed pane messages or a precisely timed free-text approval. Relay authorization with exact words, pane and approximate time. Seven cumulative source calls, none further planned, $0/$5 runtime eval spend. No further live calls/further writes or over-cap spend without fresh word THROUGH DISPATCHER; future smoke/pilot cannot reuse prior approval. Docs-only until reviews converge plus Main assignment. Flattened text does not prove numeric accuracy/column association. Permission-only record, no technical scope/cap change; fresh reviewers w7E:p30 and w7E:p41 are reviewing the technical plan.

Remaining gate and rollout requirements

Required gates: package test scripts from agent-runtime, connector-netdocs-mcp, analytics and mcp-server; root bun run lint, bun run typecheck, bun run eval:gate:cheap. Reported results above are scoped and dated, not a claim that all gates are complete. Include unchanged NetDocs handler-enforcement/absence/access/window suites and browser-import checks through the normal build gate if model-export changes require it. No UI/browser E2E changes: there is no UI change; product toolbelt plus MCP protocol tests and the real model-loop scenario cover the entry surfaces. Only the owner's isolated fixture grant covers the reported infrastructure proof; no live DB target, schema or migrations.

Two fresh Astra-medium reviewers independently audit implementation against this design, including untracked files and deletions; verify/fix findings and re-review to convergence. The historical 3aa016ea comparison is not current-main coverage: current base is 6bd0ae57, and Main must integrate origin/main before PR. The two scoped clean delta reviews above do not establish full acceptance. Gates seat runs commands and records exact output. Main owns status advancement and git; no worker commits. Flag remains off until acceptance and authorized pilot smoke pass. Rollback hides only investigation on both surfaces, rejects new/resume calls, keeps direct tools and lets state expire. No merge, deploy or enablement is authorized by this plan.

Historical scope ledger; retained stop and teardown contracts

The 20/22-file choices, 1,800-LOC ceiling, source-only assessment and worker assignment below are September 8 history, not current scope or execution status. Preserve their authority quotes and placement rationale; reported actuals remain 33 files / 5,230 gross production additions under the newly approved 34 / 5,600 ceiling above. The local-teardown, mutation-proof and safety requirements remain binding.

Exactly 20 anticipated implementation files: runtime netdocs.ts, netdocs-investigation.ts, netdocs-investigation-prompt.ts, netdocs-investigation-state.ts, netdocs-investigation.test.ts, netdocs-investigation-state.test.ts, netdocs-capability.test.ts, index.ts, model.ts, model.test.ts, model-pricing.ts; connector client.ts, types.ts, test/client.test.ts; analytics src/index.ts; MCP tools/netdocs-deep-search.ts, tools.ts, test/netdocs-tools.test.ts; runtime scripts/eval-netdocs-investigation.ts; canonical contract Markdown. Paths are expanded in steps above. Estimated 1,500 production LOC + bounded eval harness, with 300 LOC contingency inside the 1,800 cap; tests separate. Typed read extraction is the largest uncertainty.

Conditional ceiling, not chosen expansion: Main relays that dispatcher obtained kwiss's exact choice “Approve 22 files” directly via AskUserQuestion in the current session, then corrected its interpretation. 22 files and the EXISTING 300-LOC contingency are conditional ceilings only; 1,800 total production LOC remains unchanged. FIRST assess confined investigation-private, connection-owning authorization in existing runtime netdocs.ts using existing pg/@types/pg/drizzle dependencies. If awaited local teardown and no subsequent dispatch are achievable, choose the original 20 files. DB-package placement was preferred, NOT technically required by dependencies. If confinement fails the contract, report the concrete failure before using packages/db/src/rls.ts plus a DB test addition; Main arbitrates placement. No placement or teardown proof is complete yet.

Required code comment and PR-body checklist: explicitly promise awaited LOCAL teardown/no later dispatch, never instant remote rollback or custody-server cancellation. Record the chosen placement reason, shared-pool queued-acquisition race, and rejection of the pool-counter workaround because it refuses warm idle pools and defeats reliability. Report ceilings versus actual files/LOC. Required biting mutation proof: disable/remove teardown and the test must fail naming the detached query. These remain acceptance requirements; the earlier Lua/PG results alone do not assert this specific mutation proof. No schema/migrations/grants/dependencies: stop BLOCKED on drift. No live calls; DB verification only within explicit isolated-target authority, not granted to this writer. Stop on unmet confinement/expansion criteria or gross production additions beyond 5,500; never omit authorization, replay protection, first-call investigation or workbook acceptance to fit. No Excel parser/download seam is needed. This docs-only assignment owns only the two HTML pages, not code, STATE, scorecard, generated indexes or validation execution.

Final Main source-assessment selection (supersedes pending placement above): confined resolver in existing packages/agent-runtime/src/tools/netdocs.ts plus already-budgeted netdocs-investigation.test.ts; original 20-file target, conditional 22 fallback unspent. Main reports public @workspace/db pool/schema and runtime pg/@types/pg/drizzle dependencies verified by source assessment, not DB execution. Scout estimates +175–250 LOC within existing 300 contingency, not measured implementation growth; 1,800 total unchanged. Chosen confinement avoids extra DB-package placement and shared-pool queued acquisition. Own a standalone client/socket, use a fail-fast slot cap and strictly await actual connect/query/local close; no shared-pool checkout. While connect is pending, destroy the socket FIRST and await connect settlement/closure; never call client.end first, since pg can suppress the pending connect callback. Required code comment/PR body disclaims remote cancellation and exact real-time guarantees under OS/event-loop stalls. Source-feasible, not behaviorally proven; implementation and biting teardown mutation proof still required. Next worker netdocs-owned-auth w7E:p44 Astra medium is prepared but NOT yet assigned implementation. Current writer remains sole writer until this docs report; no product edits or further source exploration now.

Historical schema-admission-only verification and handoff

The following is the original two-file implementation snapshot, including the failed 222/224 run and corrected 224/224 run. Its “actual implementation is only,” “not implemented,” $0/$5 and status-command statements are historical, not today's implementation, spend or writer permissions. Current verification and STOP are at the top; no acceptance criterion is discharged by relabeling this history.

Main owns behavioral verification. Actual implementation is only connector client.ts/test/client.test.ts schema admission: 28 production additions, 4 removals, 0 moved, two implementation files. Every outbound search argument must be explicitly declared and compatible on normal/bound-page calls; no scope stripping. Main ran bunx vitest run test/client.test.ts from packages/connector-netdocs-mcp without dotenv/live access: first 222/224 passed; two supported attribute_filters positive fixtures lacked its declaration. Test-local declarations corrected them with all original assertions and fail-closed production logic unchanged. Main's rerun passed 224/224 in 8.10s. Full investigator, remaining typed/model seams, cancellation, state, prompts and registration are not implemented; isolation sentinels, first-turn/transitive ACL, durable actual-cost settlement, request cap and measured cold positive-case fit remain acceptance work. No implementation-review or full-feature green claim. Current docs-only update sets canonical metadata in-progress through the permitted spec-status command because scoped product implementation now exists; it runs no product validation or other scripts. Historical spent-live/UI-selection provenance below/above remains unchanged: 7 source calls, $0/$5 eval, flattened readable text is not numeric-association proof.