All specs and plansNetDocuments

NetDocuments investigation — design spec

2026-09-17 addendum — evidentiary grounding and v5 continuation (implemented in the lane worktree)

Opened in progress and implemented on 2026-09-17; final verification and paid synthetic adjudication are recorded below, with historical metadata and enablement unchanged. This addendum supersedes historical v4 receipt wording, not authorization, concurrency, accounting, retention or request limits. Exact quotation is not entailment: definitions, obligations and conditional scenarios do not establish approval, pending/non-approval or adoption. Explicit dated acts/status records support their corresponding scoped conclusions; a factual budget answer need not exhaust the corpus.

Final verification: runtime 5,194 Vitest + 139 and 77 Bun tests passed, 171 skipped; five scoped suites 369 passed with one opt-in PostgreSQL proof unrun; connector limiter eight passed; eval 72 passed; 14 temporary mutants failed as intended and were restored. bun run lint && bun run typecheck && bun run eval:gate:cheap exited 0. Cheap router gated cases passed 4/4; four nongated boundary cases scored 0/5 versus an earlier 5/5: variability, not a claimed regression fix or all-router-case pass. Memory passed 9/9 with a 48-active versus 20-cap warning. Final prompt pressure proof: system 4,974 bytes, rejected preflight 120,241, largest accepted preflight 101,952, largest actual request 85,689; limit remains 120,000. See contract §10.

Paid rerun c74b2c690f0018c353de174b53555e6e12100189c3af6359fb27b9a6477b2e72: four fixtures, six reports, six answers and 21 findings; independent Astra LOW and Fable MEDIUM semantic PASS, with no report-truncated evidence. Corrective generation 1 produced new current synthesis with one new research read and two model calls. Replay used zero models, zero research reads and $0 model cost, while accounting for four revalidations.

The first paid run failed whole-report acceptance only on unread-content characterization; a generic three-line prompt correction fixed that discipline, and the latest run contains no such extrapolation. An immediately self-corrected lead ID is a minor presentation defect, not a materially false final claim. The factual zero-hit find result remains independently unverified because query/hits are not exposed; only the operation's occurrence is corroborated. Exact source-length assertions are not uniformly independently verified. Semantic PASS is not blanket metadata verification; final gates passed with the nongated diagnostics noted above, and no CI/PR, merge, deploy or enablement is claimed.

2026-09-14 top addendum — direct-first research with a shared server-side research record (implemented in the lane worktree 2026-09-14; default-off)

Status and authority. Approved design, amended 2026-09-14 after two independent reviews (omp Astra HIGH and a fresh Fable HIGH in Claude Code); the design and amendments below are implemented. kwiss approved implementation and the fixed 30-minute retention; the two former open product questions are closed. Implementation status, 2026-09-14: implemented by the Fable HIGH worker (normal Claude Code subscription) in the lane worktree, uncommitted, default-off behind APP_NETDOCS_INVESTIGATION_ENABLED. Verification: independent focused suites 216 passed, 2 skipped, 0 failed (research 45 passed; investigation 124 passed, 1 skipped; state 47 passed, 1 skipped); full runtime 5554 passed; cheap eval coverage 11/11, retrieval 27/28, router 4/4, memory 9/9; actual MCP smoke 2 direct reads, save/get, Opus 5 low escalation with 3 findings, 3 citations, 2 documents in 109.7s; exact-quote save/get 4 calls, 2.8s on a 22911-character window; web production build passed. Test seat deviation recorded: independent non-Claude tests are written by the omp @write Astra HIGH seat at its unchanged configured model and effort, because no dedicated test-writer seat exists. kwiss approved the direction (the host's own model researches with the ordinary NetDocuments tools, saves and retrieves research state, and escalates to deep search when direct work is insufficient) and asked for a concrete design before any worker starts. Agreed scope: direct-first research, one shared server-side research record, escalation to north__netdocs_deep_search defaulting to Opus 5 low on the approved Bedrock ZDR route. Out of scope: new product features or UI, ingestion, a generic platform, third-party services, cloud writes, wider model selection. Motivating evidence: the same private budget question passed once through actual MCP with Opus 5 low (both expected workbooks, 5 validated citations, 81 s, $0.926140) and once through a fresh direct Fable low host (both workbooks in 67 s, 7 of 8 quotes exact). That is one budget-discovery case; nothing here claims full-workbook, latest or approval coverage, or release readiness. Seat deviation, recorded with its reason: this design was authored by Devin Fusion (Claude Fable 5.1 high with an SWE-2 medium sidekick) because kwiss explicitly asked for it, in place of the @plan Astra seat; no product code was written. The plan review ran as Astra HIGH rather than the seat table's effort because custody and paid lineage span two stores (the research record and the investigation state); the Fable HIGH review is the normal Claude Code subscription.

Source of truth. The 2026-09-14 lane tree and scorecard govern; the 2026-09-11 traversal, limit tables and frozen-exposure paragraphs below are history. Preserved from the current code, unchanged: access checks on every new typed read, one deduplicated revalidation of carried history on an actual resume, fresh exact cited-window re-reads at publication and replay, the immutable transcript, three money-turn slots, $5/$15/$50 money guards, charge ledgers without TTL, the 524,288-byte context, fixed 30-minute expiry, three live investigations per actor. No per-model traversal returns; no money key or unknown liability is reset when a content format changes.

Review amendments, 2026-09-14 — implemented corrections that replace any contradictory draft statement below

  1. Decisions closed. Retention is the fixed 30-minute TTL. The existing model selector stays; a record-started escalation only defaults to Opus 5 low. Bounds, money guards and unknown-money liabilities are unchanged.
  2. Foreign handles never mutate. A foreign, invalid or mismatched-binding handle fails without touching any owner record or slot set; nothing is deleted on a foreign-handle lookup. Only owner-authorized, server-observed source revocation can invalidate owner state.
  3. Retained observations are revalidated on actual retrieval. Before get, the escalation import or a later resume exposes retained source metadata, quotes or notes, the server runs one bounded, deduplicated fresh validation of the carried source observations, the same policy as the investigator's resume: revoked or materially changed observations withhold the old content; an outage or deadline withholds it and keeps the private state. This runs on retrieval and resume only, never around each model thought. Imported direct observations live in an explicit, non-citable researchImport field of the investigation context, never as fabricated native tool messages or sources[] entries, so later resumes cannot bypass it; stripping title and url is insufficient because IDs, notes and citations can also disclose retained material. Revalidation of a direct observation re-reads its ORIGINAL exposed range and fingerprint (up to 60,000 characters); native observations keep their 3,500-code-point window semantics; a direct fingerprint is never truncated to the first 3,500 code points to fit native validation.
  4. Actionable typed leads are kept. The record saves the server-observed typed search result IDs, the response next_page_token and reported counts, find-hit offsets and the source args and fingerprints, marked historical; no rendered-MCP parsing and no reliance on host notes to reconstruct leads. Aggregate bytes are bounded and an omission is an explicit record_full / omitted status, never silent loss. Every direct operation that influenced the research is captured when a handle is supplied, including browse and list_attributes schema provenance; their ordinary no-handle behaviour is unchanged.
  5. Paid handoff is idempotent. Before any investigation is created, one server-generated stable investigation nonce plus the pinned selection and request are persisted atomically in the record. state.start gains a trusted create-or-recover form keyed by that nonce that checks identity, request and selection, atomically seeds the immutable researchImport descriptor and snapshot into the NEW investigation context on creation, never overwrites an existing context, and never overwrites money or unknown liabilities. A retry after pending-before-create, create-before-link or a lost response recovers the same session, and a lost create response before the first claim recovers generation zero with its import, never an empty ordinary-resume context; the link is persisted before any model execution; concurrent calls cannot create separate lineages. Eviction protection begins atomically with that persistence: a record with a pending creation or linking, or a persisted link, is never evicted, and a concurrent fourth open is refused when no unprotected record exists. A context seeded with researchImport expires no later than the record it came from (fixed 30 minutes from the record's open); the trusted create form pins that bound with the nonce, request and selection, recovery never extends it, ordinary starts keep their own 30 minutes, and the money and late-settlement records keep their ordinary TTL. Ordinary continuation semantics are unchanged.
  6. Stable linkage and read-only inspection. The record keeps its link to the investigation. get resolves the latest persisted active, continuation or completed generation, including after a lost finish reply, and returns current checked progress and the handle without a model call and without touching budget or admission counters; a duplicate escalation uses that latest state, never a generation-zero replay. During an active handoff, direct mutations of the imported snapshot are refused so the model never resumes a moving import; no-handle direct reads stay available. The getter returns safe structured progress, never private model reasoning.
  7. Confidentiality stated as it is. No full server-read bodies or search snippets are stored; bounded citation quotes and the untrusted host notes and pending text are stored as plain private Redis JSON. Sentinel tests target source observations, not all host text. Retrieval checks run under bounded deadlines: an unfinished check says unavailable, preserves state and never implies verified. Eight windows per save is a maximum, not a guarantee that all eight finish within the 45 s tool budget; there is no two-window cap.

Why a separate typed read journal, not the investigation transcript

The investigation context (investigationContextSchema) is model-coupled: every sources[] entry points at a stored LangChain tool reply (message index, the "Dangling source" refinement), citation resolution reads the quote out of that stored body, a claim takes the per-actor executing lock and charges active time against a 600,000 ms lifetime, and the envelope allows three generations. A direct-first host has no server-side model turn: its reads are many, often parallel, spread over minutes, and return up to 60,000 characters that the transcript cannot hold (one such window at four bytes per code point exceeds the 120,000-byte wire cap by itself). Forcing direct reads into that transcript would mean synthesising tool turns, truncating bodies the host actually saw (so an honest host quote fails as "unobserved"), and spending the investigation's time and slot budgets on non-model work. A separate research record is therefore genuinely simpler: a typed, append-only provenance journal with no model transcript, no money and no generations. It reuses the source schema fields, investigationFingerprint, the HMAC handle scheme, the identity binding, the production Redis client and the typed read client; it does not reuse the envelope, the claim/lease protocol or the Lua money block.

The record

Operations: one MCP tool, north__netdocs_research, and one optional field on three direct reads

  1. open {request}{research_handle, expires_at, procedure}. Authorizes (binding), creates the record, returns the research procedure text below. No provider read.
  2. Recorded direct reads. north__netdocs_search, north__netdocs_fetch and north__netdocs_find_in_document accept an optional research_handle. With it, the adapter runs the read through createNetDocsReadClient (netdocs.ts:164), which invokes the same inner tool under readScope so the typed result is taken from the existing captureRead hook (netdocs.ts:128), never by parsing the rendered MCP text; a new option authorization: "pooled" keeps the ordinary RLS pool path rather than the raw-socket investigation path and its four slots, and the read() result gains the rendered host text next to the typed data. Consequences, stated plainly: a recorded read is always a fresh provider read (inside readScope the 10-minute document session cache is bypassed), the served-ID mismatch refusal applies, the binding is checked against the record on every append, and the reply's first server-authored line is research: recorded or research: not_recorded reason=<record_full|omitted|expired|binding_changed|not_enabled>; the read result itself is returned either way. north__netdocs_browse and north__netdocs_list_attributes also accept the handle and are recorded as schema provenance when it is supplied; without it they behave exactly as today. Without the field every direct read behaves exactly as today; the inner LangChain schemas and the product toolbelt are unchanged, the field exists on the MCP definitions only.
  3. save {research_handle, revision, notes?, pending?, citations?}{revision, citations[]}. Revision CAS (stale on mismatch; the host re-gets). Each citation {document, quote ≤ 300} is verified only against a recorded fetch source of that document: the server re-runs that source's typed fetch fresh (same offset and served length, up to 60,000 characters), requires the quote as an exact substring of the fresh text, and stores {document, quote, offset | null, window: {offset, served}, status, verified_at}, where offset is the absolute code-point position when the provider confirmed the window offset. Statuses: verified; unobserved (in no recorded window, or not exact); changed (recorded fingerprint differs); revoked (access withdrawn: the record is deleted and the call fails closed); unavailable (outage or deadline: record kept, nothing claimed). At most 12 citations and 8 distinct windows per save; more is refused before any read. Text the host pasted into notes is never checked as a citation.
  4. get {research_handle} → the record, after authorize() and a binding match (mismatch or another actor: restart, nothing disclosed, nothing mutated). Before any retained source metadata, quote or note is exposed, the carried source observations are revalidated once, fresh and deduplicated, exactly as the investigator's resume does: revoked or materially changed observations are withheld, an outage or deadline withholds them as unavailable and keeps the state. Sources remain provenance, and the host must fetch to quote. When the record links an investigation, get also resolves its latest persisted generation (active, continuation or completed, including after a lost finish reply) and returns checked progress and the current handle without a model call or any counter change. Citations are re-verified fresh, a located one through a 3,500-code-point window at its offset and an unlocated one through its originating window, deduplicated by document and window; a deadline mid-way leaves the rest unavailable. This is the one deduplicated revalidation on an actual resume in the direct flow. Also returned: a coverage summary (searches with their observed next_page and reported marked historical, documents with windows read), notes, pending, expires_at, and the procedure text.
  5. Escalation. north__netdocs_deep_search {research_handle, model?, effort?} is a third start form, never mixed with request or continuation_handle. Order: load and authorize the record as in get; atomically persist a server-generated stable investigation nonce with the pinned selection and request into escalation before anything is created; state.start in its trusted create-or-recover form keyed by that nonce (identity, request and selection checked; the immutable researchImport descriptor and snapshot seeded atomically into the new context on creation; an existing context, money and unknown liabilities never overwritten), default pair opus/low when omitted (explicit supported pairs unchanged); persist the link before any model execution. A retry after pending-before-create, create-before-link or a lost response recovers the same session, and concurrent calls cannot create separate lineages. A later call on a linked record continues from the latest persisted generation, never a generation-zero replay: no new slot, no new reservation. While the handoff is active, appends and saves against the imported snapshot are refused; no-handle direct reads stay available. The first DATA message carries the import block: recorded sources as provenance, the fresh citation statuses from the import-time recheck (the same bounded recheck as get, charged as refresh attempts), notes and pending as UNTRUSTED caller text, and the statement that nothing imported is evidence and that located citations name the netdocs_fetch offset to read. Nothing from the record enters the transcript as a citable body, so the investigator's existing exact-quote, revocation and change guards hold unchanged; a revoked cited document at import stops authorization/research_cited_revoked with its content withheld; the record and its link are retained. The imported provenance lives in the context's non-citable researchImport field, so later resumes of the escalated investigation revalidate it once, at their original exposed range and fingerprint, and cannot bypass it.

Byte and trust mismatch, resolved

Direct fetch serves up to 60,000 characters, possibly from the in-process session cache; the investigator retains 3,500-code-point provider windows. The record takes neither body: it stores served length, offset, a fingerprint and, for citations, the bounded quote the host asked to verify, and every verification is a fresh provider read. A host may therefore quote from a 60,000-character window and still obtain verified, because the check re-reads that window rather than a truncated copy; the escalated investigator never sees direct text at all and re-reads 3,500-code-point windows at the located offsets. Recorded reads bypass the session cache by construction, so source: "session" never appears in a record.

Host procedure (returned by open and get; three lines in the server instructions point to it)

Open a record before the first read and pass its handle on every search, fetch and find; only recorded reads exist for the record. Search for an anchor, fetch promising IDs, treat identity as a hypothesis until a fetched passage confirms it, cite only from fetched text, save notes, pending and citations after each milestone and before a turn ends, and report a quote as verified only with the server's verified status from the latest save or get. Objective escalation signals, readable from receipts: six or more recorded searches with no fetched window containing the sought material; a needed document whose totalCharacters exceeds 60,000 and whose find_in_document returned no hit; more than four pending leads at turn end; identity unresolved across two or more candidate matters after fetching one document each. When one holds, call north__netdocs_deep_search {research_handle} and follow its continuation handles. These are instructions and signals, not enforcement: North cannot make a host obey them and does not claim to detect an untruthful host.

Rejected additions

Host-supplied provenance or snippets as sources; window text in the record; writing research into north__memory_save (generic durable memory); implicit association through the MCP session ID (dies with the host process, merges topics); a per-thought traversal of the record; sliding or longer retention; a separate verify tool (folded into save); seed tool turns synthesised into the investigator transcript (leads are cheaper and keep one evidence rule); a research skill in the legal skills catalogue; any change to model selection, money caps, HTTP caps or the three-slot investigation cap. Releasing an actor slot when an investigation finishes (a finished generation holds its slot for 30 minutes, which caused start_busy during the benchmark) is not needed for this scope and is deferred explicitly: the concrete fix is a ZREM of the session from the three slot sets on the terminal finish with money keys untouched, to be designed on its own.

As built, 2026-09-14 (names that differ from the draft prose above)

Leaves: netdocs-read-provenance.ts (provenanceOf over raw typed data, full served range, plus the shared canonicalArgs/hitMaterial helpers the investigator now imports), netdocs-research-record.ts (limits, schemas, researchImportOf, a Lua store with three ops: create with protected-aware eviction, get, and revision-CAS put; the record travels as one opaque JSON string, so Lua never re-encodes an array) and netdocs-research.ts (runNetDocsResearch(...).open/save/get/record). One LangChain tool netdocs_research (createNetDocsResearchTool in netdocs.ts, MCP-only) carries open|save|get for the host and the server-side record op that the five direct MCP reads reach when they carry research_handle; the reply's first line is research: recorded id=r<n> revision=<n> or research: omitted reason=<record_full|restart|escalated|stale|deadline>. Escalation and linking are two put writes on the record (nonce, model, effort first; the generation-0 handle second); the state module's start takes an origin (nonce, seeded researchImport, bounded expiry) and gains a read-only inspect op. Revalidation reserves the connector's own operation budget before each fresh read and reports the rest unavailable when the 45 s tool deadline cannot fit it. Withheld observations are persisted with only operation, args and fingerprint. No drop op, no analytics event, no product-toolbelt exposure.

Decisions taken by kwiss, 2026-09-14

Retention is the fixed 30 minutes from open, non-sliding. Record-started escalation keeps the existing model selector and only defaults to Opus 5 low. Neither question is open.

2026-09-14 top addendum — what the first actual run of this design showed (in progress, default-off)

Observed (pre-R2 server, Opus 5 low, approved Bedrock pin): the cumulative context works as designed (26 entries, two generations, 12 model calls, no snapshots), but the run produced no answer: 11 useful reads (8 search, 2 find-in-document, 1 fetch) against 154 refresh reads and 24 denials; one partially fetched ~39k-character source; neither budget workbook reached; a negative/limited-coverage finalization refused for a 1,035-character coverage_statement against the schema-disclosed 800 maximum; the generation-1 receipt never reached the client within 360 s for a cause not observed. Ledger $0.850860 actual, $0 unknown. Not a semantic pass; not a causal comparison with the Fable CLI experiment (different model and route).

Policy question raised, not decided: the strict pre/post-model full-history revalidation costs about fourteen refreshes per useful read at this corpus size. A candidate relaxation (fresh access on new reads, on resume and at citation publication only, same actor-bound history) would leave a mid-generation source edit or revocation in already-read context until the next model call. The strict policy stays in force until kwiss decides. R2 and R3 repairs (coverage, double-resume, text-only ACL) are verified by suites and converged reviews; the final tree has not rerun the question. Feature stays default-off; PR #423 draft with a known conflict.

2026-09-14 top addendum — native research agent with cumulative context (in progress, default-off)

Approved direction. kwiss approved a fast-track cutover from the deterministic snapshot loop to a real internal research agent: the model holds one cumulative assistant/tool conversation and chooses every typed read itself; deterministic code executes each call under the existing guards, records it, and never rewrites an observed turn. Basis: a 67-second fresh Fable CLI experiment (20 read-only NetDocs calls, 4 cited documents, 7/8 exact quotes). This supersedes the 2026-09-11 "model chooses direction; code executes the frontier" ownership below, and the earlier rule that forbade transcript replay: the replacement guarantee is per-source fresh re-verification before any model replay. Implemented by a Fable LOW worker (recorded deviation from Fable high); reviews and acceptance are still open.

Revised policy

Behavior proofs added

Consumer-contract tests cover: answer over a 5-turn cumulative transcript with the whole history on the wire; a refused fifth call; malformed arguments, unknown tools, unobserved quotes and text-only turns corrected without spending reads; the 24-call ceiling and context pressure as terminal stops; an interrupted batch resumed with truthful replies and re-verified history; cancellation with no later dispatch and a settled liability; changed cited material and a revoked cited source at publication; unknown settlement after transport loss; token/total drift preserving evidence while changed text stops replay; finalized-receipt replay without model or research. The Lua contract (reference transport and actual Lua) proves quote/document/session forgeries are refused and the batch record is bounded to stored turns. The eval matrix runs the real investigator over the synthetic JSON-RPC corpus, including forced continuation. Real-provider acceptance is not claimed.

2026-09-11 top addendum — approved production integration of the deterministic loop

In-progress / default-off; approved for implementation after plan-review convergence, not implemented or semantically accepted. The user explicitly approved integration and continuation without interruptions unless post-implementation tests fail; no repeated user-approval gate. This addendum and the matching integration plan supersede prior unproven replacement acceptance, completion-driven architecture choices and stale limits wherever they conflict. Older sections remain historical evidence, including failures, not a green baseline. Read-only acceptance under the established lane identity and $50 lane allowance is already authorized. No merge, deploy or enablement is authorized.

Evidence boundary

The private reference source has SHA256 7948acc3ba06dfe0301041103fbd3932c55d88a4529aacb46e80203ac7090c40 (hash checked in this planning pass). Its clean unhinted end-to-end run reports 3 model calls, 39 research reads, 38 refresh reads, 20 pages, 257 unique candidates, 12 distinct documents, 396,866ms and 418,225 microUSD; zero failures or throttles, all accepted citations freshly validated. Six pending pages and 249 pending reads remained. This establishes one successful scenario, not exhaustive search, general reliability, a controlled speed comparison or success through production deep-search. Private cached continuation windows were withheld, never treated as fresh evidence. No confidential names, IDs, quotes, corpus or private script artifacts are copied into this addendum.

Ownership: model chooses research direction; code executes the frontier

Port the reference behavior into packages/agent-runtime/src/tools/netdocs-investigation.ts:createNetDocsInvestigator. The model proposes bounded searches, prioritizes exact returned document IDs, requests explicit revisits and interprets exposed evidence. Deterministic code discovers pages and fetches candidates. Replace the old per-read model tool dispatch and step/execute orchestration; do not retain it as a parallel fallback. Keep the existing typed five-operation read allowlist internal, with no recursive deep-search, REST, background platform, auth/connector-storage change, target-specific search hints or titles-to-ID aliases.

  1. Queue every candidate from a successful typed search page, deduplicated by provider document ID with memberships in its originating pages. Ranking changes order, never eligibility: zero-score candidates remain queued. Equivalent searches reuse their traversal rather than resetting it.
  2. Alternate page and document work when both are available. Page work rotates round-robin across active queries; explicit model priorities precede generic relevance within document work. Preserve the fairness position across yields. Empty or duplicate-only pages with advancing cursors are progress; repeated/cyclic cursors are incomplete, not completion.
  3. Keep each in-flight item pending until its result, candidate memberships, unprocessed page remainder, next cursor/window offset and counters are checkpointed together under the current fence. Persist unprocessed work before advancing a cursor. A storage refusal must leave the last durable page position recoverable, not silently skip its remaining candidates.
  4. Continue document windows only from provider-confirmed progress. No hidden per-query page cap or per-document window cap may masquerade as completion. Global budget/storage limits retain the recoverable frontier and report partial coverage. Missing offset, uncertain end or lost process-local cursor requires an explicit caveat/restart path with deduplication, never an invented span or uninterrupted-coverage claim.
  5. Revisits live-read known, currently authorized IDs before re-exposure, including material omitted by the bounded model projection. Model priority cannot authorize an unknown ID. Persisted IDs are server bookkeeping, not evidence or permission: refresh their originating dependencies before exposure or derived dispatch.
  6. Model-only exhaustion or admission failure disables further model calls, not already-authorized mechanical traversal. Drain pending reads under remaining read/HTTP/time/access/state and applicable financial guards. Decision five may be followed by more reads, never a sixth model call; preserve the frontier and report partial only when guards block that work.

State and finite retention

Extend netdocs-investigation-state.ts:investigationContextSchema, emptyInvestigationContext and invalidateInvestigationMaterial, not a new store. Store the deduplicated queue, query/page provenance, exact dependency fingerprints, pending cursors and page remainder, window continuation, fairness position, priorities/revisits, cooldown and cumulative progress. Dependency fingerprints cover exposed IDs, title, metadata, content and relevant pagination/completeness facts, not text alone. Keep missing/changed/revoked dependency state durable through interruption and invalidate dependent queries, candidates, observations and receipts transitively.

Port the proven fetch/exposure settings: 12,000-character fetched windows, 3,500-character exposed slices, at most 8 exposed windows and 12 exposed pages, with a 120,000-byte serialized wire guard. Current product retention is 6,000 code points per window: increase it to preserve the proven fetched window, not silently halve the evidence or expand exposure. Define production characters consistently as Unicode code points using the existing provider-facing offset/length and slicing semantics; bytes remain UTF-8 serialized bytes. The private loop uses JavaScript slices for some clipping, so name and justify that Unicode-safe normalization in review and verify non-BMP boundaries. Any other difference from the successful reference must likewise be named and justified in review.

Dynamically account for actual UTF-8 serialized queue/provenance, retained evidence, receipts, observations and replay bytes within existing 524,288-byte context and 2,097,152-byte envelope guards. Independent maxima need not fit simultaneously: 12 windows × 12,000 code points × 4 bytes = 576,000 before receipts. Reserve pending provenance and stop/continuation space before optional evidence. Preserve resumable unprocessed page remainder/cursor work under pressure; evict optional material or yield partial with honest omissions, never silently lose candidates or impose fixed smaller search quotas. If a pending dependency cannot remain available, park the work as unverifiable without losing its provenance.

Feasible example, not a new cap: 8 × 12,000 × 4 = 384,000 window bytes plus 100,000 measured bytes for all other serialized context fields/escaping = 484,000, leaving 40,288 for reserved headroom including stop space. Wire example: 6 × 3,500 × 4 = 84,000 plus 20,000 measured bytes for pages/framing/other fields = 104,000, leaving 16,000 below 120,000. Actual admission must measure serialization and reserve headroom; the example assumes those measured costs, not universal text sizes. At most 8 exposed windows/12 pages are upper bounds, not promised occupancy. Retained/exposed selections and envelope snapshots fit dynamic bytes, not the product of independent maxima.

Retain current identity/connection/scope/lease binding, HMAC handles, fixed 30-minute expiry, fenced claim/revision CAS, concurrency and replay-only claims. Preserve the existing THREE money-turn record slots, generation count and attempt.turn semantics, including old late settlement. A 600,000ms lifetime across these slots suffices for the planned two 300,000ms calls; do not add generations. Content-version incompatibility fails closed for content, but every cutover preserves money keys, historical liabilities and original-turn late settlement; no DB migration or ledger reset.

Budget contract: whole investigation, not fresh allowances on resume

Inspected sourceApproved integration
INVESTIGATION_LIMITS.activeMs = 300000, lifetimeActiveMs = 300000Keep 300,000 active ms per call; use approved 600,000 active ms lifetime because the proven run needed 396,866ms. Resume spends remaining lifetime. Bootstrap authorization, cooldown, reads, model work, refresh and replay consume active time; unknown crash-time reservations stay charged.
12 whole-investigation model calls, 16 research admissions; 204 refresh admissions per generationMirror the experiment's whole-investigation ceilings: 5 model decisions including correction/finalization, 80 useful research reads, 200 typed read attempts total including refresh and limiter denials. A correction or final refresh is not free. Retain an independent refresh safety guard; replace its obsolete 12×16 rationale.
440 HTTP/generation, 488 HTTP/lifetimeRetain independent actual-HTTP admission, including protocol, fallback and recovery traffic. Typed attempts are not HTTP counts. Either bound can stop work sooner; do not claim the experiment proves fit with production overhead.
15,000ms cleanup; deep MCP default activeMs + 2 * cleanupMs, maximum 360,000msPreserve existing MCP cleanup headroom and owned settlement: 330,000ms default envelope, not a 600,000ms RPC. Bound core by shorter outer deadlines; yield early enough for fresh publication/checkpoint while active time remains. Cleanup is drainage/durable finish only, never extra research.
$3.03 reservation; $5/generation, $15/lifetime, $50/organization-day; production eval comparison $5 vs separate $50 aggregate evaluation ledgerKeep production financial guards unchanged, including full reservations, actual/unknown settlement and historical aggregate spend. Main resolves the eval discrepancy from tool/repo-provided configuration as an engineering accounting check against the already-authorized $50 lane allowance, not a new product decision or implicit cap increase.

Admission work is required, not a constant-only edit. Current commandSchema admits research/refresh before reading; Lua increments research at dispatch, and has no useful-outcome/total-typed-attempt protocol. Review a fenced admission/outcome transition across createInvestigationState, investigationStateLua, runtime read and eval createEvaluationState. Reserve attempt capacity before every typed dispatch and conservatively reserve potential useful work; settle classification idempotently without refunding the attempt. Unknown outcomes remain charged. In the reference, “useful” means a non-rate-limited research attempt, including a non-throttle failure, not necessarily a successful document. Preserve that explicit meaning rather than label attempted reads as successes. Retry-after cooldown preserves the same pending item and consumes time/attempts; permission denial invalidates access and is not retried as a throttle. No reported retry-after or insufficient remaining time yields an honest partial stop.

Before each model call, reserve typed-attempt capacity for the exact frozen exposed post-decision refresh set, including dependency traversal, and wall time for bounded model plus refresh (reference loop.ts:326–329). Protect that capacity from unrelated dispatch. If admission cannot fit, disable further model calls and continue mechanical work instead; this is the reference contract, not another strategy. Apply the same persisted cooldown policy to EVERY typed read, including pre-exposure/post-decision/final/replay refresh, replacing immediate refresh-throttle budget stops. Preserve pending item and phase across resume; feasible retry-after waits remain bounded by remaining guards. Throttle never means permission withdrawal.

Live access, frozen exposure and evidence-bound publication

Reuse netdocs.ts:createNetDocsReadClient typed results, not the private rendered-MCP-text parsers. Source verification: execute runs under readScope.run; readDocumentSession returns undefined while that scope exists (lines 1903–1904), and fetchNetDocs refuses returned-ID mismatch (lines 138–150). Scoped reads propagate the active deadline, signal and HTTP admission hook. Thus investigation fetches bypass session material; ordinary direct reads keep their cache. authorize is custody/binding only, never a document permission check. Retain explicit source/offset confirmation in the typed path and reject session-served material defensively rather than infer freshness from a recent timestamp.

  1. Select a bounded prospective exposure. Traverse originating dependencies with a visited set and cycle/missing-dependency failure handling; live-refresh stale retained pages, discovery and document windows before exposing their data. Persist invalidation as it occurs so an interrupted refresh cannot leave a reusable stale snapshot.
  2. Freeze exactly the successfully validated exposure, including the exact displayed substrings and dependency fingerprints. Removed/unverifiable items do not get replaced by unrefreshed inventory after selection. Protected source data stays inside UNTRUSTED_DATA; server counters and policy remain outside.
  3. After the model returns, freshly revalidate the exact frozen exposed dependencies before any derived search, fetch priority, observation or finalization is accepted. Title/metadata changes count even when text is unchanged. Reject stale decisions; never substitute newer evidence into the old decision's dependency snapshot. Recheck dependencies again before later queued derived work as needed.
  4. Bind finalization's document-ID schema to the frozen exposed IDs. Quotes must be exact exposed substrings, not merely present in larger retained text; no title alias, fuzzy repair or invisible retained citation. Locations are published only when the provider confirmed them. Separate internal code-point matching from a claimed provider location; update investigationCitationSchema, resolveInvestigationCitation, result types and consuming renderers where needed.
  5. Lua resolve at netdocs-investigation-state.ts:514–531 currently accepts full retained text. Bind durable acceptance to the actual frozen exposed slice as well as TypeScript validation; update persisted exposure/citation schema, commands, Lua and eval mirror together. A quote found only after the exposed boundary is invalid even when the retained window contains it.
  6. Final, deterministic partial and replayed publication all freshly validate exact-quote evidence and relevant dependencies, under remaining time/attempt/HTTP/money guards. If validation cannot finish or evidence changed, withhold affected findings and report safe partial status; no cached permission fallback. Checkpoint invalidation and the surviving frontier before finish; publish a handle only after durable finish and owned connection/body drainage. Cancellation or failed drainage suppresses provisional output.

Redis prune at netdocs-investigation-state.ts:722–735 currently conflates window eviction with revocation and empties snapshots on routine retention changes. Separate ordinary retention eviction from explicit revoked/changed dependency invalidation. Routine eviction must not itself erase valid snapshots; replay still independently live-validates every cited source and relevant dependencies. Failure to validate withholds affected evidence; no stale replay fallback.

Unproven branch and acceptance boundary

Inspected current refresh compares window text but not title; invalidation is accumulated until the refresh walk ends, and step dispatches model-derived reads without an exact post-model frozen-dependency refresh. These are concrete replacement acceptance cases, not assurances from prior green suites. Historical verified review causes also remain regression obligations: cold pre-investigator authorization escaping owned time; SSE waiting for EOF or implicit model streaming; persisted unfetched IDs exposed without fresh validation; sibling work lost on evidence replacement; advancing empty/duplicate pages stranded; MCP answers reduced to counts; supported search fields lost; local pre-dispatch body rejection retaining an unknown model reservation. Repairs on the unfinished branch need fresh evidence, not assumed closure.

Review adjudication: Main verified and kept all six corrections: mechanical drainage, retention/revocation separation, durable exposed-slice citations, unchanged three-slot money semantics, pre-model refresh admission and every-read cooldown. Byte-budget planning is accepted; the claim that all independent maxima must fit together or contextBytes must automatically rise is rejected. The aggregate byte guard is intentional.

Required behavior proofs: fifth decision then further authorized reads without a sixth model call; insufficient exact-refresh attempts or model-plus-refresh time skips the model but continues traversal; persisted cooldown through each refresh phase versus true permission denial; ordinary eviction preserves replay snapshots while revoked/changed sources fail fresh validation; quotes only after the exposed boundary fail TS/Lua/eval acceptance. Exercise Unicode/non-BMP and escaped-text state pressure, actual context/envelope/wire bytes with reserved stop space, honest omissions and resumable unprocessed cursor work without candidate loss. Prove three-slot accounting and original attempt.turn late settlement survive content-version cutover. Use the matching plan's existing behavior-test surfaces; no product implementation or test execution is claimed here.

Real acceptance is the original unhinted question through actual north__netdocs_deep_search with model: opus, effort: medium, following real returned continuation handles, plus the existing Astra acceptance flow with unchanged model semantics. No private bypass, direct fallback or transport-only exit code passes this gate. The answer must distinguish supported budget contents, project identity, filing scope, latest and approval; flattened text cannot establish workbook cells or ambiguous numeric associations. Pending work and bounded omissions must remain visible even when the specific question is answered. Read-only acceptance is already authorized under the established lane identity and $50 allowance; this documentation seat does not execute it.

Approved continuation, engineering decisions and release

The approved integration is 600,000ms active lifetime, existing 300,000ms per call and 5/80/200 whole-investigation ceilings, not a latency SLA or renewed spending request. Continue after plan-review convergence without repeated approval or interruptions unless post-implementation tests fail. Main resolves configuration/accounting without increasing financial caps; dynamic retention must prove byte fit and preserve pending provenance, while three money-turn slots remain unchanged. Public model/effort support stays pinned across resumes; no model switch remedies failed acceptance. Keep default-off pending fresh reviews, gates and separate enablement; no merge/deploy authorization. Rollback disables investigation start/resume on both surfaces while direct tools remain available; no legacy model-driven fallback, extra search strategy or retry platform.

Latest superseding design — completion first; design only

In-progress / default-off; not implemented or semantically accepted. The user's ok let's try that approved evidence-first planning, identity reconciliation and normal continuations; we don't care about the limits to be honest it's a critical feature so we need that supersedes the 34-path/5,600-LOC and 6-model/8-research/35-second/3-turn/105-second acceptance ceilings. Reliability and completed research govern acceptance, not fitting those numbers. This section takes precedence over every historical “current”, cap, routing and first-call requirement below; history and safety evidence remain intact.

Authority is not unlimited execution or spend. This assignment permits four existing documentation files only. Keep authorization, fresh ACLs, source-injection isolation, cancellation, owned local drainage and durable financial admission unchanged. Preserve $5/turn, $15/investigation and $50/organization UTC-day pending separate financial authorization; the approved EVAL authority is $50 aggregate, including historical actuals, unknowns and reservations, never $50 extra. No paid/live calls, services, workers, schema changes or enablement here.

Observed cause and chosen module

Source: packages/agent-runtime/src/tools/netdocs-investigation.ts:modelDecision sends request, guidance, correction, omissions and UNTRUSTED_DATA, but no authoritative remaining allowances. The prompt already says fetch promising documents, follow cited aliases and inspect both alternatives. Main's latest real pair nevertheless spent all eight research reads without a fetch. More prompt reminders or changing 8 to 80 alone do not establish completion.

Deepen createNetDocsInvestigator and createInvestigationState, retaining their start/resume interface and existing product/MCP adapters. No background platform, callback service, recursive agent or second connector. Replace private loop termination and structured planning inside these modules; keep Sonnet 4.6 initially. No evidence supports a stronger-model cure.

Evidence-first decisions and progress

  1. Maintain explicit research obligations: identity unresolved/linked/contradicted; project budget contents; observed filing scope; latest-version evidence; approval evidence. Each supported claim has exact read citations and source dependencies. “Answered” concerns the supported question, not complete project coverage; unknown latest/approval must stay unknown.
  2. Search for an anchor, then deliberately read a promising identity document before expanding similar-name searches. Link identities only from a retained passage; inspect contrary parcel/workstream evidence instead of merging names. Search the evidenced linked identity, then inspect each relevant financing alternative's contents. Two alternatives are an acceptance case, not a universal hardcoded two-document quota.
  3. Each decision names its unresolved obligation, expected evidence gain and short serial read plan. Executor validates allowed operations, observed IDs and dependencies; a search-only repeat with no new scope, cursor, evidence or recovery reason returns structured stagnation feedback and requests a changed plan or honest yield. No blanket search quota or ban on justified refresh, discovery recovery, pagination or alias search.
  4. Build progress before presentation truncation from freshly validated structured state: completed reads/windows, unresolved obligations, failed/pending traversals, omissions and whether an equivalent search added any usable evidence. Replace repeatedSearchSameLeads cleanly, including its presentation-order-dependent comparison; do not turn per-query counts into global novelty. Unknown/truncated comparisons remain unknown, never “exhausted”.
  5. Ephemeral discovery's single exposure is deliberate ACL safety. Never persist/re-expose unfetched candidate IDs as authorized memory. A current decision may fetch a currently exposed lead; after a yield, rerun its authorized discovery query if necessary. Evidence-dependent plans, aliases, scope and feedback are pruned transitively on revocation. Source excerpts remain UNTRUSTED_DATA, not planner authority.

Authoritative feedback: extend fenced state replies with a server-generated allowance snapshot: claim/revision, absolute research-yield and hard deadlines, expiry, consumed research/model/ACL/HTTP work, remaining financial admission headroom after actuals plus unresolved reservations, and limiting reason. If no fixed count cap exists, say deadline/financial/rate-limiter constrained rather than inventing “remaining reads”. Include a trusted control object outside source data on every model decision; refresh after admission/settlement and ACL work, not from model arithmetic. It is informational, never an authorization token; atomic admission still decides each dispatch.

Actual defaults → proposed policy

Historical (inspected 2026-09-10): the limits in the left column are the measurements of that date and are retained as evidence; the current transport bound is the 120,000-byte cap recorded in the canonical contract §6/§7 and in the transport addendum below.

Inspected source/defaultProposed change and reason
netdocs-investigation.ts:run: 35s active, 40s cleanup; model 2,000 output tokens, zero retries, 48KiB actual request cap; graph recursionLimit 9.Initially retain 35s hard active + 5s local cleanup as an operational slice, not a completion SLA. Add cooperative research yield at 30s (or earlier when the next operation cannot leave the 5s active publication reserve). Use that reserve for fresh receipt validation and checkpoint; cleanup is not research time. Replace recursionLimit 9 as a hidden completion ceiling with the module's explicit deadline/financial/progress-controlled serial loop. Retain wire/output/retry guards.
netdocs-investigation-state.ts:investigationStateLua: 6 model/8 research/64 HTTP per turn; 3 generations, 192 HTTP and 105s cumulative active lifetime.Remove fixed research/model/HTTP-count and three-generation/105s completion stops; continue while evidence can advance within the current slice, fixed state expiry, financial admission and existing account limiter. Keep actual counters and pre-dispatch admission for every HTTP request, including ACL and protocol overhead. No unlimited tight loop: stop on expiry, cancellation, storage/auth failure, financial denial or unresolved stagnation; blocked uncertainty is not absence.
State version 2: three precomputed handle hashes, three money-turn entries, 45s fenced execution lease, fixed 30-minute content expiry.Version continuation content for demand-minted generations and dynamically keyed turn accounting; retain one executing turn/caller, 45s lease and fixed 30-minute non-sliding expiry initially. Keep only a bounded replay window (latest three receipts), not three executable turns; older handles return generic restart without research or disclosure. Compact authorized evidence within current storage caps and report omissions, not a silent proof of coverage.
apps/mcp-server/src/tool-deadline.ts: 45s default, APP_MCP_READ_TIMEOUT_MS clamped 5–55s; deep adapter uses min(start+35s, outer−5s), cleanup min(start+40s, outer).Keep chassis timeout and direct tools unchanged. Derive one shared slice policy for product/core/MCP; shortened outer deadlines shrink yield and publication time too. Do not extend a single RPC beyond host compatibility to satisfy end-to-end research. A later measured slice change must update state lease and adapter/core together.
netdocs.ts: 15s normal / 25s window operation, 15s window request; connector client.ts: 6s normal request.Retain transport defaults and invocation-scoped connection ownership. Bound each operation by remaining slice time; defer work that cannot safely finish, preserving the specific pending obligation. Do not add hidden retries or weaken scope-schema admission.
48KiB wire, 256KiB state, 8 retained document windows, 64 spans, 200 opaque candidates, 32 frontier entries; 3/caller, 30/org, 300 global admission.Retain memory/concurrency guards initially, not semantic shortcuts. Prioritize cited identity and budget windows; evict with dependency pruning and explicit omissions, and allow reacquisition through normal reads. If acceptance exposes evidence loss at these bounds, report that constraint rather than declare the question answered.

Accounting cutover: the inspected production Lua eval comparison is still 5,000,000 microUSD, whereas the eval harness and approved authority are 50,000,000. Reconcile the configured evaluation ceiling through the trusted state seam; never trust model/caller tool input or silently lift ordinary runtime caps. Preserve $3.03 full serial reservations, late/unknown charges, original UTC-day buckets and idempotent settlement. Versioned content expiry must not discard old monetary records; decode existing three-turn financial entries into the dynamic representation without reset, refund or loss of late settlement access.

Useful continuation, not timeout-shaped failure

Preserve public start {request} and resume {continuation_handle, guidance?}. On cooperative yield, stop new research, fresh-validate all contributing evidence, persist the pruned frontier and semantic obligations with fenced CAS, finish/rotate the opaque handle, then await owned connection/body drainage before publication. Return supported findings, identity-chain citations, pending actions, coverage gaps, limiting reason, expiry and a concrete continuation instruction. Distinguish cooperative yield from a hard deadline; no model call is required merely to manufacture a fallback receipt.

The host may resume the same tool normally, or present the useful partial result and offer continuation when its own turn ends; no automatic background work or guaranteed host obedience. Duplicate current handles replay only after fresh validation, never rerun charged research. Lost responses, stale writers, crash takeover and changed guidance keep fenced semantics. Old receipts outside the replay window give generic restart; unresolved financial attempts remain charged even then.

Cancellation/hard expiry always wins over publication, including during close: abort nested I/O, await owned local settlement, dispatch nothing later, and suppress unauthorized evidence or an undurable handle. Five seconds of publication headroom is an initial scheduling choice, not a guarantee that ACL/network drainage completes. If it cannot, return only safe generic status if deliverable; an existing handle may recover after lease expiry. A canceled first call can lose its handle. No remote rollback or exactly-once provider execution claim.

Semantic acceptance — still open

Use the unchanged natural question: the original private question (held only in the private oracle outside the repository; no confidential names, addresses or IDs appear here) Independently recover a cited identity chain and read both relevant budget contents. Separate “project budget exists in NetDocuments” from “filed under this exact matter”, and both from latest/approved. A partial 43-of-55 listing cannot establish filing under a third matter or absence. Do not invent amount/column associations from flattened numbers. Private expected IDs/addresses remain evaluation evidence, never production hints or fixture special cases.

Prove this through the full supported start → useful yield → resume → supported answer path, without the experiment's one-deep-call/12-tool ceilings. Direct fallback success and transport exit 0 are separate observations, not deep-search acceptance. Reuse syntheticCorpus's alias memo, different parcel, two financing alternatives, shuffled/held-out/no-alias/injection and long-window cases. Measure end-to-end and per-slice runtime, model/research/ACL/actual HTTP counts, actual plus unresolved cost, progress and unsupported claims; no paid matrix or new live authority is implied.

Transport addendum — 2026-09-11 approved multi-model bounded transport

Design decision. The investigator's model seam is selected per start, not per environment. A start may name modelastra | opus | grok | muse and an effort; the pure selector (tools/netdocs-investigation-selection.ts) maps the friendly id to a canonical registry id and the registry alone owns supported efforts and defaults: Astra (openrouter:openai/gpt-6-astra) and Opus 5 (openrouter:anthropic/claude-opus-5) low, medium, high, xhigh, max, defaults low / medium; Grok 4.6 (openrouter:x-ai/grok-4.6) low, medium, high, xhigh, default low; Muse Spark 1.2 (openrouter:meta/muse-spark-1.2) minimal, low, medium, high, xhigh, default low. An omitted model preserves the env/default direct-Anthropic policy; a direct-Anthropic model runs with reasoning disabled (internal effort none, not a public selector). Fable and other catalogue entries are not selectable. An unsupported pair is refused before authorization, state or budget on both surfaces; resume accepts only the existing handle and guidance and rejects a model or effort. The effective pair is pinned on the envelope for every new production or eval start, echoed by every claim and re-validated before any work, so later env changes cannot switch it; envelopes predating pinning (both fields absent) resolve today's default, a partial or unsupported stored pair stops configuration. Accounting reports the effective pair. All four routes are client-approved and never fall back (Astra and Opus 5 are utility/eval-only; Grok 4.6 and Muse Spark 1.2 additionally keep their existing, intentionally preserved ordinary-chat support, which this addendum neither adds nor withdraws; the investigation's bounded effort stays utility/eval on every route): unavailability is a configuration stop, not a silent default. No chat rollout, default change or Anthropic-path change. The completion-driven redesign above is design-only and untouched by this addendum.

Routing. Astra Azure-only ZDR; Opus 5 Amazon Bedrock-only ZDR; Grok 4.6 xAI ZDR under the shared policy; Muse Spark 1.2 the sole approved non-ZDR exception (Meta, zdr: false) with data_collection: deny and no fallback retained. Every other unpinned route stays refused.

Bound. One bounded request shape for every route: complete serialized body at most 120,000 UTF-8 bytes (49,152 was the historical cap when this addendum was first written; the canonical contract §6 carries the current figure), zero retries at both SDK layers, non-streaming, the pinned effort sent as explicit native reasoning.effort, 4,000 total output tokens including reasoning for OpenRouter and 2,000 for direct Anthropic; effort raises no limit. Temperature follows the registry (Astra and Opus omit it; Grok and Muse keep the supported setting). A final guard on the SDK-serialized body pins slug, the model's provider policy (deny / no fallbacks / single pin / zdr except the Muse exception), usage accounting, tool_choice: auto and parallel_tool_calls: false, rejecting caching, aliases, routing and output overrides before dispatch. The unchanged 3,030,000 microUSD state reservation is the admission unit; a distinct bounded-price capability (wire bytes as the input-token ceiling at conservative regional ceilings 11/55, 5.5/27.5, 2/6 and 1.25/4.25 USD per million, verified 2026-09-11 separately from the global price date) must fit inside it: at the current 120,000-byte cap with 4,000 output tokens, Astra 1,540,000, Opus 5 770,000, Grok 4.6 264,000 and Muse 167,000 microUSD, direct Sonnet (2,000 output tokens) 390,000 microUSD, all inside 3,030,000 (the historical 760,672 / 380,336 figures were computed against the 49,152-byte cap and are kept as history only; the corresponding pricing regression is the core worker's to update). OpenRouter actuals are the payload's inclusive usage.cost only; invalid or missing cost is an unknown attempt. Owned fetch and body drainage settle before the attempt does.

Evidence and limit. Real MCP smoke on frozen source, 2026-09-11, env-selected: Astra low 3 receipts / 11 models / 18 research / 28 HTTP / $0.327326; Opus 5 medium 3 receipts / 14 models / 13 research / 23 HTTP / $1.120465; zero unknown, zero reserved, correct accounting.modelId. This proves transport, start/resume and accounting only and remains valid for those pinned routes. Per-start selection through actual tool calls on a server defaulting to Sonnet: grok/high and muse/minimal, three receipts each, canonical model and effort correct throughout, both inconclusive with no findings. Those runs exposed a pre-dispatch accounting defect (the local body guard can retain an unknown reservation), being fixed separately; no financial proof or final counts are claimed for those routes until Main re-runs. No run produced a finding, so Semantic acceptance — still open above stands as failed. Feature remains in progress and default-off.

Historical record below — prior ceilings and sequencing superseded

Current disposition — invocation reuse implemented; paid core acceptance FAILED

Invocation-scoped reuse remains implemented and offline verified. Approved four-line prompt correction: independent Astra medium nd-prompt-a p6E / nd-prompt-b p6F CLEAN, source-only; no cross-model or semantic acceptance. Main's post-prompt root lint exit0/52.25s, typecheck exit0/64 of 64/52 cached/Turbo1m53.273s/wall129.23s. Feature remains in-progress/default-off; prompt correction INEFFECTIVE on the actual MCP run.

Latest Main-personally-executed /tmp/netdocs-deepsearch-main.RBZWm93I/client.ts --paid called actual north__netdocs_deep_search, unchanged budget question: capture.jsonl proves 1 deep call/6 nested dispatches/1 receipt; deep-search-receipt.json is partial/budget, 6 models/7 research/19 ACL/57 HTTP/19,673ms. Repeated original query; budgets not fetched. Full command 21.76s/exit0 is transport ONLY, not semantic acceptance or tool-research success. Increment190839 microUSD; cumulative 2563944 microUSD/81 actual/0 unknown/0 reserved; paid retries STOPPED, original $50 TOTAL unchanged. Prior paid-core failure remains historical:189687 incremental,2373105 cumulative,75 actual,0 unknown/reserved.

Exact captured-response replay through actual SDK/core/fixture/ACL and a removed cloned ledger is RED: 6 models / 7 research / 19 ACL / 57 HTTP, exit1, 2.39s wall with fixed accounting clock. Requests3–6 retain source linkage and contradiction; SDK translation preserves the repeated search. Per-query noProgress=false does not establish global novelty. Command and source citations: scorecard, latest paid-core diagnosis.

Prior natural-host localhost synthetic-MCP direct-tool success found both budgets but did not select deep_search; it is not deep acceptance. Prior prefix counterfactual 5 models/6 research/16 ACL/35 HTTP/1,518ms with bothBudgets=true remains offline feasibility only.

User approval now authorizes exactly four generic lines after the existing Follow aliases line in netdocs-investigation-prompt.ts; this supersedes the earlier no-prompt-edit restriction only for those lines. Main's baseline 5,562 + 4 = 5,566/5,600 gross production additions, 34 paths, 34 remaining; no other product/test, fixture or runtime/financial-limit change.

Hermetic investigator 187 passed/1 skipped, 5.98s; prompt Prettier/ESLint exit0 and prior private offline TS/catalog/preflight/SDK checks remain preparation evidence. Main's paid flow exercised actual SDK client → MCP handler → factory → core/model SDK → source reads → receipt, with synthetic auth/principal/state/source, not real-user documents. All old failed captures and prior direct-host success remain distinct. Read-only per-query feedback proposal and LOC estimate are in the scorecard, awaiting Main/user approval before any contract edit; no product fix or extra calls here. Metadata in-progress/default-off, generated STATUS/index and canonical contract unchanged; no PR/commit.

Latest approved implementation: model-projection-only repeatedSearchSameLeads compares all supplied arguments and equal observed lead sets only among retained validated certain complete first-page records; excludes failed/pending/restarted/cursor-history/truncated or 200-ID records. No IDs, receipt/state schema, noProgress, pagination/retry gate or prompt change. +25 production: 5,591/5,600, 34 paths, 9 remaining; +75 existing-test lines. RED before fix; final hermetic investigator 191 passed/1 skipped (6.23s), scoped format PASS/lint0 errors12 warnings, fresh private strict TS/catalog/SDK offline PASS. Main-only paid after fresh dual review using /tmp/netdocs-repeat-main.oIeTlbwr/client.ts; not run. Ledger unchanged 81 actual/2,563,944 microUSD/0 unknown or reserved. In-progress/default-off; no semantic acceptance. Prior failed captures remain binding; scorecard records commands and capture decision.

Historical preparation below — superseded sequencing, preserved evidence

The following DESIGN READY, not-implemented, review-pending, prior-spend and next-review statements describe earlier handoffs, not the current disposition above. Their measurements and failures remain evidence; technical contracts and acceptance criteria remain binding.

Status: in-progress, default-off; connection correction DESIGN READY for fresh dual review, not implemented, accepted or enabled; paid execution remains STOPPED. Historical R3 CLEAN TO IMPLEMENT reviews A=w7E:p30 and B=w7E:p41 preceded Main's STEP1 assignment, not this correction. Author: Astra medium under the existing explicit Astra-only/medium override, not cross-model review. Updated: 2026-09-10 (original design 2026-09-08). Main's recorded base/HEAD is 6bd0ae57, with origin/main 11 commits ahead; Main integrates before PR, no current-main verification. Earlier c789e830/3aa016ea rebase and eight byte-identical docs remain historical evidence.

Related: canonical provider contract; implementation plan; existing federation, firm-wide connection and long-document-read specs. Substantial: public tools, trusted prompts and resumable authorized state, despite a short plan.

Current scope and evidence — 2026-09-10

Approved authority: kwiss's exact Approve invocation-scoped ok let’s go / yes ok: invocation-scoped reuse and 34 implementation files / 5,600 total normally formatted gross production additions. Main's baseline remains 33 actual files / 5,230 additions; 370 remaining, no deletion credit or feature cut. Earlier 33/5,500 and smaller allocations are historical. This assignment is design/docs only; no schema, grants, migration, dependency or product edits. The scorecard preserves prior authority, failures and measurements.

The owner granted isolated fixture execution. Prior full runtime verification reported 5,375 passed / 111 skipped, trajectory script exit 0, and earlier actual Lua/PG verification 2 passed. All precede the current cap change; no rerun or current-main acceptance is implied. Latest cardinality repair uses native auto plus disable_parallel_tool_use, preserving one-decision, refusal and serial-plan semantics. Affected verification: 249 passed / 1 skipped; root lint and typecheck each exited 0. Two fresh independent Astra-medium delta reviews were scoped clean under the explicit Astra-only override, not cross-model review or full-feature acceptance.

September 10 kwiss authority: yes tell it it’s ok to spend more like 50$. EVAL ceiling: $50 TOTAL across all cases/repetitions, including historical actual cost, unknown costs and reservations; not an additional $50 or authorization for Fable. Runtime budgets remain $5/turn, $15/lifetime, $50/organization UTC-day. Main reconciled 64 actual calls totaling 2,025,969 microUSD before today's first case; no ledger reset.

First positive synthetic core-gate invocation today — failed acceptance: 22.00s, 5 model calls, 64 actual HTTP attempts, stop reason budget, findings 0, bothBudgets=false. Raw SDK calls: first search, then fetch memo plus access document, then repeat search. No later fixtures or resume ran; the wrapper intentionally stopped after the first invocation. CLI exit 2 / row unrun is not full paid acceptance. Added 157,449 microUSD; cumulative 2,183,418 microUSD, 69 actual calls, 0 unknown, 0 reserved. Paid work is STOPPED for unpaid captured-replay diagnosis. No acceptance, PR/CI, merge, deploy or enablement authority; no live calls allowed.

Approved invocation-scoped ownership — corrected design ready, implementation pending

Latest disposition: unpaid replay rejects source-context loss: reconstructed SDK requests 3–5 retain memo/linked alias and conflicting access text. Original outgoing requests are unavailable, so historical byte identity is unproven. Fetch coverage returned=0 is a count ambiguity, not text loss; storage and fresh ACL policy must not be patched to suppress the observed failure. This supersedes the diagnosis-pending sequencing above.

Actual-prefix feasibility FAIL: retain the captured initial search and BOTH memo/access fetches, then use the read passage's alias for search and returned IDs for both budget fetches. Real SDK, unchanged synthetic reads/state: 4 model calls, 6 research reads, 10 revalidation attempts, 64 HTTP, 1.216s active elapsed, stop budget, grounded fallback findings 1, bothBudgets=false, exit 1. Budget texts arrive at HTTP 49/61; the next memo ACL handshake consumes 62–64, denying its tool call before the fifth model/finalization. Counts are 48 initialize/notification/GET, one catalogue, two searches, four full fetches and nine dispatched ACL fetches. No weaker memo-only prefix, fixture easing or semantic-pass claim.

Native lifecycle evidence: tiny offline SDK experiment PASS, exit 0, 0.47s: two serial RPCs reuse one handshake; native close aborts/rejects before a held fetch drains; a closed transport cannot restart, requiring a new Client/transport. Explicit pending-work drainage remains necessary. The proposed owner's behavior and real-model latency/semantics remain unproven.

Corrected interface: connector withNetDocsConnection<T>(run: (close: () => Promise<void>) => Promise<T>) plus root export; private ALS, one serial generation, no raw SDK/context/resource bag or global close seam. Wrap the existing outer timing callback's run call and add a required close argument to private run, not an inline wrapper reindenting 191 lines. Its existing final finally awaits terminal idempotent close/drain with timer/listener live, then revokes provisional published/continuation on signal/deadline before cleanup. Findings and duration/event follow; wrapper finally awaits the same close promise as backstop.

Successful RPC explicitly finishes/cancels and drains operation POST bodies without waiting EOF/timeout or aborting retained transport; GET work stays generation-owned. Sticky GET/native failure fails the active SAME-generation operation, retires idle immediately and awaits retirement before permitted replacement; old callbacks cannot poison new work. Ignore only intentional POST completion, expected GET405 and already-retiring close callbacks. No idle SDK dispatch, including automatic ping/error replies. Preserve every fresh ACL, binding/lease/rotation check, limiter admission, HTTP reservation, deadline, byte limit and direct default; no hidden retry or remote rollback promise.

Planning ceiling, not measured fit: client 320 + export 1 + investigator 33 + canonical contract 8 = 362, plus 8 unallocated = 370; 5,230 + 370 = 5,600. Root export consumes approved file 34. Count formatted replacements/reindents without deletion credit; stop if implementation exceeds authority. No prompt edit, new repository file or raised runtime limit planned. The private connection-design correction specifies exact lifetime/races and added proof cases; fresh independent Astra dual review remains pending.

All acceptance criteria remain binding: actual-prefix first-call both-budget grounding within 6 model / 8 research / 64 HTTP / 35s; serial isolation, fresh ACLs, guarded recovery/rotation, current-operation admission/deadlines and held-close/persistent-GET teardown proof. No raised cap, input special case, weakened schema, freshness waiver or policy gate suppressing legitimate reads. Prior root gates remain historical after the cap change; no fresh root-gate claim.

Cap-boundary test correction, already approved and independent of connection repair: Main's affected run was 248 passed / 1 failed / 1 skipped, 75.30s; the old unknown-failure test relied on the smaller cap to stop after one dispatch, while USD 50 permitted nine. Seed valid historical settlements, each no greater than a reservation, leaving exactly one full reservation. The existing test now proves one dispatch, retained unknown charge, later cases unrun and unchanged historical attempts at the cumulative cap. Replace incidental reason wording with ledger-state assertions; no production accounting/stop policy change or weakened expected dispatch count. Writer ran the targeted test once post-fix: 1 passed / 64 skipped, 2.69s, then both affected suites: 249 passed / 1 skipped, 68.80s, exit 0. Hermetic success does not reverse the failed actual-prefix counterfactual or establish semantic acceptance; connection design remains NOT APPROVED.

Final test-only verification: a local attempt variable keeps historical seeding within 13 added test lines without changing behavior. Formatter unchanged; both affected suites re-run on the final form: 249 passed / 1 skipped, 69.25s, exit 0 (combined formatter/test wall time 70.83s). The earlier 68.80s pass is retained as history; targeted post-fix execution remained once. No production, financial or connection-design change.

Next: fresh dual review, not implementation or financial reapproval. Main accepted both reviewers' four findings: publication before drain, successful non-EOF POST, sticky generation-bound GET/native failure and gross reindent cost. Earlier NOT APPROVED/pending-allocation statements are retained historical dispositions, superseded by this approval. Failed paid/replay/counterfactual evidence and all acceptance remain binding. Financial cap $50 TOTAL; $2.183418 actual / $0 unknown / $0 reserved, 69 settlements; ledger/captures untouched. This correction ran no tests, paid calls, services or git and changed no product, prompt or unrelated telemetry.

Historical September 8–9 repair handoffs

“Current,” “latest,” allocation, ownership and unrun-proof statements in the next three paragraphs refer to their original handoff dates, not September 10. Their failures and bounded successes remain evidence; current scope, executed proofs and STOP above supersede their sequencing and ceilings without weakening any acceptance contract.

2026-09-09 current update, superseding the following historical slice allocations: Main relays kwiss's exact Allow 5,500 formatted additions selection via functions.ask; the hard cap is 31 implementation files including contract and 5,500 normally formatted gross production additions, with no feature cut. Initial eval harness repairs are locally verified except the retained empty-page-long first-call grounding assertion, which reaches the 64-HTTP allowance during receipt validation. A separate real typed-read/graph regression uses omitted-length A60000/B44878 windows, answers from both sources, and keeps observed model requests within 48KiB; this is not production Lua/resume or configured-model proof. Current combined tests: 219 passed / 1 failed / 1 skipped. Remaining proof expansion is stopped for Main's test-allocation reassessment, not production-cap overflow; all durability, protocol, infrastructure and acceptance requirements remain. See the scorecard for measured counts and the preparation-versus-historical-TLS distinction. No implemented status, acceptance or enablement.

Latest bounded writer handoff: The preceding RED/allocation-stop account is historical. Main reassessed test additions to 1,600; focused hermetic verification now passes 236 tests with two actual-target skips (20.08s). Script/test are normally formatted at 1,487/958 lines. Durable eval/reference/report/fixture work and guarded real-PG/Lua test code are written; actual proofs remain unrun. Omitted-length first-call grounding succeeds with maximum serialized wire 40,216 bytes; resume dispatches but stops storage in the bounded reference adapter. No oversized-wire core defect is confirmed and no cap was weakened. This is neither semantic acceptance nor completion: Main still owns real execution, paid evaluation, full scope measurement, independent reviews and gates. Details and exact opt-ins are recorded in the scorecard.

kwiss's scope selection, recorded with its exact wording and source in the scorecard, authorizes the full feature up to 31 implementation files and 3,600 gross production additions, not a target. This hermetic eight-repair slice is limited to 890 additional production additions from 1,744, with no new eval files or infrastructure execution. No behavior or ledger scope is cut. The private socket-owning helper promises awaited LOCAL teardown and no further dispatch, not remote rollback. Main owns final gates, separate isolated infrastructure verification and later eval assignment; both independent implementation reviews still require convergence. The earlier scope and schema-only notes below are historical evidence, not current limits or completion claims.

Historical schema-admission implementation and conditional authorization decision

This section preserves September 8 schema-only progress, conditional 20/22-file authority and source-assessment decisions. “Only,” “current session,” unassigned-worker and source-only proof claims below are historical, not current implementation or execution status. Placement rationale and teardown requirements remain binding; current scope and isolated execution evidence are above.

Only historical finding 4 is implemented in connector client.ts and test/client.test.ts: every outbound search argument requires an explicit compatible declaration on normal and bound-page paths; scope is never stripped. Production delta: 28 additions, 4 removals, 0 moved. Main ran bunx vitest run test/client.test.ts from packages/connector-netdocs-mcp without dotenv/live access: initially 222/224 passed; two positive fixtures omitted the supported attribute_filters declaration. Test-local compatible declarations corrected those fixtures without weakening assertions or admission. Main's rerun passed 224/224 in 8.10s. This is not full STEP1, investigator, cancellation, review or release acceptance. Earlier draft/unassigned and finding-4-survives statements below are historical snapshots superseded only to this extent.

Exact new authority, relayed by Main: dispatcher obtained kwiss's choice “Approve 22 files” directly via AskUserQuestion in the current session, then corrected its interpretation: 22 files and the existing 300-LOC contingency are a CONDITIONAL ceiling only; 1,800 total production LOC is unchanged. FIRST assess investigation-private, connection-owning authorization inside existing netdocs.ts using existing pg/@types/pg/drizzle dependencies. If it truly awaits local teardown and prevents subsequent dispatch, choose the original 20 files. DB-package placement was preferred, NOT technically required by dependencies. If confinement cannot meet the contract, report the concrete failure to Main before using packages/db/src/rls.ts plus a DB test addition. Placement is not selected or proven by this record; Main owns integration and the scout only reads. Scout Astra medium replaces smol Luna medium solely under kwiss's Astra-only override; other recorded routing remains unchanged.

Required implementation comment and PR-body checklist: promise awaited LOCAL teardown and no later dispatch, never instant remote rollback or custody-server cancellation. Explain the chosen placement and distinguish actual files/LOC from conditional ceilings. Reject the pool-counter workaround: it refuses warm idle pools and defeats reliability; address the shared-pool queued-acquisition race rather than treating a counter snapshot as ownership. Require biting mutation proof: disabling/removing teardown must make the test fail naming the detached query. These remain acceptance conditions; earlier Lua/PG results alone do not assert this specific mutation proof. No schema, migrations, grants or dependencies: stop BLOCKED on drift. No live calls; DB verification requires explicit isolated-target authority, not granted to this writer. This assignment changes only the two HTML pages, not status metadata, indexes, scorecard, STATE or product code.

Final Main source-assessment selection (supersedes pending placement above): choose the confined resolver in existing packages/agent-runtime/src/tools/netdocs.ts with already-budgeted netdocs-investigation.test.ts. Original 20 files remain the implementation target; the conditional 22-file fallback is unspent. Main reports public @workspace/db pool/schema and existing runtime pg/@types/pg/drizzle dependencies verified by source assessment. Scout estimates +175–250 production LOC within the existing 300-LOC contingency, not actual growth or behavioral proof; 1,800 total remains unchanged. Confinement avoids extra DB-package placement and shared-pool queued acquisition while retaining connection ownership. Use an owned standalone client/socket, fail-fast slot cap and strict awaited actual connect/query/local close; no shared-pool checkout. Critical pending-connect edge: destroy the socket FIRST, then await connect settlement/closure; do NOT call client.end first because pg can suppress the pending connect callback. Code comment and PR body must also disclaim remote cancellation and exact real-time guarantees under OS/event-loop stalls. Source-feasible only, pending implementation and biting teardown mutation proof; no DB verification claimed. Prepared next worker netdocs-owned-auth w7E:p44, Astra medium, is not yet assigned implementation; current writer retains sole ownership until this docs report. No further source exploration or product edits in this update.

One deep module, two adapters

Add netdocs_deep_search / north__netdocs_deep_search. Input is a strict union: start {request} (1–4,000 characters), or resume {continuation_handle, guidance?} (guidance at most 2,000 characters). Unknown fields fail validation. The server owns instructions, tool selection, budget and state. Initial invocation actually pursues an answer; neither discovery nor continuation is a mandatory first phase.

Retain browse, list_attributes, search, fetch and find_in_document for direct follow-up, preserving their public interfaces except rejecting undeclared provider scoping arguments (historical finding 4 below). The host can delegate again or follow evidence itself; no automatic callback, host obedience or automatic resumption is promised. No new continuation tool, general agent platform, package, managed-agent-harness code, ingestion or database migration.

Return one versioned result on both surfaces: status (answered / partial / inconclusive / refused / unavailable), stop_reason, evidence-linked findings, verified aliases, candidate matter scopes with limitations, coverage, unanswered questions, specific continuation guidance, general research guidance, and optional opaque continuation_handle + expires_at. IDs and source links are retained when supplied; absent links are not invented. A useful answer may still be partial.

Findings cite server-assigned evidence IDs resolving to document ID, observed title/link, retrieval time, exact quotation and code-point span; sheet/tab labels only when extraction actually exposes them. Search snippets are leads, not proof of budget contents. Alias verification requires a read passage linking identities; a similar name alone is insufficient. Narrative synthesis remains fallible even when citation existence and quoted text are mechanically checked.

Observed seams, not a second connector

Historical six-point check at 3aa016ea (after #387)

Read the supplied dormant-lane diagnosis as historical evidence only; no other worktree inspected. Factory ownership clearance is granted. Current source, not the old line numbers, determines these dispositions; Main tracks findings separately.

Historical findingCurrent disposition and evidence
1. Lookup keys / workspace-container / MCP document IDs conflatedClosed by #387: rest.ts and workspace-lookup handler removed; netdocs.ts netDocsTools (2536–2542), renderCabinets and renderListAttributes expose cabinet/attribute metadata, not workspace IDs. tool-copy.ts:77–86 explicitly denies folder descent; fetch remains search-document based. No old REST ID handoff remains.
2. Attribute-table REST call mislabeled/misparsed as workspace lookup (old rest.ts:368–374)Closed by #387; old parser is not applicable. client.ts parseListAttributesResponse (701–780) parses federated cabinets/attributes/skipped IDs; netdocs.ts readCabinetAttributes (1576–1616) calls listAttributes, not a REST lookup or get-or-create endpoint.
3. Impossible workspace-to-search guidance (old browse-copy.ts:7,13)Closed by #387: browse-copy.ts removed. tool-copy.ts NETDOCS_BROWSE_DESCRIPTION / NORTH_NETDOCS_BROWSE_DESCRIPTION name cabinets; list-attributes descriptions name attribute_filters. Both strict search schemas declare those arguments (netdocs.ts:1461–1492; MCP netdocs-search.ts:13–44), not generic filters.
4. Undeclared scoping arguments forwardedSurviving: client.ts safelyAcceptsSchema:1431–1445 rejects unknown keys only when additionalProperties is false; resolveListedTool:1548–1575 uses that check, and callTool:1726–1729 forwards toolArgs. Current capture declares cabinets/attribute_filters, so today's capture is not proof of an ignored filter; catalogue-drift admission remains unsafe. Step 1 requires explicit declaration for every outbound search argument, never dropping scope to search unscoped.
5. Captured field descriptions lost: historical 16, current 18Closed by #387: September 8 capture has 18 fields and 18 descriptions, including cabinets/attribute_filters. tool-copy.ts describeNetDocsSearchFields:44–66 attaches captured guidance on runtime netdocs.ts:1490–1492 and MCP netdocs-search.ts:42–44. Existing netdocs-tools.test.ts:171–207 checks exact capture parity on both surfaces (read, not run).
6. Unsupported document claim on REST failureClosed for removed REST paths by #387; not an absence claim in the historical source but an unsupported existence claim. netdocs.ts unavailableReceipt:655–656 still has document-specific copy; client.ts parseSearchResponse:503–513 and parseListAttributesResponse:708–712 now use generic provider_unavailable. The only connector producer of provider_format_unsupported is validated fetch text at client.ts:856–866. Preserve that fetch behavior; old REST failure is not applicable.

Bounded investigate → assess → refine

A private StateGraph alternates model decisions and serial allowlisted reads, checkpointing after each decision/read. It starts from the unchanged natural request. Discover attributes when scoping needs them, but plain searching survives discovery failure. Search distinctive terms, inspect promising passages, follow evidenced aliases across workstreams, assess contradictory identities, refine terms/scopes and paginate while useful. No recursive deep_search, knowledge, mail, web, writes or model-selected endpoints.

The system prompt distinguishes provider text from instructions, workstream from project, financing alternative from revision, and query traversal from project coverage. Coverage records each cabinet/filter/query, returned and deduplicated counts separately from provider totals, skipped cabinets, unreadable/truncated windows and outstanding cursors. Empty pages with continuation remain pending; cycles/no progress consume budget and terminate honestly. A completed query is never proof of complete project coverage or of the latest approved version.

The model emits strict validated decision/finalization objects through bound tools. Cap argument sizes, number of tool calls and evidence references before dispatch. Invalid schemas, unknown tools or unsupported references permit at most one corrective model call within the same budget; refusal, provider error, truncated/blank output or repeated invalid output returns a deterministic evidence-led partial receipt, not fabricated findings. A finalizer cannot mark unanswered coverage complete merely by saying done.

Only authored policy occupies the system message. Request/guidance are caller data; retrieved titles, metadata, quotes and provider errors are delimited untrusted tool data. Final output separates server research guidance from untrusted source excerpts; documents cannot set continuation policy, mutate the allowlist or supply handles. Validate quotes/spans against retained reads and URLs against observed provider metadata before output.

Defaults and accounting

ControlProposed default and rationale
Runtime modelAGENT_NETDOCS_INVESTIGATION_MODEL, default claude-sonnet-4-6 via createModel(utility/client/tools). Author-seat Astra override does not select runtime model. Require registered, policy-approved model with verified hard-cap pricing/limits; invalid configuration fails closed.
Wall clock35s active work + 5s checkpoint/receipt reserve inside MCP's 45s default; product uses the same envelope. Add optional caller signal to NetDocsOperationContext and optional outer deadline through every typed read, window/fallback, lease preparation/recovery, limiter wait/retry and provider fetch. Use min(local deadline, outer active deadline); compose caller/SDK/per-request signals, retaining abort through response-body consumption. Recheck after waits/reservations and immediately before dispatch; cancellation aborts nested I/O, not just the adapter. Preserve direct defaults (15s normal / 25s window / 6s provider request). Tests must prove canceled nested requests and no post-deadline dispatch.
Depth6 model calls including finalization/repair, 8 research reads per turn; 3 turns, 18 model calls, 24 research reads lifetime. At most 8 retained documents; authorization revalidation reads are separately counted but share all HTTP/time/lifetime limits. Serial dispatch, depth one. Optional investigation-only atomic reservation/attempt hook immediately before actual baseFetch in createBudgetedFetch counts initialize, notifications, each catalogue page, tool call and retry; limiter.consume currently counts executeAttempt, not HTTP, and the existing direct limiter remains unchanged. Replace 24/72 with 64 actual HTTP attempts/turn and 192/lifetime, including replay revalidation; no reset on replay. A cold positive path of 5 research + 3 authorization reads at 5 HTTP each leaves 24 attempts for additional pages/retries; this is a sizing example, not a provider guarantee. Measure the full exposure/revalidation sequence; cap exhaustion returns honest partials, never skips ACL checks.
Context/storage48 KiB COMPLETE serialized request per model call: authored prompt, original request, authorized structured state/guidance/read results, tool schemas and tool-call framing included. Enforce before dispatch for input/memory control, NOT financial token proof; no 200k proof or application pricing-input-limit API. Rebuild context, never replay retained freeform assistant/tool transcripts. No images, automatic history or hidden request additions. Constructor output cap 2,000 tokens; 256 KiB state including receipts, 64 evidence spans across at most 8 documents, 200 opaque internal candidate IDs, 32 pending queries. Trimming exposes generic omissions.
Money$5 hard turn / $15 lifetime / $50 organization UTC-day. Verified Sonnet 4.6 contract: truthful 1M input capability, 128K maximum output, $3/$15 base per MTok. Main's simplification, independently accepted by both R2 reviewers: keep resolveModelBudgetCapability's existing interface and full $3.03 reservation per attempt with constructor 2k output; only add verified provider metadata to model-pricing.ts. Calls SERIAL. Atomically require settled actual + unresolved reservations + new reservation <= each turn/lifetime/org-day/eval cap before dispatch. Exact resolved model policy/pricing still required; unsupported configuration fails closed. Constructor retry zero (other callers unchanged), prompt caching AND thinking explicitly off. An unknown $3.03 leaves at most $1.97, so no further default model call that turn: return deterministic authorized partial. Five calls settled at $0.10 then a sixth $3.03 reservation total $3.53 and fit $5; six unknown calls do NOT fit. Six calls remains a ceiling, not a guarantee.
AdmissionExisting NetDocs pilot/kill-switch plus new APP_NETDOCS_INVESTIGATION_ENABLED, default off. Redis atomically caps 3 unexpired investigations/caller, 30/organization, 300 globally; $50/organization/UTC-day reservations and one executing turn/caller. State limits bound memory and cap repeated starts, not just one handle.

Durably reserve under a unique attempt ID before invocation. Fenced, atomic, idempotent settlement replaces that reservation with VALIDATED actual cost once; failure, missing/invalid usage, cancellation or crash retains the full reservation. Output-schema validity and accounting are independent: schema-invalid output can still carry valid billable usage. Late settlement preserves the original UTC-day bucket. Unexpected usage above the bound stops work and records a violation and actual cost, never clamps it. Disable retries below the model seam as well as in the graph. Charge nested usage once to existing llm/anthropic.ts diagnostic channels; all typed count/enum-only events use the injected sink, with lazy production analytics defaults and preserved direct settlements. No prompts or document text in telemetry.

Continuation is authorized state, not pagination

Mint 256-bit random handles, store only their hashes as keys; fixed 30-minute lifetime from start, never sliding. State includes original request, verified evidence/aliases, frontier, coverage, pending questions, budget reservations, schema/prompt version and last receipt. Provider cursors are internal frontier data only. No access/refresh tokens or credential-bearing headers in state.

Existing connector next_page_token values are process-local nd1 handles (client.ts encodeBoundPageToken/decodeBoundPageToken, 15-minute TTL), not durable raw cursors. On token loss/expiry after restart, restart that saved query from page one, deduplicate by document ID, and mark traversal restarted/incomplete. This consumes the remaining lifetime budget; never claim uninterrupted pagination or reset evidence/budget state.

Bind organization + actor + credential owner + connector + accessMode + accessScope + endpoint + existing leaseId. Freshly resolve access and acquire a lease before opening caller-bound state, each nested read and evidence return. netDocsLeaseId(accessToken) (connector-auth netdocs.ts:111–115; token load tokens/encrypt.ts:390) is stable for the same bearer and changes on rotation. Never bind authorization to connectors.updated_at: normal lease mint reserves/advances that custody generation (connectors-api netdocs-mcp.ts:1092,1798,1952), invalidating its own reads and concurrent callers. Conservatively restart on tuple/lease change, including token rotation; no new custody protocol, schema or hash layer. accountId remains connector-derived, not remote user identity.

Connector validity is not document ACL: the lease can stay unchanged after revocation. Gate EVERY investigator model request and receipt, including FIRST-TURN cached fetch/find results, resume/replay and finalization. Investigation-only fresh provider ACL validation bypasses document/window caches and requires returned ID = requested ID; direct cache/default behavior stays unchanged. Validate every contributing retained document (at most 8); an immediately preceding fresh read may satisfy that exposure, never a persistent grant. Track explicit source-document dependencies for ALL retained derived state: aliases, queries/frontier, scopes, coverage, guidance, pending questions, findings and receipts. Prune transitively before exposure AND derived-frontier execution; drop uncertain derived state conservatively. Rebuild model context from authorized structured state, never retained freeform assistant/tool transcript replay. Keep unvalidated candidate IDs opaque/internal and omit their titles/snippets/attributes, not 200 fetches under an 8-doc cap. Fresh provider search results authorized at read time may be leads, NOT budget proof; cached leads do not inherit that grant. Revocation, timeout, ID mismatch, ambiguous response or exhausted budget suppresses the source and all dependents; return independently authorized partials with generic omissions. All checks consume HTTP/time/lifetime budgets, including replay, without TTL extension. Test direct cache populated → remote ACL revoked → first deep call, and alias memo revoked under same lease → no pending-query execution, coverage or old-assistant-content leak. No general provenance/ACL service.

Atomic Redis claim/revision CAS consumes each handle generation at most once, takes a 45s execution lease with fencing token, and reserves work before side effects. Concurrent calls return busy. Settlement rotates the handle; replay claims a fenced revalidation-only operation and filters its cached receipt after fresh access checks, without rerunning research/model work or refunding reservations. Replay is not zero work: it shares its original turn's remaining 64-HTTP allowance and a 105s cumulative active-work lifetime allowance (three 35s turns, including replay; unknown elapsed reservations retained). No counter reset on replay; exhaustion yields only generic restart/partial output, never unchecked evidence. Keep at most three generation receipts within state cap; same-handle lost-response retries supported, changed-guidance replay refused.

Checkpoint each completed operation with its reservation and fence; stale owners cannot write after takeover. After a crash, wait for lease expiry, keep unresolved reservations charged and resume from the last durable frontier under a new fence. No lease renewal beyond the turn deadline. Expired/corrupt/incompatible/mismatched state produces a generic restart receipt without revealing its existence to another caller.

Redis failure before reservation means no model/provider work. Failure mid-turn stops further work; return only currently authorized verified evidence with storage-unavailable and no newly promised handle. Cancellation may prevent delivery altogether: existing-handle callers can retry after lease expiry, while a disconnected first call may lose its response and must start again. No claim of exactly-once provider execution or durable job delivery. Revocation/identity change suppresses previously stored evidence and handle replay.

Acceptance, risk and release

Workbook format gate passed; loop behavior unproven

The September 8 captured fetch description promises document text and character windows, not XLSX worksheet coverage. client.ts fetch:1925–1943 calls only federated MCP; parseFetchResponse:782–898 requires text, bounds/redacts it and refuses binary/control content. netdocs.ts netDocsFetchTool:1921–1980 slices provider text or its cached window; find_in_document searches that same text. Neither downloads bytes nor invokes a spreadsheet parser. The comment at client.ts:854–855 describes a parked REST/extraction fix, not an available capability.

Main reports text-read feasibility PASSED at 3aa016ea using exact production netDocsSearchTool/netDocsFetchTool: first run 3 top-level reads, 8.09s, exit 0; two additional shape-only runs each 2 reads passed in 5.14s and 6.25s, totaling 7 successful top-level source calls. First run: workbook A 60,000 characters truncated; B 44,878 complete. Both retain DEVELOPMENT BUDGET, SOURCES AND USES, numeric and financing-option markers. A's first 60,000 characters are FLATTENED single-line text: zero physical newlines/tabs, zero literal backslash-n/r/t, zero pipe Markdown rows, no HTML table or JSON-like start; headings are not standalone lines and have no Sheet/Worksheet prefix. This is NOT preserved worksheet rows/cells. Exact numeric accuracy, column/value association, full workbook/sheet coverage and new-loop behavior remain acceptance, not verified by readable markers. Existing text path works; this does not automatically require a parser/download seam.

Historical September 8 probe authority — not current eval spend

The following $0/$5, docs-only and pending-assignment statements describe the original probe and permission relay. The September 10 EVAL ceiling does not renew the spent live exception; no live calls are allowed.

The probe used its own pinned actor through NORMAL authorized firm-shared access; a different credential owner is legitimate, not an alternate actor. Earlier three diagnostics stopped before reads due an overstrict owner==actor probe guard, not app/provider failure. Authorization covered normal operational lease/rate/analytics writes only, no document mutation/upload/manual DB writes/DB tests/alternate actor. Safe evidence payload was counts/booleans only; no real values, IDs, titles or text stored. Model fixtures remain fully synthetic. Final shape metadata is received: STOP live checking, no further probe or waiting. No probe by this fixer; $0 model spend of approved $5 TOTAL.

Binding permission record — LIVE EXCEPTION SPENT. Dispatcher steering relayed in Main w7E:p3F during the current September 8 session, after the final shape metadata; dispatcher independently confirmed with kwiss: “That word does not generalise. Any further live call under his actor, any write beyond lease/rate/analytics rows, or any spend past the $5 cap needs a fresh word through me — not a re-reading of the same approval.” Original UI selections were exactly “Allow normal reads” and “Allow up to $5” in Main w7E:p3F shortly before the 21:23 UTC plan-fix launch, NOT typed pane messages; no precise approval timestamp is asserted. Authorization relays must quote exact words, pane and approximate time. Seven cumulative source calls, none further planned; $0/$5 runtime eval spend. No further live calls/further writes or over-cap spend without fresh word THROUGH DISPATCHER; future smoke/pilot cannot reuse prior approval. Docs-only until reviews converge and Main assigns implementation. Flattened text feasibility is not numeric accuracy/column association proof. Permission record only; technical scope/caps unchanged.

Unchanged synthetic acceptance and release requirements

Synthetic core fixture: unchanged question “What is the project budget for 410 Example Avenue?” A status memo links it to 88 Sample Street via parcel SYN-P17; two budget extracts contain DEVELOPMENT BUDGET headings and distinct financing options. Extracts MUST include flattened inline headings/numeric groups and ambiguous column/value-association cases; do not invent preserved newlines, tabs, cell coordinates or verified numeric associations. The discovered matter is access/license-only; a similar-name matter is another project. First call must follow the memo, read both budgets and answer supported content with evidence and unresolved latest/approved status; ambiguous amounts/associations must remain explicitly unanswered rather than guessed. No supplied IDs or mandatory second call.

Vary identities, topic and document order; include no alias, contradictory parcel, false similar-name lead, multi-cabinet schema differences, empty page with real cursor, inflated totals, missing budget, long/truncated reads and source-injection cases. A deterministic provider transcript tests dispatch and receipts; a real configured model against the same synthetic provider tests judgment from the unchanged question. A separate unguided host run measures selection of deep_search versus direct tools without requiring selection for server correctness.

Historical R3 plan reviews A=w7E:p30 and B=w7E:p41 independently returned CLEAN TO IMPLEMENT under the Astra-medium override; Main accepted convergence and assigned STEP1 after rebase. This is not cross-model review or implementation-review convergence. Factory has cleared all NetDocs surfaces. Release still requires implementation reviews, behavioral/auth/replay/failure tests and independently reviewed rollout; the current two scoped clean delta reviews do not discharge full acceptance. The current $50 TOTAL synthetic-model EVAL ceiling includes all historical actuals, unknown costs and reservations across all repetitions/scenarios. Stop rather than exceed that aggregate to finish the matrix; paid work is currently STOPPED for unpaid captured-replay diagnosis, not authorized to continue merely because headroom remains. Host selection is a separate measure. Pilot smoke approval remains separate from the spent two-workbook format-read exception above. Rollback disables only deep search, retaining direct tools and fixed TTL expiry. No PR/CI, merge, deploy or enablement authority.

Main owns verification and integration before PR. Enablement still requires evidence for the complete implementation and fresh reviews, measured DB-import/init isolation on success/rate-limit/recovery/error (sentinel asserts ZERO attempts even if analytics swallows errors), first-turn/transitive ACL and durable accounting proofs, complete serialized 48 KiB guard and nested cancellation including awaited local teardown, cold-catalogue/ACL-positive-case fit within 35s/64 HTTP, synthetic model behavior and separate host/pilot measures. Dated scoped passes above are not full acceptance: today's first positive core-gate invocation failed to reach both budgets; no later fixtures or resume ran. Paid eval uses one persistent $50 TOTAL ledger across ALL cases/repetitions, never reset; report unrun matrix if the next full reservation will not fit. Runtime $5 turn/$15 lifetime/$50 organization-day limits are unchanged. Current measurement is 33 files and 5,230/5,500 gross production additions with 270 headroom; original 20/22-file and 1,800-LOC limits are historical. Retain safe confinement, all first-call/alias/flattened-numeric ambiguity, authorization/ACL, limits, teardown and coverage requirements. Stop on unmet conditions or cap overflow; no safety reduction or silent expansion.