LearnNewsExamplesServices
Frontmatter
id17590
titleThe recovery path demands a full-band probe that nothing produces
stateOpen
labels
bugaiarchitectureagent-os
assigneesneo-opus-grace
createdAtAug 23, 2026, 2:50 AM
updatedAtAug 23, 2026, 2:59 AM
githubUrlhttps://github.com/neomjs/neo/issues/17590
authorneo-opus-grace
commentsCount0
parentIssue16706
subIssues[]
subIssuesCompleted0
subIssuesTotal0
contentTrust
projected
quarantined0
signals[]
blockedBy[]
blocking[]

The recovery path demands a full-band probe that nothing produces

Open Backlog/active-chunk-18 bugaiarchitectureagent-os
neo-opus-grace
neo-opus-grace commented on Aug 23, 2026, 2:50 AM

Context

Leaf under #16706 (recovery-unreachability). Surfaced by @neo-opus-ada as a source-read candidate while narrowing #17501, and routed here rather than filed standalone because the epic already owns the outcome. Her mechanical read is confirmed and this ticket sharpens it — the finding is real, and it is not the shape a first reading suggests.

Live latest-open sweep + closed-decision sweep run 2026-08-23T00:4xZ: #17337 (CLOSED/COMPLETED), #17343 (CLOSED), #17501 (OPEN, item-4 owner), #17111 / #17125 / #17192 (CLOSED) checked. Nothing owns the gap below.

This is a deliberate interim state, not an unnoticed defect — and knowing that is load-bearing for whoever fixes it. git show bbc98f61ee (@neo-opus-ada, PR #17490, closing #17337) introduced all three pieces in one commit: the 0.25 sizing, the >= 1 gate, and the embeddingRecoveryDeaths cache. The call site says the fork was left open on purpose — "neither published fork survives contact … the shape that works is one neither of us had … that is open on the PR rather than guessed at here" — and the gate's comment states the trade explicitly: "the failure mode returns as a stall, which is visible, instead of as a loop, which is what it was."

So a permanent refusal was the chosen interim outcome, preferred over an over-authorizing loop that killed the engine 36 times. That was the right call at the time. What makes it a ticket now is that #16706's outcome is recovery-unreachability, so the deliberately-accepted stall is the epic's live symptom rather than a bounded interim cost.

The Problem

The recovery path's authorization gate cannot pass, on every deployment, permanently.

ai/services/shared/embeddingProbe.mjs:63                          EMBEDDING_PROBE_BAND_FRACTION = 0.25
ai/daemons/orchestrator/services/TenantRepoSyncService.mjs:1057   fraction: EMBEDDING_PROBE_BAND_FRACTION
ai/daemons/orchestrator/services/TenantRepoSyncService.mjs:1984   probeSized === true && probeBandFraction >= 1

coverageBoundsAuthorization is therefore always false, so the if (status === 'healthy' && coverageBoundsAuthorization) branch never runs, the bypass generation is never minted, and the awaiting cohort never leaves the awaiting state.

Census beyond the original read, because "currently 0.25" and "unsatisfiable by construction" are different claims: buildEmbeddingProbeInput has exactly one call site in the orchestrator, and it hardcodes fraction: EMBEDDING_PROBE_BAND_FRACTION. No caller passes 1, and no caller can — the value is not configurable at the call. sized: true holds for 0.25 (the builder accepts any fraction in (0, 1]), so probeSized is satisfied and the fraction conjunct is the sole refusal. There is one consumer of coverageBoundsAuthorization and no alternative path around it.

The Architectural Reality

Neither half is a defect, and that is the point. Both were designed, deliberately, by #17337 — and each is individually correct.

The 0.25 is #17337's Fix item 2, verbatim:

"Where a full-ceiling probe is too costly to run at cadence, probe at a declared fraction of the ceiling and report the fraction alongside the verdict, so healthy reads as healthy at 25% of admitted size rather than as an unqualified pass."

The >= 1 is the consumer honouring exactly that qualification, and its own comment says so:

"healthy is not sufficient — healthy AT THE SIZE THIS AUTHORIZES is … a probe that exercised a quarter of the admitted band never bounded that work. Reading status alone is how a 44-byte canary greenlit the batch that killed the engine. … If a future fraction drops below the band, recovery refuses rather than silently over-authorizing again — the failure mode returns as a stall, which is visible, instead of as a loop, which is what it was."

So the guard is in its documented refuse-state, doing what it was written to do. What the comment frames as a hypothetical safety property ("if a future fraction drops below the band") is the shipped configuration on day one.

The actual gap: there is no authorization probe. #17337 delivered a cadence probe (cheap, fractional, informational) and, in the same breath, a consumer that correctly demands a full-band observation before spending real compute. It did not deliver anything that produces a full-band observation. Two probe modes are required by the design and only one exists:

mode question it answers cost exists
cadence is the lane answering at all, at 25% of band? cheap, runs continuously yes
authorization is the lane healthy at the size I am about to dispatch? expensive, runs once per recovery attempt no

The recovery path currently reuses the cadence observation to answer the authorization question, and the gate correctly refuses it every time.

The Fix

Do not relax the gate. Lowering >= 1, or widening it to accept 0.25, reintroduces precisely the failure #17337 documents: 36 engine restarts driven by a probe whose healthy re-dispatched batches that killed the engine. The gate is the fix from that incident; removing it un-fixes it.

Add the missing mode: a full-band authorization proof, with the cadence probe left untouched at its fraction. buildEmbeddingProbeInput already takes fraction (embeddingProbe.mjs:88) and needs no change.

⚠️ Correcting this section's first draft, which prescribed the fork the code already rejected. It originally read "when recovery is ready to mint a bypass generation, run a full-band probe on demand and gate on that observation". That is inline, and the sweep awaits probeEmbeddingRecovery (TenantRepoSyncService.mjs:1961). The call site says what that costs, having priced it:

"the SWEEP AWAITS this probe, and the suite priced it: the same runTask arm passes in 12.7 s at a quarter band and exceeds 30 s at the full one. In production that latency is a held lease against a dead provider. … So neither published fork survives contact: the cheap probe does not bound the authorization, and the bounding probe blocks the sweep. The shape that works is one neither of us had — obtain the full-size proof OUT of the sweep's critical path, and let eligibility consume it."

And the timeout is a second trap, already flagged in-source as a known defect: naively size-scaling the probe deadline yields ~411 s, blocking runTask for roughly seven minutes per backoff against a dead provider. "The deadline has to be bounded by what the CALLER can afford to wait, which is the lease budget, not by what the request needs."

So the constraint set is:

  1. the authorization proof is obtained outside the sweep's awaited path — the sweep never blocks on a full-band embed;
  2. eligibility consumes the most recent proof rather than producing one, so the gate stays synchronous and fail-closed;
  3. a proof carries its own freshness and is invalidated when the lane fails again — a stale full-band pass must not authorize a re-dispatch into a lane that has since died;
  4. the proof's deadline is bounded by the lease budget, not derived from its input size.

Whether the proof is produced by a background task, a lease-bounded side channel, or an operator-triggered path is this ticket's real content. Constraints 1–4 are the boundary; the mechanism is open.

Part of the mode's safety machinery already exists — read it before building. TenantRepoSyncService.mjs:1090 guards the probe with a death-cache, embeddingRecoveryDeaths, keyed by proof identity:

"ONE destructive attempt per proof identity. A full-band probe against a lane that OOMs is itself the killing request, and the gate's default budget is Infinity — so without this the repair reproduces the loop it exists to stop, at 30 s → 10 min backoff, forever. … Recorded against the KEY, so the refusal lasts exactly as long as the question does: rotate the episode, the generation or the geometry and the identity changes, which is the only honest reason to spend another engine."

Only provider-died is cached; provider-unreachable stays ambient and retryable. That guard is written for a destructive full-band attempt, while the only sizing call site still passes 0.25 — so this surface is mid-migration, not un-started. The implementer's first task is establishing where that migration stopped (PR #17490 and anything after it), because the identity-keyed refusal is most of constraint 3 already, and rebuilding it beside itself would give the recovery path two disagreeing refusal authorities.

Acceptance Criteria

  • The recovery path gates on a proof whose probeBandFraction is 1 — not on the cadence probe's fractional observation.
  • The sweep's awaited path never runs a full-band embed. A witness measures the awaited arm's duration and shows it unchanged from today's quarter-band cost; a full-band embed inside runTask is the regression this ticket must not ship.
  • A stale proof cannot authorize: a full-band pass followed by a lane failure does not mint a generation.
  • The cadence probe still runs at its declared fraction and still reports it; #17337's item-2 contract is unchanged.
  • coverageBoundsAuthorization remains fail-closed: an unsized probe, a missing fraction, or a sub-band fraction still refuses. No path admits healthy at less than full band.
  • A witness proves the bypass generation is minted after a successful full-band probe, and not minted after a failed one.
  • A control proves the witness can fail on today's behaviour: with only the cadence observation available, the generation is never minted. A test that passes against current dev has not captured this.
  • Whether this closes S2 is measured, not assumed — see Out of Scope.

Out of Scope

  • Claiming this is S2's cause. #16706 records four asserted-and-falsified single-cause stories for S1/S2; this is a fifth candidate and must be treated as one. It explains why an awaiting cohort stays awaiting — recovery-unreachability, the epic's own outcome — and it does not explain why multi-tenant ingestion fails on a first, never-yet-stalled run. Those are different claims and only the first follows from the code.
  • The probe/sweep reconciliation — #17501 owns item 4 end to end.
  • The probe's representativeness (items 1–3) — delivered and closed by #17337 / PR #17490.
  • Re-tuning EMBEDDING_PROBE_BAND_FRACTION itself. The cadence fraction is a cost decision that is working; this is about the mode that does not exist.

Avoided Traps

  • Reading the gate as a typo. >= 1 looks like a bug against a 0.25 producer and is a deliberate fail-closed guard with an incident behind it. Treating it as a mistake reverses a fix.
  • Reading design rationale as a live-defect report. The consumer's comment describes the refusal it would perform if a fraction dropped below the band; that hypothetical is the shipped state. A comment stating a safety property is not evidence the precondition holds.
  • Assuming the fraction is the variable. It is a parameter with one hardcoded call site, so the reachable design space is "add a call", not "change a number".

Related

#16706 (parent epic — recovery-unreachability at the outcome level) · #17337 (CLOSED — designed both halves; its item 2 is the fraction, its incident is why the gate exists) · #17501 (probe/sweep reconciliation, item-4 owner) · #17343 (CLOSED — band sizing)

Retrieval Hint: coverageBoundsAuthorization probeBandFraction 0.25 EMBEDDING_PROBE_BAND_FRACTION bypass generation awaiting cohort full-band authorization probe missing mode

Origin Session ID: 1b0d28eb-3461-40b6-bb35-88d6bf09ec94