LearnNewsExamplesServices
Frontmatter
id17410
titleA tenant checkpoint can only be initialized by a pass that completes, so a corpus larger than one slice never checkpoints
stateClosed
labels
bugai
assigneesneo-opus-vega
createdAtAug 20, 2026, 10:09 AM
updatedAtAug 20, 2026, 10:16 AM
githubUrlhttps://github.com/neomjs/neo/issues/17410
authorneo-opus-vega
commentsCount1
parentIssuenull
subIssues[]
subIssuesCompleted0
subIssuesTotal0
contentTrust
projected
quarantined0
signals[]
blockedBy[]
blocking[]
closedAtAug 20, 2026, 10:16 AM

A tenant checkpoint can only be initialized by a pass that completes, so a corpus larger than one slice never checkpoints

neo-opus-vega
neo-opus-vega commented on Aug 20, 2026, 10:09 AM

Context

Observed live on an external pull-mode tenant deployment, 2026-08-20. One repository has been re-materializing its entire envelope every sweep since it was configured, embedding 1–2 chunks per sweep, and persisting nothing. Its outstanding count has not moved.

Four consecutive sweeps, ~7–11 minutes apart:

materialized: envelopeFiles=1586 envelopeDeleted=0 ingested=86947 deleted=0 embeddings=2 errors=0
partial-progress: slice budget reached, checkpoint held at none ingested=86947 embeddings=2

materialized: envelopeFiles=1586 envelopeDeleted=0 ingested=86947 deleted=0 embeddings=1 errors=0
partial-progress: slice budget reached, checkpoint held at none ingested=86947 embeddings=1

embeddings = 2, 1, 1, 1 across the four. Five embeddings in 26 minutes against 86,946 outstanding chunks.

The deadlock

tenantRepoCheckpointValidity.mjs:383:

if (!normalizedState?.lastIngestedRev) {
    return TenantRepoCheckpointStatus.UNINITIALIZED;
}

lastIngestedRev is written only by a pass that completes. A pass that yields on the slice budget reports partial-progress and leaves it null. So:

  1. Sweep admits the repo with sliceBudgetMs (default 5 * 60 * 1000, configBase.mjs:2182).
  2. The embed phase dispatches chunks. One oversized chunk consumes the entire budget on its own — measured below.
  3. createSliceBudgetPredicate fires after that single chunk.
  4. partial-progress → lastIngestedRev stays null.
  5. Next sweep: classifyTenantRepoCheckpoint returns UNINITIALIZED, there is no resume point, and materialization starts from zero.

The checkpoint's initialization precondition is the completion of the pass that the budget prevents.

Corrected 2026-08-20: what consumes the budget is ONE chunk, not materialization

My first framing blamed envelope materialization. Falsified by the provider's own slot log:

375.04.628  release  n_tokens = 545
375.05.587  release  n_tokens = 218
379.33.302  release  n_tokens = 13725    <- 4m28s gap, no other completion

Sub-1k-token chunks release ~1 s apart. A single 13,725-token chunk took ~4.5 minutes — the whole sliceBudgetMs. Rate check: 4.5 min/chunk implies 0.22 chunks/min against 0.19/min observed across four sweeps, so this single cost accounts for the whole stall.

Measured cost curve on this lane: ~400 tokens ≈ 1 s, 13,725 tokens ≈ 268 s — 268× the time for 34× the tokens, i.e. cost ∝ n^1.58. That is quadratic attention plus a linear term, which is what an unfused attention path gives. Consequence: the same content split into ten ~1,370-token chunks costs ~70 s instead of ~268 s and parallelizes across the four idle slots.

So the corpus's chunk-size distribution — sitting against our own 14,336-token admission ceiling — is the proximate cause, and the checkpoint deadlock is what makes it permanent rather than merely slow. Below a corpus size where one pass fits in one slice this is invisible; above it, the repo cannot ever initialize, and every sweep re-buys the identical materialization.

Why the existing fixes do not cover it

All of these are merged AND present in the observed deployment's revision (verified with git merge-base --is-ancestor against the plane's deployedRevision):

ticket what it fixed why it cannot help
#17112 batch ingestion persisted embeddings only per full slice VectorService — persists completed embeddings, not a sweep checkpoint
#16826 a lease yield discarded completed provider chunks same layer, same reason
#17062 embedding admission had no aging TextEmbeddingService — admission order, not checkpointing
#17132 a repo held its concurrency slot until its corpus was exhausted introduced the slice yield this defect now depends on
#17349 a clean partial slice accrued a failure streak correctly reports streak 0 -> 0; the repo is not failing, it is looping
#16822 kbSync checkpointed later than its fairness bound kbSync, which is disabled on pull-mode deployments

Each is correct and none of them can initialize a checkpoint. The defect is in the interaction, which is why it survived them.

Not the provider, and not the network

Ruled out on the live plane so the fix is not aimed at the wrong layer:

  • Provider throughput. launch_slot → release is ~16 ms per task, all four slots in use. The embedding container consumed 12.7 h CPU in 5.8 h wall — 2.18 of its 6-core quota — while the host sat at 2–12% of 64 threads. The lane is idle waiting for work.
  • Network. Host eth0 receive held at 9 Kb/s across three samples 2.5 minutes apart. Re-fetching 1,586 blobs per sweep would be megabits. The partial-clone mirror is fully backfilled (#16557), so materialization reads locally.
  • What the lane is actually serving. The observed inputs are 9–13 tokens with f_sim_best = 1.000 — the healthcheck's fixed 'neo-kb-healthcheck-embedding-canary' probe (HealthService.mjs:61), re-embedded on a cache hit. A 4-slot lane is predominantly answering its own liveness check.

So the cost of the loop is orchestrator CPU (one saturated Node event loop, cpus: "1.0") and local disk writes, not provider time or bandwidth.

Acceptance Criteria

  • RED-PROOF: a repo whose materialization cannot complete inside sliceBudgetMs advances a durable resume point across sweeps. Asserted by driving two consecutive sweeps over a fixture corpus sized above the budget and showing sweep 2 starts from sweep 1's position — not from zero.
  • SILENT ARM: a repo whose pass completes inside one slice still reports complete with lastIngestedRev set, so the change does not convert healthy repos into permanently-partial ones.
  • A partial pass is distinguishable from an uninitialized one. UNINITIALIZED currently means both "never ran" and "ran four times and got nowhere". The classifier must separate them, because the second is a defect and the first is a new repo.
  • One chunk cannot consume the whole budget. A slice that admits a chunk whose projected cost exceeds the remaining budget must either decline it to a splitter or carry it across sweeps — never spend an entire slice on one input and checkpoint nothing. Asserted on a fixture chunk sized to exceed the budget alone.
  • Cost is projected from token count, not assumed linear. The admission decision uses the measured n^1.58 shape for this lane, so a chunk 34× larger is budgeted as ~268× the cost rather than 34×. A linear projection is what admits a 4.5-minute chunk into a 5-minute slice.
  • No re-materialization without a cause. A sweep that resumes an unchanged revision does not rebuild the envelope. Asserted on call counts, not on timing.
  • Regression arm for the interaction itself: a fixture at 1× and at 20× the slice budget, so the size threshold where the deadlock appears is pinned rather than rediscovered.

Out of Scope

  • The truncation failure on a sibling repo (KB_VECTOR_EMBED_INPUT_TRUNCATED, separate ticket to follow) — different stage, different cause.
  • Lease starvation under this holder — #17398, fix already open as PR #17399.
  • Provider geometry (UBATCH, context-per-slot, per-call timeout ceilings) and enabling fused attention — real wins on the measured cost curve, but deployment tuning rather than this defect.
  • The corpus chunk-size distribution itself. Reducing the parser's maximum chunk size is the largest single lever here (n^1.58), but it changes parserVersion and therefore every chunk id, so it is a separate decision with a re-ingest cost.
  • Raising sliceBudgetMs. A larger budget moves the corpus size at which the deadlock appears without removing it.

Avoided Traps

  • Do not blame envelope materialization. I did, in the first version of this body, and the slot log falsified it. Network was also ruled out (9 Kb/s receive, mirror fully backfilled per #16557) before reaching for it.
  • Do not "fix" this by widening the provider request. BATCH_EMBEDDING_CHUNK_SIZE=1 against a 4-slot lane is wasteful and worth changing on its own, but the pipeline only requests 1–2 embeddings per slice, so widening the request while the checkpoint is broken changes nothing measurable and will read as a fix that did not work.
  • Do not add CPU to the embedding container. It runs at 2.18 of 6 cores. Starvation is upstream.
  • Do not infer the rate from the collection count. Total chunk count grew 22,313 → 29,347 overnight, which projects to a healthy-looking ETA; that growth was a different repo completing while this one made no progress. Per-repo corpusOutstanding is the honest observable.