Resolves #17162
Branched off dev, not stacked on #17160 — the stacked-ticket lint correctly refuses a branch whose commits claim a ticket the PR does not declare. The two touch the same mint step, so whichever merges second takes a small conflict; this one is one line plus a comment, so it is the cheaper rebase.
The gap
#17148 established the principle for this pipeline: an optional DevIndex intake failure must not decide whether the corpus publishes. It fixed that at the stage level. The same defect sits one layer up at the step level, and it is now the larger exposure.
| line |
step |
| 69 |
Mint Publisher installation token |
| ~102 |
Mint Intake installation token ← no guard |
| ~125 |
Checkout repository |
| ~164 |
Setup Node.js |
| ~176 |
Run bounded Data Sync emission and publish |
The intake mint had no continue-on-error and runs before checkout, Node and emission, so any intake credential fault killed the job before a single corpus byte existed. The corpus is emitted under the reader identity and never consumes the intake token — an enrichment credential was deciding whether the whole repository published.
⚠ PREMISE ROUND-TRIP, RESOLVED — the premise above stands as filed. After opening I claimed two fault injections could not make the mint fail, and re-scoped this PR as thin on that basis. Both claims were wrong. Both probes did fail the mint: 31877595849 → HTTP 422 "The permissions requested are not granted to this installation"; 31879164402 → 404 Not Found. I had read conclusion from the jobs API — the field this PR's own continue-on-error guard rewrites to success on a failed step. The guard masked the failures that justify it, and I read that masking as evidence the guard was unnecessary. @neo-gpt falsified it. The re-scoping is withdrawn: this PR is better evidenced than at opening.
So a permission cleanup on neomjs-devindex-intake does fail the mint — HTTP 422, "The permissions requested are not granted to this installation" (run 31877595849) — and a removed installation fails it with 404 (run 31879164402). #17160 adds permission-contents: write, so workflow line and App grant are a coupled pair whose drift takes the mint down, not merely the stargazer read. A rotated or expired private key is a third trigger, untested. All three are exactly what this guard exists for, and 31879164402 is the live L4 proof it works: mint failed, job continued through checkout and Node to the preflight.
Why one line is the whole fix
continue-on-error: true needs no new error path behind it, because #17148 already built the landing zone:
- failed mint ⇒
steps.intake-token.outputs.token is empty
- empty
DATA_SYNC_INTAKE_TOKEN ⇒ assertDataSyncAccess throws "No intake token was provided"
emitGeneratedData defers a preflight failure ⇒ every stage runs, corpus publishes
- the deferred error is rethrown ⇒ run still exits non-zero
The failure moves later; it does not disappear. That fourth step is the load-bearing one — without it this would trade a loud freeze for indefinitely green runs over a dead intake, which is strictly worse. The Publisher mint stays fatal: it is the publish credential, and without it there is nothing to salvage.
Deltas
.github/workflows/data-sync-pipeline.yml — continue-on-error: true on the intake mint, with the rationale and the asymmetry against the Publisher recorded inline.
test/playwright/unit/ai/buildScripts/DataSyncPipeline.spec.mjs — one test pinning the runtime half.
Rejected
- Marking the step non-fatal and stopping there. Without the preserved non-zero exit, a revoked intake permission yields green runs forever over a dead intake. The point is to move the failure, not remove it.
- Guarding the Publisher mint too. It is the credential that performs the publish; degrading past it would produce a run that cannot do the one thing it exists for.
- Handling the empty token explicitly in
dataSyncPipeline.mjs. assertDataSyncAccess already rejects a falsy token with a specific message, and #17148's deferral already carries it. A second path would be a second thing to keep in sync.
Evidence: L2 (unit, complete) + L4 (a live dispatch that fails the mint for real — dispatched, still queued at time of opening) → L4 is the class the workflow half needs and no unit test can reach. Residual: the L4 observation itself, posted as a comment when the run lands; do not merge on the L2 half alone.
Test Evidence
npx playwright test -c test/playwright/playwright.config.unit.mjs \
test/playwright/unit/ai/buildScripts/DataSyncPipeline.spec.mjs
45 passed (3.6s)
New test: an ABSENT intake token degrades exactly like a denied one (#17162) — asserts the absent-token error is deferred rather than thrown, and that both reader-scoped stages (--emit-only, --include-labels) still execute. It exists because the workflow change otherwise trusts an unasserted path: if that case ever threw instead of deferring, a credential fault would resume killing publication while the workflow guard still looked correct.
Live falsifier — IN FLIGHT at time of opening, result to be posted as a comment either way. It is queued behind the first post-#17149 scheduled run (shared concurrency: data-sync-pipeline), so do not read this section as a completed observation yet.
The mint made to fail for real, not simulated. A throwaway branch requested permission-deployments: write, which neomjs-devindex-intake does not hold, so create-github-app-token fails the step genuinely: run 31877595849. The observable claim is that the job proceeds past the failed mint to checkout, Node and preflight, rather than dying at the mint. (A preflight_only dispatch then correctly aborts at the preflight, which is that mode's contract — it is asserting the credential topology, and the topology is genuinely broken. The point being proved here is reaching the preflight at all.)
An accidental first attempt put the bad permission on the Publisher mint instead (run 31877572286, superseded/cancelled) — recording it because it is the control this PR's asymmetry claims: the Publisher mint has no guard and must stay fatal.
Post-Merge Validation
Observable only on a scheduled dev run. Owned by the standing alarm rather than this PR's close target, which will be closed:
Residual-Owner: #17131
⚖️ Ada · @neo-opus-ada · Claude Opus 5 · Claude Code
Resolves #17162
Branched off
dev, not stacked on #17160 — the stacked-ticket lint correctly refuses a branch whose commits claim a ticket the PR does not declare. The two touch the same mint step, so whichever merges second takes a small conflict; this one is one line plus a comment, so it is the cheaper rebase.The gap
#17148established the principle for this pipeline: an optional DevIndex intake failure must not decide whether the corpus publishes. It fixed that at the stage level. The same defect sits one layer up at the step level, and it is now the larger exposure.The intake mint had no
continue-on-errorand runs before checkout, Node and emission, so any intake credential fault killed the job before a single corpus byte existed. The corpus is emitted under thereaderidentity and never consumes the intake token — an enrichment credential was deciding whether the whole repository published.So a permission cleanup on
neomjs-devindex-intakedoes fail the mint — HTTP 422, "The permissions requested are not granted to this installation" (run 31877595849) — and a removed installation fails it with 404 (run 31879164402).#17160addspermission-contents: write, so workflow line and App grant are a coupled pair whose drift takes the mint down, not merely the stargazer read. A rotated or expired private key is a third trigger, untested. All three are exactly what this guard exists for, and 31879164402 is the live L4 proof it works: mint failed, job continued through checkout and Node to the preflight.Why one line is the whole fix
continue-on-error: trueneeds no new error path behind it, because#17148already built the landing zone:steps.intake-token.outputs.tokenis emptyDATA_SYNC_INTAKE_TOKEN⇒assertDataSyncAccessthrows "No intake token was provided"emitGeneratedDatadefers a preflight failure ⇒ every stage runs, corpus publishesThe failure moves later; it does not disappear. That fourth step is the load-bearing one — without it this would trade a loud freeze for indefinitely green runs over a dead intake, which is strictly worse. The Publisher mint stays fatal: it is the publish credential, and without it there is nothing to salvage.
Deltas
.github/workflows/data-sync-pipeline.yml—continue-on-error: trueon the intake mint, with the rationale and the asymmetry against the Publisher recorded inline.test/playwright/unit/ai/buildScripts/DataSyncPipeline.spec.mjs— one test pinning the runtime half.Rejected
dataSyncPipeline.mjs.assertDataSyncAccessalready rejects a falsy token with a specific message, and#17148's deferral already carries it. A second path would be a second thing to keep in sync.Evidence: L2 (unit, complete) + L4 (a live dispatch that fails the mint for real — dispatched, still queued at time of opening) → L4 is the class the workflow half needs and no unit test can reach. Residual: the L4 observation itself, posted as a comment when the run lands; do not merge on the L2 half alone.
Test Evidence
New test:
an ABSENT intake token degrades exactly like a denied one (#17162)— asserts the absent-token error is deferred rather than thrown, and that both reader-scoped stages (--emit-only,--include-labels) still execute. It exists because the workflow change otherwise trusts an unasserted path: if that case ever threw instead of deferring, a credential fault would resume killing publication while the workflow guard still looked correct.Live falsifier — IN FLIGHT at time of opening, result to be posted as a comment either way. It is queued behind the first post-#17149 scheduled run (shared
concurrency: data-sync-pipeline), so do not read this section as a completed observation yet.The mint made to fail for real, not simulated. A throwaway branch requested
permission-deployments: write, whichneomjs-devindex-intakedoes not hold, socreate-github-app-tokenfails the step genuinely: run 31877595849. The observable claim is that the job proceeds past the failed mint to checkout, Node and preflight, rather than dying at the mint. (Apreflight_onlydispatch then correctly aborts at the preflight, which is that mode's contract — it is asserting the credential topology, and the topology is genuinely broken. The point being proved here is reaching the preflight at all.)An accidental first attempt put the bad permission on the Publisher mint instead (run 31877572286, superseded/cancelled) — recording it because it is the control this PR's asymmetry claims: the Publisher mint has no guard and must stay fatal.
Post-Merge Validation
Observable only on a scheduled
devrun. Owned by the standing alarm rather than this PR's close target, which will be closed:Residual-Owner: #17131
chore(data): Hourly data sync pipeline update [skip ci]commit.continue-on-errorannotation appears.⚖️ Ada ·
@neo-opus-ada· Claude Opus 5 · Claude Code