LearnNewsExamplesServices
Frontmatter
titlefeat(ai): diff declared config across revisions (#16765)
authorneo-gpt-emmy
stateMerged
createdAtAug 21, 2026, 2:01 PM
updatedAtAug 21, 2026, 3:39 PM
closedAtAug 21, 2026, 3:38 PM
mergedAtAug 21, 2026, 3:38 PM
branchesdev ← codex/16765-config-revision-diff-v2
urlhttps://github.com/neomjs/neo/pull/17459
contentTrust
projected
quarantined0
signals[]
Merged
neo-gpt-emmy
neo-gpt-emmy commented on Aug 21, 2026, 2:01 PM

Resolves #16765

Related: #16489 · #15294 · #16777

Neo can now answer the deployment question “which declared configuration inputs changed between revision A and revision B?” without checking out or executing either revision. revisionConfigDiff.mjs resolves immutable commits, discovers Tier-1 plus each revision's template-backed MCP config surfaces, reads git objects, statically parses the declared leaf() trees, composes the existing added/removed differ, adds same-path default/env/type changes, and emits one revision-config-diff.v1 JSON receipt. Added leaves carry a four-way target applicability verdict; bad refs, unreadable objects, unsupported declarations, pre-horizon ranges, and malformed current-model surfaces fail loud with distinct diagnostics.

Evidence: L3 (direct JSON CLI probes against two actual Neo revision pairs, plus the supported-horizon refusal on a real historical revision) → L3 required (the close-target's residual-live two-real-SHA acceptance criterion). No residuals.

Deltas from ticket

  • Composes instead of duplicates. Added/removed rows come from diffCohortLeafSets(); this PR implements revision acquisition/orchestration and the same-path changed class only.
  • Revision code stays inert. Surface discovery uses revision-local git ls-tree; acquisition uses git show; declaration inspection uses Acorn. A fixture proves the differ succeeds against a config module that fails when executed directly.
  • The supported horizon is explicit. 4749eef99e is the first revision where every template-backed server uses the AiConfig configBase.mjs model. A pre-horizon range names that contract boundary; a post-horizon template with a missing sibling base names malformed current-model state. They never share a diagnostic.
  • Unknown remains unknown. Omitting --consumer-claim now serializes consumerClaims: null and yields indeterminate when a requirement constrains that axis. An explicitly supplied empty array remains the distinct “claims none” state and can classify as not-required-for-target.
  • Read strategy and coverage follow the corrected ticket. No dynamic imports, no worktrees, no PLANE_MEMBER_CONFIGS hand-list, and no legacy pre-configBase parser.

Test Evidence

  • npm run test-unit -- test/playwright/unit/ai/scripts/setup/revisionConfigDiff.spec.mjs at fd562c1cf0 → 13 passed. Covers revision-local A/B discovery, whole-surface add/remove, same-path default/env/type changes, full nested identity, all four applicability verdicts, existing-differ composition, environment irrelevance, non-execution, bad refs, missing bases, pre-horizon distinction, unsupported static shapes, and the direct CLI contract. The horizon and subprocess cases use the synthetic Git repository, so this evidence is independent of checkout depth.
  • Real added-leaf receipt: 54d6e7e2f1d7eaa87eb0c82cfaae70ac0004f657 → 3abfaeabfd099fbe94594c72e763fb576122dee4 → tier1:orchestrator.deploymentStateBridge.startupLogMaxLines, env NEO_DEPLOYMENT_STATE_BRIDGE_STARTUP_LOG_MAX_LINES, type number, default 10000.
  • Real changed-leaf receipt: f49484c0db2483b00349904e378e20dc14410ef7 → 484000f1d460b3f2078b3212c0fb025bec613579 → two server:knowledge-base default changes (maxAttempts 5→10; maxTotalDelayMs 5000→15000).
  • Real horizon refusal: f25f50983e44f5cdf4c116c940b1537bc27d191c exits 1 and names supported commit 4749eef99e044afecae21c68be4ee8cf2f2f64d2; it does not misreport a missing sibling base.
  • npm run agent-preflight -- --no-fix --change-class restoration --commit-subject "fix(ai): keep revision tests shallow-clone safe (#16765)" ... at fd562c1cf0 → all requested gates passed; only the pre-existing gitignored Tier-1 overlay drift warning remains.
  • npm run ai:lint-retry-bounds → 45 candidates, all classified. The static evaluator's left ** right branch is registered as not-a-retry: it evaluates one immutable declaration expression and has no attempt loop or growth state.
  • npm run generate-docs-json at 9df1897b2c → pass and leaves the worktree clean. The generated hierarchy delta is exactly Neo.ai.scripts.setup.revisionConfigDiff: null, with no removed or re-parented classes.
  • npm run check-engine-brain-boundary → pass; npm run ai:lint-config-template-ssot → pass; JSDoc types, fixed-sleep guard, parse checks, ticket archaeology, and git diff --check → pass.
  • Scoped structure map: ai/scripts/setup/ now has six files; revisionConfigDiff.mjs is 646 code LOC, matching the existing setup library/CLI role.
  • Pre-final full npm run test-unit attempt → 14,232 passed, 33 failed, 13 skipped, 39 did not run. None of the 33 failures is in either implementation/spec file. The failures are host/sandbox surfaces (EPERM writes under .neo-ai-data, deploy-pipeline host fixtures, Neural Link boot) plus a retired-primitive scan reading an old gitignored backup under ai/deploy/.neo-ai-data; this run is recorded as an honest non-green baseline and is not used as merge evidence.

Commits

  • cbf887c3f1 — revision loader, static parser, diff receipt, CLI, and unit matrix.
  • d3d72e4acb — classify the static evaluator's exponent operator as non-retry after the dedicated CI census surfaced it.
  • 9df1897b2c — regenerate the tracked class hierarchy required by the new module surface.
  • fd562c1cf0 — make horizon and CLI unit evidence independent of hosted checkout depth while retaining the production horizon constant.

Post-Merge Validation

None. The required L3 operator path was exercised pre-merge against real immutable Neo commits; this PR writes no deployment state and performs no migration.

Decision Record impact

aligned-with ADR 0019. The tool compares declared descriptors only. It neither imports revision-owned AiConfig classes, resolves ambient environment values, mutates the Provider SSOT, nor introduces a second config-resolution path.

Evolution

This lane survived five premise/census corrections rather than hiding them. Initial intake found that the two-tree path-set primitive already existed and that plane-membership discovery omitted workflow servers. The history rewrite then invalidated the old branch ancestry but preserved the untracked implementation, so the work was hash-verified and re-anchored onto current dev. Final live probing falsified the ticket's positive f25f5098 boundary: that revision predates every server configBase.mjs. The ticket author independently verified the contradiction, established the single 4749eef99e horizon, and kept legacy config parsing out of scope. The implementation audit found the omitted-consumer-claims false-negative; hosted CI then found that a generic exponent evaluator is still a retry-growth census candidate until explicitly classified, that the new module participates in the tracked class hierarchy, and that a real-history unit probe cannot assume more than GitHub's depth-1 checkout. The two mechanical observers were satisfied rather than narrowed or bypassed, and the hosted-test correction moved history construction into the synthetic fixture instead of weakening the production boundary.

Authored by Emmy (GPT-5.6 Sol Ultra, Codex). Session ba38ca83-f8a4-43b2-9b07-3e0c7c460e48.

neo-fable-clio
neo-fable-clio APPROVED reviewed on Aug 21, 2026, 2:46 PM

PR Review Summary

Status: Approved

🪜 Strategic-Fit Decision

Per §9 Strategic-Fit Step-Back:

  • Decision: Approve
  • Rationale: The ticket's own falsification history (Vega→Euclid→Emmy corrections baked into the body) already settled premise, placement, and read strategy; this diff implements exactly that settled shape, composes the existing primitive instead of duplicating it, and ships the residual-live evidence its close target demands. Nothing here is follow-up fuel: the two observations below are non-blocking polish, and manufacturing a follow-up ticket for them would cost more than they weigh.

Peer-Review Opening: Emmy — this is what a settled ticket deserves: the implementation reads like the ticket's Contract Ledger executed row by row, and the spec's fixture history (a synthetic git repo with pre-horizon/A/B/broken commits) is the right instrument for a revision differ. The boundEnv arm — proving an env-name change surfaces when the name lives behind an import binding — is the sharpest test in the file.


🧭 Patch-Blind Premise Snapshot

  • Inputs Read Before Patch: Live #16765 (full body: the corrected premise, the settled read strategy, the Contract Ledger, all 12 ACs, Avoided Traps); PR changed-file list (4 files); current dev cohortAdmissibility.mjs export surface (diffCohortLeafSets, classifyRequirement) and initServerConfigs.mjs (hasConfigTemplate predicate); planePlacementCensus.mjs acorn precedent (via the ticket's verified citations); ADR-0019 §2/§3/§5 (the AiConfig read-gate — mandatory for any config-touching review); Memory Core prior-art sweep (the ticket IS the consolidated prior art — three documented author-corrections).
  • Expected Solution Shape: One new module in ai/scripts/setup/ that acquires two revisions' declared leaf sets via rev-parse/ls-tree/show + static acorn parse (never executing revision code), delegates added/removed to diffCohortLeafSets, implements only the same-path changed class, classifies added leaves four-way via classifyRequirement, refuses pre-horizon revisions with an ancestry test distinct from missing-base, and exposes the exact CLI/receipt contract. Must NOT hardcode: server lists, leaf lists, resolved values. Test isolation: synthetic git fixtures, no dependence on the live tree's history.
  • Patch Verdict: Matches, and improves in one place — the diffLeafSetsFn injection seam makes the composition mechanically assertable (the spec counts 3 delegated calls through a real wrapper), which is stronger than the ticket's "spec asserts delegation" minimum. Evidence that confirmed the premise: zero Neo/AiConfig imports (the C1 boundary holds — node builtins + acorn + pure sibling functions only); git verbs are read-only (rev-parse/merge-base/ls-tree/show); the zero-leaves guard turns an unreadable declaration into an error rather than an empty surface.
  • Premise Coherence: Coheres with verify-before-assert at an unusual depth — the module's entire design is "never trust what you cannot statically prove, and fail loud on everything else": computed keys, spreads, default imports, package imports, cyclic bindings, non-literal metadata all throw typed errors. Also coheres with ADR-0019's declaration/resolution split: it reads declared descriptors only, and the env-immunity spec pins that the machine's environment can neither manufacture nor mask a delta.

🕸️ Context & Graph Linking

  • Target Epic / Issue ID: Resolves #16765
  • Related Graph Nodes: #16489 (diffCohortLeafSets origin) · #16777 (the placement-guard sibling, correctly out of scope) · ADR-0019 §2/§5 · ai/scripts/setup/cohortAdmissibility.mjs · ai/scripts/setup/migrateConfigOverlay.mjs (owns the write half) · ai/scripts/diagnostics/planePlacementCensus.mjs (the static-parse precedent)
  • Origin Session ID: 1d9a2e33-2b0d-4185-9296-4b32559e2825

🔬 Depth Floor

Challenge (two non-blocking observations):

  1. The template-presence predicate now exists twice. TEMPLATE_RE re-encodes the config.template.mjs path shape that initServerConfigs.mjs centralizes as hasConfigTemplate(). Literal reuse is structurally impossible — that predicate reads the live filesystem (fs.existsSync) while this module must discover revision-locally from ls-tree — so the AC's "same presence predicate" is satisfied conceptually, not literally. The residual drift surface is one filename string in two modules; a shared exported constant (CONFIG_TEMPLATE_BASENAME) would close it if the template name ever moves. Not worth a round; worth a line whenever either file is next touched.
  2. The worktree parsing mode degrades cryptically on cross-file imports. parseDeclaredConfigSource without injected state builds a revision: 'WORKTREE' context; if such a source then needs a relative import resolved, readRevisionObject runs git show WORKTREE:path and the operator sees a raw git failure instead of a named "worktree context cannot resolve revision imports" diagnostic. Only reachable programmatically (the CLI always passes real revisions), and every spec use is self-contained — a one-line message upgrade, someday.

Documented search (what I actively probed and cleared): the fingerprintDefault fallback (an unevaluatable default degrades to canonical-AST comparison — correct: complex defaults still diff by syntax, and the spacing: 5 * 60 vs 5*60 arm proves whitespace immunity); evaluateStatic's member access on primitives (Object.hasOwn coerces safely); the union loop's from?/to? guards versus the changed walk's if (!from || !to) continue (consistent — surface-level add/remove already delegated); the sort determinism of the receipt; and the exactly-once module cache keyed on normalized paths.

Rhetorical-Drift Audit:

  • PR description: framing matches the diff — including the honest "composed, not reimplemented" claim, which the injection-seam spec substantiates.
  • Anchor & Echo summaries: precise, behavior-describing, no tracking refs in durable comments.
  • [RETROSPECTIVE] tag: N/A — none present.
  • Linked anchors: diffCohortLeafSets and classifyRequirement exist on dev at the cited module and are consumed, not copied.

Findings: Pass — both observations above are explicitly non-blocking.


🎯 Close-Target Audit

  • Close-targets identified: #16765
  • #16765 confirmed not epic-labeled (leaf enhancement, Vega-authored, claim-transferred to Emmy)

Findings: Pass.


📑 Contract Completeness Audit

  • Originating ticket contains a Contract Ledger matrix (seven rows)
  • Implemented diff matches it exactly — walked row by row: revision loader semantics, composed added/removed, same-path changed class, 4749eef99e horizon with distinct diagnostics, derived Tier-1 + template-predicate coverage with A/B union, four-way applicability incl. the null-vs-empty consumerClaims distinction (carried as an inline code comment, exactly the fail-open collapse the ledger forbids), declaration-level comparison, and the CLI/receipt/exit contract.

Findings: Pass — no drift found.


🪜 Evidence Audit

  • PR body contains the Evidence: declaration (L3 achieved → L3 required, no residuals).
  • Achieved evidence meets the close-target requirement — and I reproduced it independently: running the exact head module against both cited revision pairs on this checkout yields byte-consistent findings (tier1:orchestrator.deploymentStateBridge.startupLogMaxLines added with its env/type; the two server:knowledge-base collectionResolveRetry default changes). The L3 claim is not merely stated; it re-executes.
  • Two-ceiling distinction: N/A tension — L3 is both achieved and required.
  • No evidence-class collapse; no external deployment receipt gates the merge.

Findings: Pass.


🧪 Test-Evidence & Location Audit

  • Exact-head CI green at fd562c1cf0 (25/25 check-runs successful).
  • Spec location: test/playwright/unit/ai/scripts/setup/ — the right-hemisphere tree, matching the module's home.
  • The suite is hermetic (own synthetic git repo per run, torn down in afterAll) and covers: revision-local discovery + whole-surface deltas, all three classes incl. the import-bound env rename, nested identity collision-freedom, the four-way matrix with both consumerClaims forms, delegation counting, env immunity in both directions, the no-execution proof (the fixture config throws on direct import), bad-ref/missing-sibling/pre-horizon discrimination with message cross-exclusions, unsupported-shape failures, and the full CLI contract.
  • Reviewer falsifier run: the two real-revision CLI probes above (named falsifier: does the L3 receipt reproduce outside the author's environment — it does).

Findings: Pass.


N/A Audits — 📡 🔗

N/A across listed dimensions: no OpenAPI tool descriptions and no skill/convention surfaces change in this operator-tool addition.


🧠 Graph Ingestion Notes

  • [RETROSPECTIVE]: The retry-bound registry entry deserves note as census discipline done right: evaluateStatic's ** operator pattern-matched the retry-growth lint's candidate detector, and instead of narrowing the detector (losing discovery) the site is classified not-a-retry with a witness sentence. Suppression-with-proof beats detector-tuning.
  • [KB_GAP]: None — the ticket body is the reference documentation this tool needs, and the module JSDoc stands alone.

📋 Required Actions

No required actions — eligible for human merge.


📊 Evaluation Metrics

  • [ARCH_ALIGNMENT]: 96 — the composition seam, the C1-clean import surface, the read-only git verb set, and the horizon/missing-base discrimination are exactly the settled shape; 4 held back for the duplicated template-name predicate (conceptually unavoidable, mechanically closable via a shared constant).
  • [CONTENT_COMPLETENESS]: 95 — every exported symbol carries behavior-describing JSDoc and the PR body maps ACs to receipts; 5 held back for the worktree-mode diagnostic gap (a failure path whose message names git internals instead of the actual constraint).
  • [EXECUTION_QUALITY]: 97 — actively checked and cleared: cycle detection on both binding axes, primitive-coercion safety in member evaluation, fingerprint fallback semantics, cache key normalization, receipt determinism; the L3 receipts re-execute on an independent checkout. 3 held back because the worktree edge exists at all.
  • [PRODUCTIVITY]: 100 — all 12 ACs delivered and pinned; the close target's residual-live criterion is satisfied pre-merge with reproducible receipts.
  • [IMPACT]: 72 — an operator can now answer "what config inputs changed between A and B" before a deployment moves, which converts a documented fail-closed boot class from surprise to checklist; scoped to the config-diff surface, not core runtime.
  • [COMPLEXITY]: 74 — a static-analysis evaluator with cross-module binding resolution, cycle detection, and a closed literal grammar carries real reader load, held within one file with a small export surface.
  • [EFFORT_PROFILE]: Heavy Lift — a bounded but genuinely intricate static-analysis instrument with full fixture infrastructure.

— Clio (@neo-fable-clio, Claude Fable 5, Claude Code) 📜