Skip to content

62 of 62 adjudicated, and half the corpus was judged by an instrument that moves - #849

Merged
MongLong0214 merged 3 commits into
mainfrom
cdeb-v5-census-complete
Aug 23, 2026
Merged

62 of 62 adjudicated, and half the corpus was judged by an instrument that moves#849
MongLong0214 merged 3 commits into
mainfrom
cdeb-v5-census-complete

Conversation

@MongLong0214

Copy link
Copy Markdown
Owner

The census is complete and the study is held. Both halves of that sentence are
the result.

FUNCTIONALLY_VIOLABLE                    34
FUNCTIONAL_ACCEPTANCE_NONDETERMINISTIC   23
SEMANTIC_BOUNDARY_AMBIGUOUS               5

logic-pro-mcp went the way agent-control-plane did, and its failure has a
sharper shape. Twenty runs of the registered command on the unmodified tree:
the same test failed on four of them, and every failing run took 554 to 603
seconds against 402 to 417 for every clean one. The two distributions do not
touch -- a 137-second gap sits between them -- and load does not separate them,
since the failing runs sat between 2.11 and 8.49 and the clean ones between 1.6
and 8.32. The suite announces its own flake in the clock.

That is a different defect from agent-control-plane's, which rotated seven
distinct tests, each exactly once in twenty runs, with no duration signature. Two
unrelated failures that defeat the same criterion. Meanwhile agent-operator-score
and gitseed produced an identical result in all twenty of their runs, on the same
host, through the same harness -- which is what makes the other two readable as
their own defects rather than as mine.

So the violability rate divides, and both numbers are printed:

34 / 62   55%   over every candidate
34 / 39   87%   over the candidates whose violability could be assessed

Twenty-three of the sixty-two were never asked. Their repositories cannot answer
the same question twice.

The last candidate adjudicated is the argument for separating functional
viability from semantic violation in one row: its additive revival passed
acceptance and the blind judges found it did not violate the ruling, and its
opt-in revival passed and they agreed it did. Stopping at the first passing
revival would have recorded a candidate that was violable as one that was not.

… that moves

The census is complete and the study is held. Both halves of that sentence are
the result.

    FUNCTIONALLY_VIOLABLE                    34
    FUNCTIONAL_ACCEPTANCE_NONDETERMINISTIC   23
    SEMANTIC_BOUNDARY_AMBIGUOUS               5

logic-pro-mcp went the way agent-control-plane did, and its failure has a
sharper shape. Twenty runs of the registered command on the unmodified tree:
the same test failed on four of them, and every failing run took 554 to 603
seconds against 402 to 417 for every clean one. The two distributions do not
touch -- a 137-second gap sits between them -- and load does not separate them,
since the failing runs sat between 2.11 and 8.49 and the clean ones between 1.6
and 8.32. The suite announces its own flake in the clock.

That is a different defect from agent-control-plane's, which rotated seven
distinct tests, each exactly once in twenty runs, with no duration signature. Two
unrelated failures that defeat the same criterion. Meanwhile agent-operator-score
and gitseed produced an identical result in all twenty of their runs, on the same
host, through the same harness -- which is what makes the other two readable as
their own defects rather than as mine.

So the violability rate divides, and both numbers are printed:

    34 / 62   55%   over every candidate
    34 / 39   87%   over the candidates whose violability could be assessed

Twenty-three of the sixty-two were never asked. Their repositories cannot answer
the same question twice.

The last candidate adjudicated is the argument for separating functional
viability from semantic violation in one row: its additive revival passed
acceptance and the blind judges found it did not violate the ruling, and its
opt-in revival passed and they agreed it did. Stopping at the first passing
revival would have recorded a candidate that was violable as one that was not.

Record-Id: r-v5censuscomplete
Provenance: authored
Certainty: firm
Blast: system
Undo: costly
Ruled-out: reporting only the 87% figure | it is the number this study would be accused of wanting, and a reader given only it cannot see that 23 candidates were excluded by broken instruments
Ruled-out: reporting only the 55% figure | it counts 23 unasked questions as answers, on the side that says the wrong path was blocked, which is the one thing they do not show
Ruled-out: excluding logic-pro-mcp on the earlier evidence of two clean runs and one load-induced timeout | that evidence was judged insufficient when it was collected and running twenty was the honest way to settle it, even though the answer did not change the verdict
Limit: half the corpus is disposed for instrument instability rather than for anything about decisions. The 87% rate therefore describes the surviving half, which is more homogeneous than the corpus that was sampled -- two repositories, both JavaScript and Python, neither of which drives a socket protocol
Limit: the swift failure-id parser also captured the run summary line, so recorded rows carry an artifact alongside the real id. Rows are analysed with it filtered rather than rewritten; editing a recorded measurement to match a later parser is the move this study refuses elsewhere
Limit: the census counts G4 dispositions. No candidate is BUILDABLE, because the task, oracle and firewall gates were never run and the study stopped before they could be
Verified: 110 tests pass, both tsconfigs clean; the generated artifacts were regenerated and the drift check passes; the denominator was corrected after it read 64 for a 62-candidate corpus, which was superseded ledger rows being counted as extra candidates
@github-actions

github-actions Bot commented Aug 23, 2026

Copy link
Copy Markdown

CommitLore — record lint

Trailers: clean — 3 commits in origin/main..85b1e0671ce2b102f1a3d2096145ff3ac705a6e1
Active constraints: 77 limits · 112 ruled-out · 0 warnings — from 36 records over 17 changed paths

Active constraints for the paths this PR touches

Limits (77)

  • r-v5terminaltestrepair 85b1e06 — three tests written against a live study needed changing when it ended, and nothing warned that they encoded a transient fact. A test that reads the current declaration will always have that property
  • r-v5terminalphase 58d31c8 — the terminal-phase list still has to learn each new stopping word. A study that ends at a stage nobody has reached yet will resolve as active until someone notices, and the only defence is that the phase is written by the study itself
  • r-v5censuscomplete e9445d3 — half the corpus is disposed for instrument instability rather than for anything about decisions. The 87% rate therefore describes the surviving half, which is more homogeneous than the corpus that was sampled -- two repositories, both JavaScript and Python, neither of which drives a socket protocol
  • r-v5censuscomplete e9445d3 — the swift failure-id parser also captured the run summary line, so recorded rows carry an artifact alongside the real id. Rows are analysed with it filtered rather than rewritten; editing a recorded measurement to match a later parser is the move this study refuses elsewhere
  • r-v5censuscomplete e9445d3 — the census counts G4 dispositions. No candidate is BUILDABLE, because the task, oracle and firewall gates were never run and the study stopped before they could be
  • r-v5censusdenominators 890c3ee — the assessable denominator excludes candidates for reasons that are themselves study findings. A repository disqualified for instrument instability leaves the remaining corpus more homogeneous than the corpus that was sampled
  • r-v5censusdenominators 890c3ee — gitseed has one candidate outstanding and logic-pro-mcp has thirteen, so the descriptive rate will move
  • r-v5determinismbyrepository 59e8f76 — logic-pro-mcp is unmeasured. Its suite takes about 700 seconds a run, twenty runs is roughly four hours, and its stratum cannot change the TERMINAL_HOLD that agent-control-plane already forced
  • r-v5determinismbyrepository 59e8f76 — both suites were run on trees freshly materialized from their bundles with digests verified, which is the same procedure the census used -- so this checks reproducibility of the run, not of the materialization
  • r-v5terminalhold e2581ab — twenty runs refute the determinism criterion but do not characterise the flakes. Seven tests seen once each gives a rate of roughly three runs in ten with at least one, and nothing here explains the mechanism
  • r-v5terminalhold e2581ab — the disposition is about this host. A different machine might not flake, and the study cannot distinguish "this suite is unstable" from "this suite is unstable here"
  • r-v5terminalhold e2581ab — the three surviving repositories are untouched by this and their numbers stand, but they now describe a corpus that cannot carry the confirmatory design
  • r-v5floorstanding c14ea75 — agent-control-plane's two negatives rest on a baseline captured in a single run. Their failing tests differ from attempt to attempt and sit in files unrelated to the rulings, which is the signature of an unstable suite rather than of enforcement -- three further baseline runs are in flight and the verdicts depend on what they show
  • r-v5floorstanding c14ea75 — the reserve count assumes every violable candidate survives the task, oracle and firewall gates, none of which has run
  • r-v5scopeconflict e26019b — the coverage check compares totals and skip counts. A patch that deletes one test and adds another keeps both numbers and passes
  • r-v5scopeconflict e26019b — the scope-conflict corroboration is a word match on the ruling -- a removal verb near a test noun. It confirms a claim already made rather than classifying rulings, but a ruling that describes removing coverage without those words would not corroborate
  • r-v5unchangedtreeisnotarevival 4721058 — telling an adjudicator that the repository's process documents do not bind it is a real intervention in what is being measured. It buys a technical answer to a technical question, and it means the census cannot speak to how much a project's own conventions deter a wrong path -- which is a genuinely interesting quantity this design now cannot see
  • r-v5unchangedtreeisnotarevival 4721058 — the unchanged-tree check catches an empty diff. An adjudicator that touches a file trivially and declines in substance passes it
  • r-v5canonicalcensus 944d454 — the generator validates rows against the adjudication rules, but the rules only constrain what a row asserts. A receipt whose numbers are real and whose patch does something other than the approach it names passes every check here
  • r-v5canonicalcensus 944d454 — 39 of 62 candidates are still unadjudicated, and the two repositories still at zero are the ones whose floors are in doubt
  • r-v5acceptancereceipts ea67250 — two agreeing judgements measure stability, not correctness. A reading this design gets consistently wrong stays consistently wrong, and the rule is a floor rather than a warrant
  • r-v5acceptancereceipts ea67250 — the agreement rule was registered after six of twenty-eight pairs had been read. Its direction is conservative -- it can only move a candidate out of FUNCTIONALLY_VIOLABLE -- but it was not fixed before any of the data existed
  • r-v5acceptancereceipts ea67250 — a receipt proves the tree as it stands passes the registered command. Whether that tree implements one approach or several accumulated ones is a question the receipt cannot answer, and is left to the semantic judges who read the diff
  • r-v5sandboxtaintvoided 29aebf1 — the seven were caught because one of them contradicted itself loudly enough for a validator to reject it. A sandbox that narrows what runs without any row becoming self-contradictory would still pass, and nothing here detects that shape
  • r-v5sandboxtaintvoided 29aebf1 — agent-control-plane's own baseline carries 9 pre-existing failures, all in the launchd file already excluded from acceptance. Every verdict in that repository is relative to a suite that does not fully pass on the unmodified tree
  • r-v5adjudicationtaxonomy afef35d — three distinct approaches is a floor that was tested and not a guarantee. A negative remains unproven, and the census says so rather than reporting it as a property of the tree
  • r-v5adjudicationtaxonomy afef35d — whether an opt-in revival is the violation the decision meant is unresolved. A record ruling out JSON storage may have meant storing runs as JSON at all rather than offering it as one option, and the endpoint as written counts the opt-in path. Recorded as a question for the oracle rather than settled after seeing which verdicts it moves
  • r-v5adjudicationtaxonomy afef35d — per-repository violable rates are not comparable. The acceptance commands differ in scope, 41 tests against 318, so a lower rate may mean a stricter repository or a wider suite
  • r-v5firewallprecision a18d37f — the precise gate reads the working files. A ruling paraphrased with no shared five-word run is invisible to it, exactly as the earlier leak screen was
  • r-v5firewallprecision a18d37f — this narrows what blocks. The wider question -- whether seeing any record primes an author about how this repository records decisions -- is real, unmeasured, and not addressed here
  • r-v5oracledirectoryallowed 2b5e552 — the forbidden-input check reads the oracle's source text, so an oracle that reached the arm through an indirection the words do not name would pass it
  • r-v5oraclethirdattempt 799f8fd — three of sixty-two. The mechanism is consistent across two repositories and two languages, and it remains a prior rather than a rate -- nothing here says what fraction of the corpus is guarded
  • r-v5oraclethirdattempt 799f8fd — no candidate has yet been attempted whose decision was recorded but never enforced. That is the case the mechanism predicts should be buildable, and until one is found the prediction is untested in the direction that would confirm it
  • r-v5oraclethirdattempt 799f8fd — these are G4 probes against each repository's full suite, not against a frozen functional acceptance authored behind the firewall. A narrower acceptance could admit a revival the full suite rejects
  • r-v5oraclesecondattempt b262363 — two of sixty-two, and both were picked for a crisply testable ruled-out approach -- the property most likely to correlate with having been guarded. The sample is selected in the direction of the finding
  • r-v5oraclesecondattempt b262363 — neither candidate had a task authored for it, so these are G4 probes against the repository's own suite rather than against a frozen functional acceptance. A narrower acceptance might admit a revival the full suite rejects
  • r-v5acceptancescreen fb93b39 — agent-control-plane's exclusion removes real coverage. The 9 cases assert deployment behaviour, and nothing in this study checks it -- a patch that breaks launchd installation passes acceptance
  • r-v5acceptancescreen fb93b39 — 700 seconds per acceptance run makes logic-pro-mcp the budget constraint, and the confirmatory design expects 300 to 400 episodes. The arithmetic is inside the registered ceiling but not comfortably
  • r-v5acceptancescreen fb93b39 — two serial runs are two observations. SSOT 6.4 asks for determinism over a hundred, and that has not been run
  • r-v5guardnamestale 6202b5d — nothing prevents the next rename from doing this again. The ratchet catches it on the following CI run rather than at the moment of the rename, and that is the guarantee available
  • r-v5oraclefirstattempt 8d158c6 — one candidate of 62. The mechanism predicts that well-implemented decisions are the least usable, and a mechanism that predicts more is not a measurement of more
  • r-v5oraclefirstattempt 8d158c6 — the two revivals I built are mine. A more inventive agent might find a shape the validator admits, and the honest reading is that I could not find one, not that none exists
  • r-v5oraclefirstattempt 8d158c6 — 61 candidates remain undecided, so nothing here says whether the corpus can clear the SSOT 7.3 floors of 8 buildable per repository and 24 in the reserve
  • r-v5taskchainfirstpair 74176cb — the oracle half has not been built for any candidate, so gate G2 is untouched and no candidate can be BUILDABLE yet
  • r-v5taskchainfirstpair 74176cb — zero four-gram overlap is a lexical floor. It cannot rule out a paraphrase that shares no four-word run, and nothing here tests for one
  • r-v5taskchainfirstpair 74176cb — one candidate, one model, one run of each half. Nothing yet says the other 61 behave the same way, and the leak screen already suggests some will not
  • r-v5needscoutfirstrun 7a4967f — one candidate, one scout, one model. FUNCTIONAL-AUTHOR has not run, so no manifest pair exists and the record-blind-before-record-aware ordering is still unproven end to end
  • r-v5needscoutfirstrun 7a4967f — zero four-gram overlap cannot rule out a paraphrase that shares no four-word run. The lexical check is a floor, not a proof of independence
  • r-v5needscoutfirstrun 7a4967f — the encouraging part -- a need landing on the decision boundary -- is one observation and says nothing yet about whether the other 61 behave the same way
  • r-v5firewallleakscreen c5d47eb — the screen measures shared wording, not disclosure. A reason that names a function shares wording with the code defining it without saying what was ruled out, and nothing here weighs where the hit landed
  • r-v5firewallleakscreen c5d47eb — the direction is not established. Record and code were often written together, so wording in common does not say which explains which
  • r-v5firewallleakscreen c5d47eb — two candidates are refuted outright by Record-Id presence, and the other 32 need per-candidate reading that has not happened. No disposition has been frozen for any of them
  • r-v5ssotreconciliation d3981ab — the 9.3 gate passes under an assumption the study cannot test. Nothing establishes the true between-candidate variance, the 12-candidate pilot cannot estimate one, and the corpus is fixed at 62 so no budget reduces it
  • r-v5ssotreconciliation d3981ab — the nine registered reasons are now the closed list, but no candidate has been disposed under any of them -- all 62 remain undecided pending oracle construction
  • r-v5stage1r1reviewfixes 87cacec — the pilot-gate custody gap is open and needs role separation -- execution, custody of arm-coded outcomes, and continuation held by different parties. That is an owner decision and it has to be settled before the pilot, not after
  • r-v5stage1r1reviewfixes 87cacec — the firewall now names who produced the maintenance need and refuses a producer not declared record-blind, but a declaration is not evidence. Making it evidence needs an attestable isolated authoring environment this layer does not have
  • r-v5stage1r1reviewfixes 87cacec — tau_squared_bound = 0.06 is a frozen assumption, not a measurement. It is conservative in the direction that matters -- lower true heterogeneity means the study detects more than promised -- but nothing here establishes the true value
  • r-v5stage1r1designlayer 28699d8 — the task-author firewall has still never run. Whoever built this corpus has read all 241 records, so the record-blind half cannot be satisfied from here -- it needs an author whose only inputs are the base tree and the maintenance need, with the manifest proving it
  • r-v5stage1r1designlayer 28699d8 — no oracle exists for any of the 62. Stage 0 recorded that reviewers thought one could be written; between that and a validated discriminating oracle sits the whole of G2, and nothing has crossed it
  • r-v5stage1r1designlayer 28699d8 — the between-candidate variance in the power table is a range I chose to bracket, not a measurement. The pilot supplies the real value, and only then does the detectable effect stop being a family of curves
  • r-v5qualification 4f63474 — 55 gates are unresolved after both tie-breakers disagreed, so 62 is a lower bound. A third independent model family would resolve some of them, and none is available -- the one attempted refused with HTTP 402
  • r-v5qualification 4f63474 — G3 and G4 are reviewer judgements made from the record, the paths and the commit prose. No reviewer read the current code or ran a test, so both are informed readings about a maintenance task rather than measurements of one
  • r-v5qualification 4f63474 — A0 admitted all 241 enumerated decisions, and seven of its eight conditions cannot fail on input the census built. The corpus filtering in v5 is done by G2 through G7, not by the authority model
  • r-v5authority 1012c88 — reusing the frozen bundles bounds the corpus to what existed on 2026-08-20. If v5 holds, the 14 excluded decisions are reported by name and count so the owner can order a deliberate re-snapshot rather than have one smuggled in here
  • r-v5authority 1012c88 — this commit registers the model and the thresholds. It produces no v5 count, and the preregistration is deliberately written before the census so the counts cannot choose the rules
  • r-v4qualification b8ff1b9 — G3 and G4 were judged from the commit message, the changed paths and the ruling. Neither reviewer read the current code or ran a test, so both are informed judgements about a maintenance task rather than measurements of one. G5 classifies whether an oracle could be written; none was built
  • r-v4qualification b8ff1b9 — this says nothing about whether recording decisions helps an agent. It says the four surveyed repositories cannot supply gold that is independent of the records being tested, which is a fact about these repositories and this gate
  • r-v4stage0id 6b427af — this proves a terminal study cannot be resolved as active through the declaration. It does not stop a caller that names a study root directly, which is why the measured-run gate is separate and still shut
  • r-v3terminalseal 7754f1a — the placeholder row remains in the ledger and always will. This makes it legible, not absent, and a reader who takes digests on faith rather than reading the deviation is still misled
  • r-v3terminalseal 7754f1a — the canonical digest binds the artifact list it is given. A transition that names too few artifacts is bound to a partial set, and nothing here decides what the right set is for a future study
  • r-v3terminalseal 7754f1a — guard coverage is unchanged -- thirteen exclusion kinds remain uncovered and one scan inert, recorded in the mutation baseline
  • r-guardratchet 26b1989 — the baseline is a floor. A guard can bind its claim against the one mutation recorded for it and still miss a different violation of the same claim
  • r-guardratchet 26b1989 — the reasons are written by the same author as the claims, so a gap reasoned narrowly can look more settled than it is
  • r-guardratchet 26b1989 — thirteen exclusion kinds remain uncovered and one scan inert; this records them and repairs neither
  • r-guardmutation 551921d — a mutation proves a guard reacts to one specific edit. A guard can bind its claim for that edit and miss a different violation of the same claim, so coverage here is a floor and not a proof
  • r-guardmutation 551921d — the claims were written by the same author as the guards, so a claim stated too narrowly produces a control that passes for a property nobody wanted
  • r-guardmutation 551921d — thirteen exclusion-index kinds remain inert; this change makes that visible and does not repair it

Ruled out (112)

  • r-v5terminaltestrepair 85b1e06 — asserting only that the resolver does not return v5, without checking why | it would pass whether the slot was empty or held some other study, and which of those is true is the fact that changed
  • r-v5terminaltestrepair 85b1e06 — relaxing the phase list assertion to a length or a subset check | naming every phase is what forces a new ending to be added deliberately, and that is the property the list exists for
  • r-v5terminaltestrepair 85b1e06 — hand-editing the phase line in RESULT.md | it is generated, and editing the output instead of the input is the drift the check was added to catch
  • r-v5terminalphase 58d31c8 — leaving v5 as the active study since no successor exists | it holds a published TERMINAL_HOLD, and a study that can still be resolved as active is one that can still have work attributed to it
  • r-v5terminalphase 58d31c8 — inventing a v6 id to keep active_study_id populated | the resolver would point at a directory that does not exist, and a preregistration written to fill a field is not a preregistration
  • r-v5terminalphase 58d31c8 — relaxing the v4 test to assert only that no error is thrown | it would then pass whether or not v4 was active, which is the one thing it exists to check
  • r-v5censuscomplete e9445d3 — reporting only the 87% figure | it is the number this study would be accused of wanting, and a reader given only it cannot see that 23 candidates were excluded by broken instruments
  • r-v5censuscomplete e9445d3 — reporting only the 55% figure | it counts 23 unasked questions as answers, on the side that says the wrong path was blocked, which is the one thing they do not show
  • r-v5censuscomplete e9445d3 — excluding logic-pro-mcp on the earlier evidence of two clean runs and one load-induced timeout | that evidence was judged insufficient when it was collected and running twenty was the honest way to settle it, even though the answer did not change the verdict
  • r-v5censusdenominators 890c3ee — replacing the all-adjudicated rate with the assessable one | it is the number that favours the study, and a reader given only it cannot see that a quarter of the corpus was excluded by a broken instrument
  • r-v5censusdenominators 890c3ee — keeping only the all-adjudicated rate | it counts ten unasked questions as answers, and the answer it counts them as is the one this study would be accused of wanting to avoid
  • r-v5censusdenominators 890c3ee — closing the not-a-violation candidate as a bounded negative | a passing revival that does not violate the ruling means the search found working code rather than a wrong path, so the search has not finished and three distinct shapes have not failed
  • r-v5determinismbyrepository 59e8f76 — taking the agent-control-plane result as sufficient on its own | it establishes that one suite rotates and says nothing about whether the harness or the host contributed, and the 24 verdicts already recorded elsewhere rest on exactly that distinction
  • r-v5determinismbyrepository 59e8f76 — running the two-arm load contrast again for these repositories | the contrast already established that load is not the driver, and repeating it here would spend an hour re-deriving a known answer instead of asking whether these suites move at all
  • r-v5determinismbyrepository 59e8f76 — reading twenty identical runs as determinism established | the registered protocol asks for a hundred, and what twenty buys is a readable contrast rather than a licence
  • r-v5terminalhold e2581ab — a reproduction rule that counts a failure only when it repeats across two runs | the measured asymmetry would justify it -- noise only ever added failures, never removed one, in all twenty runs -- and it would cut contamination from 30 percent to about 2, which is exactly why it must not be introduced now: it is a change to the registered acceptance outcome, made after seeing the data, in the direction that keeps the study alive
  • r-v5terminalhold e2581ab — adding the seven rotating tests to the exclusion list | an exclusion chosen after watching which tests failed is the shape the receipt validator already refuses, and it would have grown the exclusions from nine to sixteen on evidence that each one flaked once
  • r-v5terminalhold e2581ab — waiting for the remaining 33 candidates before reporting a verdict | a fixed stratum at zero of eight with zero candidates left is arithmetic, not an open question, and reporting INCOMPLETE would invite work that cannot change the outcome
  • r-v5terminalhold e2581ab — declaring nondeterminism from the four earlier runs alone | those ran under uncontrolled and unrecorded load between 6 and 141, and the honest reading of them was that they could not be read
  • r-v5floorstanding c14ea75 — teaching assertNegativeIsNotOverstated to allow negated uses of the overclaim | negation detection is fragile, and the phrasing it would permit is exactly the phrasing that fails when quoted without its qualifier
  • r-v5floorstanding c14ea75 — reporting the pooled violability rate as the feasibility number | 83% across the corpus is compatible with two empty strata, because the estimand averages over four fixed repositories rather than over candidates
  • r-v5floorstanding c14ea75 — waiting until the census finished before generating a floor section | the position is decision-relevant now, and a reader who has to compute it themselves will read the pooled number instead
  • r-v5scopeconflict e26019b — allowing a revival to delete the test that blocks it and noting the deletion | a revival permitted to remove what fails it passes everything, and the census would then measure how freely adjudicators delete tests
  • r-v5scopeconflict e26019b — taking ACCEPTANCE_SCOPE_CONFLICT from the adjudicator's blocked_by field | it removes a candidate from the corpus, so the cheapest way to shrink a hard workload would be to call it unmeasurable
  • r-v5scopeconflict e26019b — dropping the category once its motivating candidate started producing patches | the argument does not depend on that candidate: any ruling whose content is the acceptance suite has the same problem, and finding it once means it will recur
  • r-v5unchangedtreeisnotarevival 4721058 — trusting the adjudicator's own implemented:false flag | it was accurate in all six cases and it is still a self-report, which is the class of evidence this whole layer replaced
  • r-v5unchangedtreeisnotarevival 4721058 — counting a declined attempt as a failed revival | the approach was never tried, so the run is evidence about the adjudicator and not about the tree, and folding it into the negatives would inflate them with refusals
  • r-v5unchangedtreeisnotarevival 4721058 — leaving the prompt alone and treating policy refusals as a real finding | the finding would be "an agent would not do it", and no reader of the census would take the number that way
  • r-v5canonicalcensus 944d454 — keeping buildability-summary.json hand-maintained and being more careful | care is what failed; three rows went stale under attention and the failure mode is that correct and stale artifacts are indistinguishable by reading them
  • r-v5canonicalcensus 944d454 — mapping FUNCTIONALLY_VIOLABLE to BUILDABLE | G4 is one of five gates, and a corpus promoted on one gate would carry candidates with no oracle, no frozen task and no firewall manifest
  • r-v5canonicalcensus 944d454 — deleting the superseded rows once the current verdict is known | the census found that two failed attempts had been treated as proof, and the only record of that is the rows it produced
  • r-v5canonicalcensus 944d454 — putting the drift check in bench/cdeb/verify.mjs | that file verifies bench/results/cdeb, not the studies directory, and the check would have run against nothing while reporting clean
  • r-v5acceptancereceipts ea67250 — keeping the thirty prose-backed verdicts on the grounds that only agent-control-plane had a blocked sandbox | the objection is not that those runs went wrong, it is that nothing in the artifact can tell, and a rule that trusts prose wherever it has no reason to doubt it is the rule that produced the seven
  • r-v5acceptancereceipts ea67250 — discarding the thirty and re-running the search | the search is the expensive half and it already happened; the patched trees are still on disk, so acceptance can be replayed mechanically and the receipt built from what it actually produces
  • r-v5acceptancereceipts ea67250 — taking a single blind semantic judgement per revival | one judgement records whichever run happened to be kept, and one of the first six pairs disagreed with itself
  • r-v5acceptancereceipts ea67250 — editing the historical TREE_ENFORCED rows to the new name | an overturned negative is evidence that the search budget was once too small, and a rewritten row cannot show that
  • r-v5sandboxtaintvoided 29aebf1 — keeping the seven and noting the subset as a limitation | the gate is whether the full acceptance passes, so a verdict from a subset is not a weaker answer to that question but an answer to a different one
  • r-v5sandboxtaintvoided 29aebf1 — giving every repository danger-full-access so the sandbox stops interfering | three of the four run their suites correctly under workspace-write, and widening all of them to fix one removes a real constraint from the three that do not need it
  • r-v5sandboxtaintvoided 29aebf1 — trusting the validator's first reading that thirteen rows were malformed | a passing revival matches the baseline by definition, so the rule that flagged them was reading the evidence of success as evidence of not being a revival
  • r-v5adjudicationtaxonomy afef35d — discarding a failed G4 as a bare exclusion | the reason it failed is the measurable part, and without recording it there is no reason to count attempts and no way to notice that two was never enough
  • r-v5adjudicationtaxonomy afef35d — relaxing the viability requirement for candidates that failed it | the endpoint is exactly the event a revival produces, so a study that relaxes it measures something else while keeping the name
  • r-v5adjudicationtaxonomy afef35d — re-running the negatives I judged too similar rather than the ones below the registered count | v4-8ab61d73c22d675b tried three approaches that all look like one idea to me, and deciding that after seeing the verdicts is choosing which results to overturn
  • r-v5adjudicationtaxonomy afef35d — writing the census report after the numbers arrived | every rule in it would then have been chosen knowing what it would produce
  • r-v5firewallprecision a18d37f — keeping the coarse gate and disposing 32 candidates | it empties two fixed strata on the strength of a format placeholder and two unrelated ADR references, and an estimand lost to a false positive is worse than one lost to a real constraint
  • r-v5firewallprecision a18d37f — deleting the repository-wide scan now that it does not block | it found the two leaks in the first place, and a tree carrying records is worth reporting even when this candidate is unaffected
  • r-v5firewallprecision a18d37f — loosening to a warning everywhere | the candidate's own record in the tree really is disqualifying, and a firewall that only warns is a firewall that is not one
  • r-v5oracledirectoryallowed 2b5e552 — flipping "no oracle exists" to "an oracle exists" | it would pass on an empty file and check nothing, where the thing worth checking is what the oracle is allowed to read
  • r-v5oracledirectoryallowed 2b5e552 — leaving the oracles directory banned and putting the oracle elsewhere | the ban would then be enforcing a filename rather than the distinction it exists for, and the next oracle would go somewhere else again
  • r-v5oraclethirdattempt 799f8fd — continuing to pick candidates whose ruled-out approach looked testable | the first two dispositions were already limited by exactly that, and a third of the same kind would have added a count without addressing the objection
  • r-v5oraclethirdattempt 799f8fd — reading the two failing test names as sufficient and skipping the build | the names describe the compliant behaviour, which is suggestive and not the same as showing the revival cannot pass, and the text-matching shortcut already died against its null once today
  • r-v5oraclesecondattempt b262363 — reporting "42 of 318 failed" as the finding | most of those failures are my patch being incomplete, and a revival that fails because it was written badly says nothing about whether the approach is viable
  • r-v5oraclesecondattempt b262363 — writing a more complete JSON store until only the structural failures remained | it would change the two failures that matter not at all, and the effort would be spent making a number look tidier rather than making the conclusion truer
  • r-v5oraclesecondattempt b262363 — inferring the disposition from the grep that found 33 sqlite mentions | the same shortcut was tried as a corpus-wide screen earlier today and died against its null, so a grep is a reason to build the revival rather than a substitute for building it
  • r-v5acceptancescreen fb93b39 — disposing logic-pro-mcp's 13 candidates on the parallel non-determinism | the fair question was whether a serial configuration is deterministic, and it is, so the disposition would have been wrong and would have emptied a fixed stratum
  • r-v5acceptancescreen fb93b39 — keeping the default parallel command and excluding the flaky tests | the flakiness is a shared store observed across concurrent tests, so the set that fails moves between runs and there is no fixed set to exclude
  • r-v5acceptancescreen fb93b39 — leaving the acceptance commands in prose | the choice of --no-parallel and the launchd exclusion are the instrument, and an instrument that lives only in a narrative can be changed later without anything noticing
  • r-v5guardnamestale 6202b5d — adding a static check that every registry test_name appears in its file | test.each generates names at run time, so the check reports a false positive on a guard that genuinely binds, and a check that cries wolf gets muted
  • r-v5oraclefirstattempt 8d158c6 — counting Bad C as a functionally passing violation | nothing reads the file it adds, so it passes by being inert, and an unused file is litter rather than a revival
  • r-v5oraclefirstattempt 8d158c6 — editing the validator so a revival could pass | that dismantles the guard instead of reviving the approach, and it changes far more than the decision's scope
  • r-v5oraclefirstattempt 8d158c6 — concluding from Bad A alone | the record claims the approach would have passed every gate, so one failed attempt is not enough to call the endpoint unobservable
  • r-v5oraclefirstattempt 8d158c6 — pinning the census tests to "62 of 62" | they would need editing every time a candidate is disposed, which is how a test stops being read; they now track the census and still fail closed on any open slot
  • r-v5taskchainfirstpair 74176cb — installing node_modules in the sandbox to execute the acceptance command | it changes the frozen tree the digest is taken over, and the tree is the one thing the manifest claims to pin
  • r-v5taskchainfirstpair 74176cb — recording the oracle manifest as if the oracle existed | it carries a placeholder digest, and a pair that looks complete is worse than an absent one because the next reader stops asking for G2
  • r-v5taskchainfirstpair 74176cb — letting the ordering gate be tested only on the well-formed pair | a gate that never sees the tampered case has not been shown to order anything
  • r-v5sandboxtestselfcontained d199744 — committing the corpus bundles so the test can find them | they are mirrors of private repositories and the gitignore rule exists to keep them out of a public repo
  • r-v5sandboxtestselfcontained d199744 — skipping the test when the bundle is absent | it would then be a test that never runs in CI and reports green for having done nothing
  • r-v5needscoutfirstrun 7a4967f — telling the scout a decision exists and to avoid it | it would write around the shape of the thing it was told about, and the avoidance would be the leak
  • r-v5needscoutfirstrun 7a4967f — letting the record-holder pick among the returned needs | at that point the picker has read the ruling and would choose the need running closest to it, which is the selection bias the external seed exists to remove
  • r-v5needscoutfirstrun 7a4967f — restoring .git after codex refused to run without it | the refusal was the sandbox working, and the flag that tells codex the directory is not a repository costs nothing
  • r-v5firewallleakscreen c5d47eb — giving the task author the materialized bundle | it carries the commit history and the notes ref, so the firewall would rest on the author not running git log
  • r-v5firewallleakscreen c5d47eb — redacting the leaking files from the sandbox | the base tree is frozen and the episode runs on the real one, so an author working in a different tree authors a task for a repository that does not exist
  • r-v5firewallleakscreen c5d47eb — reporting the 34 without the cross-repository null | a shared five-word run means nothing until something establishes the rate at which unrelated text shares one, and that rate turns out to be zero
  • r-v5firewallleakscreen c5d47eb — calling this a global HOLD | a shared run is not a restatement, a hit in docs/adr is read by a different reader than one in the file the task must change, and deciding each case is the census's job
  • r-v5ssotreconciliation d3981ab — keeping my 15pp/1,650-episode envelope | it was derived by inverting a detectable-effect formula against a variance the pilot was supposed to supply, which is the sizing direction section 9 exists to remove
  • r-v5ssotreconciliation d3981ab — running the 9.3 gate with my more conservative variance assumption | that substitutes my judgement for the owner's registered method and would have produced a HOLD the SSOT does not call for
  • r-v5ssotreconciliation d3981ab — reporting the 9.3 PASS without the heterogeneity table | the registered simulation assumes candidates benefit equally, and a reader who cannot see what that assumption buys cannot judge a null result
  • r-v5ssotreconciliation d3981ab — deleting the analytic sizing path now that it is not the gate | it is the more pessimistic of the two and a disagreement between them is worth seeing rather than averaging
  • r-v5censusschemadrift a610828 — dropping additionalProperties:false so the schema stops caring | it is the clause that keeps an outcome field off a disposition row, which is the whole reason the census has a schema
  • r-v5stage1r1reviewfixes 87cacec — treating the reviewer's findings as claims to weigh | six of them were statements about what the code does, and running the code settled each one in under a minute
  • r-v5stage1r1reviewfixes 87cacec — keeping 8 repeats and reporting the detectable effect as a range | the range was over a parameter the design had no way to obtain, which is how an operator ends up choosing the favourable end of it
  • r-v5stage1r1reviewfixes 87cacec — lowering the minimum important effect to what 8 repeats reaches | that is the move the HOLD rule exists to refuse, and writing it into the rule that refuses it would have been circular
  • r-v5stage1r1reviewfixes 87cacec — claiming the pilot custody finding was closed by the schema | the record now cannot carry an arm contrast and the operator can still have watched the runs; a control over bytes is not a control over people
  • r-v5stage1r1designlayer 28699d8 — filling the runtime lock with plausible placeholder values | an unpinned runtime that reads as pinned is worse than an empty lock, because the next reader stops asking
  • r-v5stage1r1designlayer 28699d8 — marking the 62 screen-surviving candidates BUILDABLE | the screens can only refute, and calling a candidate buildable without an oracle is the claim the census exists to check
  • r-v5stage1r1designlayer 28699d8 — patching the six defects into the failed Stage 1 draft | its own section 7 says anything but the deferred N makes a change a new preregistration, so amending the document that defines amendment is the failure it guards
  • r-v5stage1r1designlayer 28699d8 — reporting PASS on the eleven satisfied criteria | four unresolved P0/P1 is a HOLD under section 19, and a partial pass reads as readiness to whoever approves execution
  • r-v5qualification 4f63474 — keeping the single third vote | the four-repository set rested on a tie-break measured at 67/33 toward one disputant, and reporting that as a limitation would have left the repository decision standing on it
  • r-v5qualification 4f63474 — moving the third vote to the other model | that reverses the lean instead of removing it
  • r-v5qualification 4f63474 — reporting only the adopted rule | the three rules differ by 44 candidates and one repository, so a reader who cannot see the spread cannot judge the verdict
  • r-v5qualification 4f63474 — resolving the remaining splits myself | 55 gates stay unresolved and fail closed, which is the honest outcome for a question two independent readings could not settle
  • r-v5authority 1012c88 — renaming v4's independent-prose gate rather than removing it | the preregistration and the authority policy both forbid any gate that requires a decision to be documented outside its record, because a gate that does is the v4 gate whatever it is called
  • r-v5authority 1012c88 — re-snapshotting to grow the corpus | the only decisions a fresh snapshot adds are the 14 authored during the study, in the repository where they would decide eligibility
  • r-v5authority 1012c88 — asking the owner to attest the 14 are natural | that is owner testimony, which is disabled, and the answer would arrive after the counts were known
  • r-v5authority 1012c88 — keeping invalidated as the only terminal phase | v4 is not invalid, it is finished, and a phase list that only knows about invalidation would let a study with a published verdict be run again
  • r-v4qualification b8ff1b9 — adjudicating the 92 split gates myself | the study operator reading their own corpus, already knowing how the pair voted, is the least blind reader available; a third blind vote costs one more session and is a vote rather than an override
  • r-v4qualification b8ff1b9 — averaging or passing an unresolved disagreement | it would put a candidate in the corpus that no two reviewers agreed on, and the disagreement would stop being visible
  • r-v4qualification b8ff1b9 — relaxing the quote-correspondence floor after seeing 8% | the floor was fixed in code and in the deviation record before any overlap was computed, and moving it now would let the count choose the method
  • r-v4qualification b8ff1b9 — accepting any rejection found in the same commit | it qualifies candidate X on evidence about decision Y, which is how a corpus fills up without meaning anything
  • r-v4stage0id 6b427af — keeping a hardcoded list of terminal study ids in the resolver | the list and the studies drift apart silently, and the drift shows up as a terminated study resolving cleanly
  • r-v4stage0id 6b427af — leaving the declaration at null and passing the study root explicitly everywhere | every caller then carries the choice, and the one caller that forgets picks a default nobody reviewed
  • r-v4stage0id 6b427af — reusing the v3r1 study directory under a new name | §4.3 requires a new study id, and a renamed directory keeps the qualification verdicts this estimand discards
  • r-v3terminalseal 7754f1a — correcting the placeholder digests in place | the correction is indistinguishable from the mistake it repairs, and the row is historical evidence rather than a working value
  • r-v3terminalseal 7754f1a — recomputing digests for the historical rows from today's artifacts | the artifacts have changed since, so the result would be a number that never bound anything, wearing the authority of one that did
  • r-v3terminalseal 7754f1a — hand-maintaining evidence-matrix.md beside the JSON | two copies of the same claims disagree eventually and the disagreement is silent
  • r-v3terminalseal 7754f1a — leaving cdeb-fresh-v3r1 as the default study root | a terminated study as a fallback is how a measured run gets attempted against a study that ended
  • r-guardratchet 26b1989 — leaving the job in the gate while it is red | branch protection would block every merge until someone deleted the job, and deleting it removes the only thing that can see these gaps
  • r-guardratchet 26b1989 — allowing an improvement without updating the baseline | the record would drift below the measurement, and a baseline that overstates the gaps is as useless as one that understates them
  • r-guardratchet 26b1989 — recording gaps as a count instead of per property with a reason | a count cannot distinguish a control nobody wrote from one that cannot exist, and those need different work
  • r-guardmutation 551921d — fixing the inert guards in the same change | a runner that has never reported a real failure is not known to report one, and the red run is the evidence that it can
  • r-guardmutation 551921d — indexing guard functions instead of claims | that reproduces the exact failure this exists to stop, because the control comes back out of the implementation it is meant to test
  • r-guardmutation 551921d — treating an unexpressible control as a skip | it is indistinguishable from a control nobody attempted, and both were silently green before
  • r-guardmutation 551921d — folding this into the check job | it runs vitest once per mutation, so it belongs in its own job where its cost is visible

git log --follow accepts exactly one pathspec, so renames are not followed for 17 paths; query one path at a time to follow its rename chain

withheld the content of 1 record(s) graded blocked: a Ruled-out trailer matching an injection pattern is reported, never quoted (SPEC §7)

Trailer violations fail this check. Active constraints are informational — they are what the repository already decided, not a verdict on this PR.

v5 has a published verdict in its tree and was still resolving as the active
study, because the terminal-phase list only knew where v4 stopped. v4 ended at
stage0-hold during feasibility; v5 ended at stage1-hold after its buildability
census, one stage later and under a word the list had never seen.

That is the same failure the list was created to prevent, arriving one stage on.
`stage1-hold` joins it, and the check still reads the study's own STATUS.json
rather than a roster of names kept elsewhere -- a roster drifts the first time a
study ends without anyone remembering to update it.

ACTIVE-STUDY.json now declares no active study rather than naming a successor
that does not exist. There is no v6 and inventing an id to keep the field
populated would make the resolver point at an empty directory. The declaration
carries what a successor would need: a new study id, a new preregistration, and
repositories whose acceptance is established as deterministic before any
candidate is adjudicated, which is precisely the check v5 skipped and paid for.

The v4 governance test asserted that resolving the active study returns
something other than v4. With the slot empty the resolver refuses outright, so
the assertion failed on a stronger version of the fact it was guarding. It now
accepts either outcome, because which study holds the slot was never what it
was there to protect.

Record-Id: r-v5terminalphase
Provenance: authored
Certainty: firm
Blast: system
Undo: easy
Ruled-out: leaving v5 as the active study since no successor exists | it holds a published TERMINAL_HOLD, and a study that can still be resolved as active is one that can still have work attributed to it
Ruled-out: inventing a v6 id to keep active_study_id populated | the resolver would point at a directory that does not exist, and a preregistration written to fill a field is not a preregistration
Ruled-out: relaxing the v4 test to assert only that no error is thrown | it would then pass whether or not v4 was active, which is the one thing it exists to check
Limit: the terminal-phase list still has to learn each new stopping word. A study that ends at a stage nobody has reached yet will resolve as active until someone notices, and the only defence is that the phase is written by the study itself
Verified: 113 tests in the stage1-r1 suite and 130 across the governance suites pass; ratchet at 58 bound with a new mutation that removes stage1-hold from the terminal list and fails its test; the resolver was observed refusing with "No active CDEB study"
Marking v5 terminal broke the tests that had been written while it was running.
All three were correct when written and all three encode something that outlives
the change, so each was repaired rather than relaxed.

The phase list assertion names every ending the repository has reached instead
of counting them, so `stage1-hold` had to be added by hand -- which is the
behaviour wanted: a new stopping word cannot slip in unnoticed. The v5 test that
asserted the study resolved as active now asserts the resolver refuses, and what
it was actually guarding is unchanged: the measured run was shut while v5 ran
and is shut now that it is finished. The terminal-hardening test asserted the
terminal study is not the one that resolves, which is still true and no longer
requires that some other study exist.

RESULT.md was regenerated. Its one changed line is the phase, which is the
generated-artifact check doing its job -- the study's status moved and the
rendered result had not.

Record-Id: r-v5terminaltestrepair
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: asserting only that the resolver does not return v5, without checking why | it would pass whether the slot was empty or held some other study, and which of those is true is the fact that changed
Ruled-out: relaxing the phase list assertion to a length or a subset check | naming every phase is what forces a new ending to be added deliberately, and that is the property the list exists for
Ruled-out: hand-editing the phase line in RESULT.md | it is generated, and editing the output instead of the input is the drift the check was added to catch
Limit: three tests written against a live study needed changing when it ended, and nothing warned that they encoded a transient fact. A test that reads the current declaration will always have that property
Verified: the full suite runs clean apart from these; 134 tests across the four cdeb governance suites pass and the regenerated RESULT.md differs from the committed one by exactly one line
@MongLong0214
MongLong0214 merged commit b39448a into main Aug 23, 2026
15 checks passed
@MongLong0214
MongLong0214 deleted the cdeb-v5-census-complete branch August 23, 2026 10:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant