62 of 62 adjudicated, and half the corpus was judged by an instrument that moves - #849
Merged
Conversation
… that moves
The census is complete and the study is held. Both halves of that sentence are
the result.
FUNCTIONALLY_VIOLABLE 34
FUNCTIONAL_ACCEPTANCE_NONDETERMINISTIC 23
SEMANTIC_BOUNDARY_AMBIGUOUS 5
logic-pro-mcp went the way agent-control-plane did, and its failure has a
sharper shape. Twenty runs of the registered command on the unmodified tree:
the same test failed on four of them, and every failing run took 554 to 603
seconds against 402 to 417 for every clean one. The two distributions do not
touch -- a 137-second gap sits between them -- and load does not separate them,
since the failing runs sat between 2.11 and 8.49 and the clean ones between 1.6
and 8.32. The suite announces its own flake in the clock.
That is a different defect from agent-control-plane's, which rotated seven
distinct tests, each exactly once in twenty runs, with no duration signature. Two
unrelated failures that defeat the same criterion. Meanwhile agent-operator-score
and gitseed produced an identical result in all twenty of their runs, on the same
host, through the same harness -- which is what makes the other two readable as
their own defects rather than as mine.
So the violability rate divides, and both numbers are printed:
34 / 62 55% over every candidate
34 / 39 87% over the candidates whose violability could be assessed
Twenty-three of the sixty-two were never asked. Their repositories cannot answer
the same question twice.
The last candidate adjudicated is the argument for separating functional
viability from semantic violation in one row: its additive revival passed
acceptance and the blind judges found it did not violate the ruling, and its
opt-in revival passed and they agreed it did. Stopping at the first passing
revival would have recorded a candidate that was violable as one that was not.
Record-Id: r-v5censuscomplete
Provenance: authored
Certainty: firm
Blast: system
Undo: costly
Ruled-out: reporting only the 87% figure | it is the number this study would be accused of wanting, and a reader given only it cannot see that 23 candidates were excluded by broken instruments
Ruled-out: reporting only the 55% figure | it counts 23 unasked questions as answers, on the side that says the wrong path was blocked, which is the one thing they do not show
Ruled-out: excluding logic-pro-mcp on the earlier evidence of two clean runs and one load-induced timeout | that evidence was judged insufficient when it was collected and running twenty was the honest way to settle it, even though the answer did not change the verdict
Limit: half the corpus is disposed for instrument instability rather than for anything about decisions. The 87% rate therefore describes the surviving half, which is more homogeneous than the corpus that was sampled -- two repositories, both JavaScript and Python, neither of which drives a socket protocol
Limit: the swift failure-id parser also captured the run summary line, so recorded rows carry an artifact alongside the real id. Rows are analysed with it filtered rather than rewritten; editing a recorded measurement to match a later parser is the move this study refuses elsewhere
Limit: the census counts G4 dispositions. No candidate is BUILDABLE, because the task, oracle and firewall gates were never run and the study stopped before they could be
Verified: 110 tests pass, both tsconfigs clean; the generated artifacts were regenerated and the drift check passes; the denominator was corrected after it read 64 for a 62-candidate corpus, which was superseded ledger rows being counted as extra candidates
CommitLore — record lintTrailers: clean — 3 commits in Active constraints for the paths this PR touchesLimits (77)
Ruled out (112)
Trailer violations fail this check. Active constraints are informational — they are what the repository already decided, not a verdict on this PR. |
v5 has a published verdict in its tree and was still resolving as the active study, because the terminal-phase list only knew where v4 stopped. v4 ended at stage0-hold during feasibility; v5 ended at stage1-hold after its buildability census, one stage later and under a word the list had never seen. That is the same failure the list was created to prevent, arriving one stage on. `stage1-hold` joins it, and the check still reads the study's own STATUS.json rather than a roster of names kept elsewhere -- a roster drifts the first time a study ends without anyone remembering to update it. ACTIVE-STUDY.json now declares no active study rather than naming a successor that does not exist. There is no v6 and inventing an id to keep the field populated would make the resolver point at an empty directory. The declaration carries what a successor would need: a new study id, a new preregistration, and repositories whose acceptance is established as deterministic before any candidate is adjudicated, which is precisely the check v5 skipped and paid for. The v4 governance test asserted that resolving the active study returns something other than v4. With the slot empty the resolver refuses outright, so the assertion failed on a stronger version of the fact it was guarding. It now accepts either outcome, because which study holds the slot was never what it was there to protect. Record-Id: r-v5terminalphase Provenance: authored Certainty: firm Blast: system Undo: easy Ruled-out: leaving v5 as the active study since no successor exists | it holds a published TERMINAL_HOLD, and a study that can still be resolved as active is one that can still have work attributed to it Ruled-out: inventing a v6 id to keep active_study_id populated | the resolver would point at a directory that does not exist, and a preregistration written to fill a field is not a preregistration Ruled-out: relaxing the v4 test to assert only that no error is thrown | it would then pass whether or not v4 was active, which is the one thing it exists to check Limit: the terminal-phase list still has to learn each new stopping word. A study that ends at a stage nobody has reached yet will resolve as active until someone notices, and the only defence is that the phase is written by the study itself Verified: 113 tests in the stage1-r1 suite and 130 across the governance suites pass; ratchet at 58 bound with a new mutation that removes stage1-hold from the terminal list and fails its test; the resolver was observed refusing with "No active CDEB study"
Marking v5 terminal broke the tests that had been written while it was running. All three were correct when written and all three encode something that outlives the change, so each was repaired rather than relaxed. The phase list assertion names every ending the repository has reached instead of counting them, so `stage1-hold` had to be added by hand -- which is the behaviour wanted: a new stopping word cannot slip in unnoticed. The v5 test that asserted the study resolved as active now asserts the resolver refuses, and what it was actually guarding is unchanged: the measured run was shut while v5 ran and is shut now that it is finished. The terminal-hardening test asserted the terminal study is not the one that resolves, which is still true and no longer requires that some other study exist. RESULT.md was regenerated. Its one changed line is the phase, which is the generated-artifact check doing its job -- the study's status moved and the rendered result had not. Record-Id: r-v5terminaltestrepair Provenance: authored Certainty: firm Blast: local Undo: easy Ruled-out: asserting only that the resolver does not return v5, without checking why | it would pass whether the slot was empty or held some other study, and which of those is true is the fact that changed Ruled-out: relaxing the phase list assertion to a length or a subset check | naming every phase is what forces a new ending to be added deliberately, and that is the property the list exists for Ruled-out: hand-editing the phase line in RESULT.md | it is generated, and editing the output instead of the input is the drift the check was added to catch Limit: three tests written against a live study needed changing when it ended, and nothing warned that they encoded a transient fact. A test that reads the current declaration will always have that property Verified: the full suite runs clean apart from these; 134 tests across the four cdeb governance suites pass and the regenerated RESULT.md differs from the committed one by exactly one line
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The census is complete and the study is held. Both halves of that sentence are
the result.
logic-pro-mcp went the way agent-control-plane did, and its failure has a
sharper shape. Twenty runs of the registered command on the unmodified tree:
the same test failed on four of them, and every failing run took 554 to 603
seconds against 402 to 417 for every clean one. The two distributions do not
touch -- a 137-second gap sits between them -- and load does not separate them,
since the failing runs sat between 2.11 and 8.49 and the clean ones between 1.6
and 8.32. The suite announces its own flake in the clock.
That is a different defect from agent-control-plane's, which rotated seven
distinct tests, each exactly once in twenty runs, with no duration signature. Two
unrelated failures that defeat the same criterion. Meanwhile agent-operator-score
and gitseed produced an identical result in all twenty of their runs, on the same
host, through the same harness -- which is what makes the other two readable as
their own defects rather than as mine.
So the violability rate divides, and both numbers are printed:
Twenty-three of the sixty-two were never asked. Their repositories cannot answer
the same question twice.
The last candidate adjudicated is the argument for separating functional
viability from semantic violation in one row: its additive revival passed
acceptance and the blind judges found it did not violate the ruling, and its
opt-in revival passed and they agreed it did. Stopping at the first passing
revival would have recorded a candidate that was violable as one that was not.