Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
251f26b
One arm would have suppressed nothing and reported that it had
MongLong0214 Aug 28, 2026
3c5b85d
Both repositories use CommitLore, so scanning the whole tree marks ev…
MongLong0214 Aug 28, 2026
451d0a5
Thirteen of twenty-five, and the other twelve fail rather than default
MongLong0214 Aug 28, 2026
5e24883
STATUS names what PR-B now has and what it still gates on
MongLong0214 Aug 28, 2026
98a72f3
The gate had two conditions on a quantity the specification never def…
MongLong0214 Aug 28, 2026
8a80b90
The numbers alone cannot show the second analysis was independent
MongLong0214 Aug 28, 2026
7e114da
Half the study would have scored zero, and the other half would have …
MongLong0214 Aug 28, 2026
7f03249
"The judge saw the final tree" is weaker than it sounds
MongLong0214 Aug 28, 2026
6bab2bb
Exit 127 is command-not-found, and four tasks were scored as if it me…
MongLong0214 Aug 28, 2026
21fe7e0
A 401 and a model that produced a bad tree both exit non-zero
MongLong0214 Aug 28, 2026
8d88f47
Installing the test that scores the episode changed what the episode …
MongLong0214 Aug 28, 2026
19ebcb4
Thirteen of seventeen tasks can carry a functionally passing violation
MongLong0214 Aug 28, 2026
c7a1e6a
The same treatment is one record in three, or one in sixty, depending…
MongLong0214 Aug 28, 2026
bbb2704
The runner had never once been run
MongLong0214 Aug 28, 2026
6bcbc52
The judge answered correctly and told us it had judged "work.judge-1"
MongLong0214 Aug 28, 2026
65871c1
Checking that the runner starts is the same act as starting it
MongLong0214 Aug 28, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 24 additions & 2 deletions bench/cdeb/studies/cdeb-fresh-v8/PRD.md
Original file line number Diff line number Diff line change
Expand Up @@ -1423,6 +1423,21 @@ manual-discovery difference
token/wall-time difference
```

RBDR is defined as follows (owner ruling 2026-08-28, v8-d012). The specification
named it and section 27 gated it twice without ever defining it; an independent
analyst reading only this plan returned null.

```text
RBDR = |{pairs : SUPPRESSED functionally revived AND ON did not}|
--------------------------------------------------------
|{pairs : SUPPRESSED functionally revived}|
```

Pairs whose SUPPRESSED arm did not revive have nothing to block and are not in the
denominator. The lower bound is the 2.5th percentile of a bootstrap over the
revived pairs. When no suppressed arm revived, RBDR is undefined and the section 27
conditions on it fail rather than defaulting.

### 23.7 Judge sensitivity

각 judge를 단독 instrument로 사용한:
Expand Down Expand Up @@ -1476,12 +1491,19 @@ Match:
raw counts exact
panel labels exact
point estimates <= 1e-12
bootstrap quantiles <= 1e-6
permutation p <= 1e-6
reliability metrics <= 1e-6
claim gate identical
```

Bootstrap quantiles and the permutation p are compared but not required to match to
a fixed tolerance (owner ruling 2026-08-28, v8-d013). They are Monte Carlo
estimates, and two independent implementations consume the random stream in
different orders from the same seed, so agreement to 1e-6 would mean the two
analysts wrote the same code -- the opposite of what this section asks for. The
comparison reports the gap against Monte Carlo error at the registered replicate
count. Measured on a synthetic seal, two independent implementations differed by
2.8e-3 on the interval and 1.5e-3 on the p.

Mismatch unresolved:

```text
Expand Down
8 changes: 8 additions & 0 deletions bench/cdeb/studies/cdeb-fresh-v8/STATUS.json
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
"calibration_candidates_completed": 4,
"calibration_candidates_required_minimum": 3,
"calibration_complete": true,
"dry_run_manipulation": "17/17",
"evidence_tier_bounded_by": "calibration label/origin confound, section 26; see v8-d010 for what it does and does not bound",
"headline_scope_in_sentence": true,
"judge_independence_enforced": true,
Expand All @@ -20,6 +21,13 @@
],
"panel_frozen": true,
"phase": "schedule-frozen",
"prb_prerequisites": {
"batch_runner": "batch.py, one worker per repository, refuses while measured_run_allowed is false",
"episode_runner": "run-episode.py, section 19's sixteen steps",
"gate_input_builder": "gate-inputs.py, 13 of 25 derived, the rest fail rather than default",
"judge_packet_builder": "episode_packet.py, built before teardown",
"suppression_identity": "all 17 resolved, dry-run verified against real trees"
},
"product_effect_rows": 0,
"red_team_p0_open": 0,
"red_team_p1_open": 0,
Expand Down
213 changes: 213 additions & 0 deletions bench/cdeb/studies/cdeb-fresh-v8/acceptance-commands.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,213 @@
{
"all_verified": true,
"commands": [
{
"candidate_id": "v4-002ffd1e428c572a",
"command": "node --test packages/schema/test/capability-evidence-locator-allowlist.acceptance.test.ts",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "node --test packages/schema/test/capability-evidence-locator-allowlist.acceptance.test.ts",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-0ecd7426eebc1cab",
"command": "python3 -m pytest -q tests/test_custom_evidence_reader_acceptance.py",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [
"interpreter -> python3 -m pytest"
],
"recorded_how_to_run": "pytest -q tests/test_custom_evidence_reader_acceptance.py",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-34aef026d81c2f6b",
"command": "node --test tests/epic-dependency-normalization.acceptance.test.mjs",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "node --test tests/epic-dependency-normalization.acceptance.test.mjs",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-377f04276465b59d",
"command": "python3 -m pytest tests/test_ci_action_pinning.py -q",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest tests/test_ci_action_pinning.py -q",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-77e1745655a235ce",
"command": "python3 -m pytest tests/test_category_manifest_evidence.py -q",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest tests/test_category_manifest_evidence.py -q",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-84cd6d391ac2fa6d",
"command": "python3 -m pytest tests/test_correction_point_lookup_acceptance.py -q",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest tests/test_correction_point_lookup_acceptance.py -q",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-8f24735524874167",
"command": "node --test tests/acceptance/schema-doctor-lane.test.mjs",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "node --test tests/acceptance/schema-doctor-lane.test.mjs",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-8fc3d2ec14b1c078",
"command": "python3 -m pytest tests/test_collect_paging_validation_acceptance.py -q",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest tests/test_collect_paging_validation_acceptance.py -q",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-9b42b1951da730e1",
"command": "npm test -w @aos/schema -- registry-level-contract-fields",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "npm test -w @aos/schema -- registry-level-contract-fields",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-c61d7c943edd8cff",
"command": "node --test packages/schema/test/capability-derivation-proof.acceptance.test.ts",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [
"took the backticked command out of its prose wrapper"
],
"recorded_how_to_run": "From the repository root: `node --test packages/schema/test/capability-derivation-proof.acceptance.test.ts`",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-cadfb63755c3f504",
"command": "python3 -m pytest -q tests/test_pipeline_collection_rate_limit.py",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [
"interpreter -> python3"
],
"recorded_how_to_run": "python -m pytest -q tests/test_pipeline_collection_rate_limit.py",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-ce2adee3c134ab03",
"command": "node --test packages/schema/test/capability-validation-result.acceptance.test.ts",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "node --test packages/schema/test/capability-validation-result.acceptance.test.ts",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-dd4a74ba2b628991",
"command": "npm test -w @aos/schema -- metric-registry-envelope",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "npm test -w @aos/schema -- metric-registry-envelope",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-e7587b2b65750306",
"command": "node --test packages/schema/test/metric-definition.public-contract.test.mjs",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "node --test packages/schema/test/metric-definition.public-contract.test.mjs",
"recorded_verified_fails_on_base": true,
"repository_id": "agent-operator-score",
"runnable": true
},
{
"candidate_id": "v4-ed878960135ff45a",
"command": "python3 -m pytest -q tests/test_observation_ordering_acceptance.py",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [
"interpreter -> python3 -m pytest"
],
"recorded_how_to_run": "pytest -q tests/test_observation_ordering_acceptance.py",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-f3c960a48273132c",
"command": "python3 -m pytest -q tests/test_evidence_reader_fallback.py",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest -q tests/test_evidence_reader_fallback.py",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
},
{
"candidate_id": "v4-f901052615fa3aee",
"command": "python3 -m pytest -q tests/test_bounded_storage_reads.py",
"exit_code_on_base": 1,
"fails_on_base": true,
"normalisation": [],
"recorded_how_to_run": "python3 -m pytest -q tests/test_bounded_storage_reads.py",
"recorded_verified_fails_on_base": true,
"repository_id": "gitseed",
"runnable": true
}
],
"counts": {
"fails_on_base": 17,
"normalised": 4,
"runnable": 17,
"total": 17
},
"document_id": "cdeb-fresh-v8-acceptance-commands",
"problems": [],
"schema_version": 1,
"study_id": "cdeb-fresh-v8",
"what_normalisation_may_do": "Take a backticked command out of a prose wrapper, and replace a leading `pytest` or `python` with `python3 -m pytest` or `python3`. Nothing else. A broader rewrite would be editing the task.",
"what_this_is": "The command each episode runs to score task acceptance, normalised from the v6 task's human-readable how_to_run and verified on a freshly materialised tree.",
"why_normalisation_was_needed": "Four of the seventeen how_to_run strings are not shell commands: one is prose wrapping a backticked command, and three name an interpreter that is not on PATH. Each exits 127, which is not a failing test -- and scored as one, those four candidates lose all twenty of their episodes in both arms."
}
Original file line number Diff line number Diff line change
Expand Up @@ -58,8 +58,13 @@ argument for a 25-condition gate rather than a p-value: the one dataset that
could have produced a false headline was stopped by a condition about mechanism,
not significance.

`known_positive` was also blocked — its RBDR is real but below the 50% floor. A
generated effect large enough to see is not automatically a claim.
`known_positive` does reach the claim, and that is the gate working rather than
failing: it was generated with a large effect and the simulation holds the fifteen
non-statistical conditions at passing values. An earlier version of this note said
no scenario reached the claim, which was true only by accident — RBDR had been
implemented before it was defined, and the invented formula happened to fall below
the 50% floor. What matters, and what is asserted, is that every scenario with no
effect or a harmful one is blocked.

## Mutation controls

Expand Down
Original file line number Diff line number Diff line change
@@ -1,10 +1,10 @@
{
"analysis_sha256": "542e3618ad34b3e57733e9d367e0d3e83760111da3f1fca1cd2d317934081db8",
"analysis_sha256": "1fc6cfb70cb667a2c01955225abeb5c3662294f6d540300c62464fff4e86fb0a",
"document_id": "cdeb-fresh-v8-analysis-code-pin",
"mutate_analysis_sha256": "da3b0351a5469b4bbd3f6de3dcc93cd144220d6c33f65372abf6b7867639ab02",
"mutate_analysis_sha256": "4a0ca1fd09e4b26a515f4396b46ab84e21dbd84441e92a2f9e9329a9ba1582a4",
"schema_version": 1,
"simulate_sha256": "447d8380e0bf95288cf5e9c88313375fa76fe53179810a8b7356147bc2c73cfb",
"study_id": "cdeb-fresh-v8",
"test_analysis_sha256": "41c1b22b04434de4cf7e8d13d3f5b88eb5ea63af4800f2eee57629f5b0f70a46",
"test_analysis_sha256": "929da077132996f868772533375217f6b8787ab89be3198aec652a881dccefe6",
"what_this_is": "Digests of the analysis and its controls at the moment the recorded simulation and mutation results were produced. A test asserts the files still hash to these, so an edit without a rerun fails."
}
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
baseline unmutated: all controls pass
caught indeterminate counts as a success <- indeterminate is not a success
caught panel label drops the PANEL_ prefix <- panel truth table matches section 9.1, an indeterminate panel counts a
caught a duplicate assignment silently overwrites <- a duplicate assignment is refused
caught a second live row for one assignment is accepted <- two live rows for one assignment are refused
caught a superseded attempt with no retry passes <- a superseded attempt with no retry is refused
caught a superseded attempt still enters ITT <- crashed: ValueError: two live rows for ('c', 0, 'ON'): section 20 for
caught incomplete episodes dropped from ITT <- incomplete is not a success
caught one judge decides the panel <- panel truth table matches section 9.1
caught repositories weighted by candidate count <- repositories weighted equally
Expand All @@ -10,13 +12,14 @@ baseline unmutated: all controls pass
caught interval uses the extremes, not percentiles <- interval sits at the 2.5/97.5 percentiles
caught randomization never swaps labels <- randomization p small under a real effect
caught randomization always swaps labels <- randomization p small under a real effect
caught RBDR divides by a zero denominator <- crashed: ZeroDivisionError: float division by zero
caught RBDR counts every pair, not the revived ones <- RBDR undefined when nothing revived, RBDR counts pairs that revived an
caught RBDR treats no revival as a perfect score <- crashed: ZeroDivisionError: division by zero
caught gate passes when any condition holds <- strong claim fails one gate at a time, missing input is a failure, not
caught gate treats a missing input as a pass <- missing input is a failure, not a pass
caught gate answers without a stated input origin <- the gate refuses inputs with no stated origin
caught AC1 uses the product of marginals like kappa <- AC1 stays high where one category dominates, Fleiss kappa collapses on
caught three-way agreement counts non-unanimous episodes <- three-way agreement counts only unanimous episodes
caught Fleiss kappa drops the chance correction <- Fleiss kappa collapses on the same data

18/18 mutations caught
code pinned: analysis.py 542e3618ad34b3e5
21/21 mutations caught
code pinned: analysis.py 1fc6cfb70cb667a2
Loading
Loading