Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions bench/cdeb/ACTIVE-STUDY.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"active_study_id": null,
"active_study_id": "cdeb-fresh-v8",
"last_terminal_study_id": "cdeb-fresh-v7",
"status": "no-active-study",
"reason": "cdeb-fresh-v7 reached TERMINAL_HOLD_FINAL before any product-effect episode. Of the fixed 17 decisions, 8 yielded a semantic boundary precise enough for deterministic oracle construction and 9 did not. The preregistered population was fixed at all 17 and unresolved ambiguity was terminal, so the population was not reduced post hoc and no episode was run. The study holds zero measured product-effect rows. This result concerns deterministic machine adjudicability, not the causal effect of CommitLore delivery. v3, v3r1, v4, v5, v6 and v7 are terminal and none may be resumed; a successor requires a separate owner decision and is not generated automatically.",
"status": "active",
"reason": "cdeb-fresh-v8 is the final blind-panel effect trial of this research line, opened by a separate owner decision after cdeb-fresh-v7 reached TERMINAL_HOLD_FINAL holding zero measured rows. v7 established that 8 of the fixed 17 decisions yield a deterministic final-tree predicate and 9 do not; that is a result about instruments, not about the product. Reading a decision and judging whether one finished implementation clearly takes the ruled-out approach is a different question from writing a predicate covering every implementation, and v8 measures the first. The primary instrument is a blinded three-judge semantic panel, so an unresolved machine boundary is neither an exclusion nor a hold reason and all 17 tasks stay in the population. The design is 17 tasks x 2 arms x 10 repetitions = 340 episodes, each judged by 3 blind judges for 1,020 primary judgements, with no pilot and no sample-size gate. Judges are calibrated against 47 controls whose labels two blind sessions agreed on; four Good controls whose judges split are retained but excluded from the key, recorded as v8-d001. v3, v3r1, v4, v5, v6 and v7 are terminal and none may be resumed. This is the final planned study; a successor requires a separate owner decision and is not generated automatically.",
"successor_requires_new_study_id": true
}
Loading
Loading