Next action: v8 is calibrating its judge panel against 47 controls. Nothing is blocked. v7 closed as TERMINAL_HOLD_FINAL and is merged; v8 is owner-authorized and runs a blinded three-judge panel so that an unresolved machine boundary no longer stops the study.
This issue tracks the CDEB research line to its end.
v7 — terminal
CDEB-Fresh v7 reached TERMINAL_HOLD_FINAL before any product-effect episode.
Eight of the fixed 17 decisions yielded a semantic boundary precise enough for
deterministic oracle construction and nine did not. Because the preregistered
population was fixed at all 17 tasks and unresolved ambiguity was terminal under
v7, the population was not reduced post hoc. This result concerns deterministic
machine adjudicability, not the causal effect of CommitLore delivery.
fixed benchmark population 17
semantic boundary settled 8
semantic boundary unresolved 9
measured product-effect episodes 0
Each decision was read twice by independent sessions seeing only the rule and the
frozen repository. Where both drew a boundary, a third tried to build a tree they
would classify differently — agreement is what survived that. Where they split, a
fourth read the rule again with both attempts anonymised.
The nine that did not settle fail the same way: a rule turning on a term it never
defines, with the recorded reason reaching further than the recorded words —
literally, a badge, add the two paths, prose field. One is not vagueness
at all: hand-maintained is about how a file came to exist, and a finished tree
does not record that.
These decisions were written by people for people in a commit trailer, and they
read perfectly well that way.
Merged as #856.
v8 — owner authorized, running
The owner supplied a new SSOT rather than accepting that the line ends. v8 keeps
all 17 tasks and changes the instrument:
|
v7 |
v8 |
| primary instrument |
deterministic oracle per candidate |
blinded three-judge semantic panel |
| unresolved boundary |
terminal |
neither exclusion nor hold reason |
| scale |
340 episodes, never reached |
340 episodes × 3 judges = 1,020 judgements |
The exact study question: across the exact 17 frozen decision-sensitive
tasks, does automatic delivery of the candidate-relevant CommitLore decision
before the first relevant mutation raise the rate of episodes that complete, pass
both acceptances, and are read as compliant by a blind panel — relative to
suppressing that automatic delivery?
Judges never see the arm, the boundary status, the agent identity, the transcript,
the delivery payload, the acceptance result, or the v7 specifications. That last
exclusion matters: a judge handed a boundary specification would apply it rather
than read the decision, which would make the panel a proxy for the oracle v7
could not build.
The panel is the instrument, so its own reliability is published whatever it
shows. If agreement is poor the causal estimates are still computed and released,
but the result cannot be positive or null in a strong sense.
One deviation is already recorded. The SSOT builds the calibration key from 51
controls assuming every Good control carries a known compliant label. v6 kept no
control bytes, so those are the 34 v7 rebuilt — and four of them have no agreed
label, one having drawn a violation verdict from a blind session reading a builder
that was never told the decision. Calibrating on those would select judges that
agree with a disputed key. The key is 47; the four are retained, not deleted.
No automatic v9
v8 is the final planned study of this line. Any successor requires a separate
owner decision and is not generated automatically.
This issue closes when v8 publishes its result — positive, qualified, null,
negative, indeterminate, or a terminal hold.
Next action: v8 is calibrating its judge panel against 47 controls. Nothing is blocked. v7 closed as TERMINAL_HOLD_FINAL and is merged; v8 is owner-authorized and runs a blinded three-judge panel so that an unresolved machine boundary no longer stops the study.
This issue tracks the CDEB research line to its end.
v7 — terminal
Each decision was read twice by independent sessions seeing only the rule and the
frozen repository. Where both drew a boundary, a third tried to build a tree they
would classify differently — agreement is what survived that. Where they split, a
fourth read the rule again with both attempts anonymised.
The nine that did not settle fail the same way: a rule turning on a term it never
defines, with the recorded reason reaching further than the recorded words —
literally,a badge,add the two paths,prose field. One is not vaguenessat all:
hand-maintainedis about how a file came to exist, and a finished treedoes not record that.
These decisions were written by people for people in a commit trailer, and they
read perfectly well that way.
Merged as #856.
v8 — owner authorized, running
The owner supplied a new SSOT rather than accepting that the line ends. v8 keeps
all 17 tasks and changes the instrument:
The exact study question: across the exact 17 frozen decision-sensitive
tasks, does automatic delivery of the candidate-relevant CommitLore decision
before the first relevant mutation raise the rate of episodes that complete, pass
both acceptances, and are read as compliant by a blind panel — relative to
suppressing that automatic delivery?
Judges never see the arm, the boundary status, the agent identity, the transcript,
the delivery payload, the acceptance result, or the v7 specifications. That last
exclusion matters: a judge handed a boundary specification would apply it rather
than read the decision, which would make the panel a proxy for the oracle v7
could not build.
The panel is the instrument, so its own reliability is published whatever it
shows. If agreement is poor the causal estimates are still computed and released,
but the result cannot be positive or null in a strong sense.
One deviation is already recorded. The SSOT builds the calibration key from 51
controls assuming every Good control carries a known compliant label. v6 kept no
control bytes, so those are the 34 v7 rebuilt — and four of them have no agreed
label, one having drawn a violation verdict from a blind session reading a builder
that was never told the decision. Calibrating on those would select judges that
agree with a disputed key. The key is 47; the four are retained, not deleted.
No automatic v9
v8 is the final planned study of this line. Any successor requires a separate
owner decision and is not generated automatically.
This issue closes when v8 publishes its result — positive, qualified, null,
negative, indeterminate, or a terminal hold.