Skip to content

CDEB-Fresh v8: blind-panel fixed-benchmark effect trial - #857

Draft
MongLong0214 wants to merge 6 commits into
mainfrom
cdeb-v8-panel
Draft

CDEB-Fresh v8: blind-panel fixed-benchmark effect trial#857
MongLong0214 wants to merge 6 commits into
mainfrom
cdeb-v8-panel

Conversation

@MongLong0214

@MongLong0214 MongLong0214 commented Aug 26, 2026

Copy link
Copy Markdown
Owner

Status

This PR resumes the owner-authorized CDEB research line after v7 terminalized with zero product-effect rows.

Current branch state:

study                         cdeb-fresh-v8
fixed tasks                   17 (8 AOS / 9 gitseed)
expected coding episodes      340
expected blind judgements     1,020
measured product-effect rows  0
panel frozen                  no

Why v8 exists

v7 established that deterministic machine boundaries were operationalizable for 8 of 17 fixed tasks and unresolved for 9. It did not measure CommitLore's effect. v8 keeps all 17 tasks and replaces the deterministic-oracle prerequisite with a preregistered, arm-blind three-judge semantic panel.

Completed on this branch

  • V8 SSOT and preregistration installed.
  • Exact 17-task population imported and verified.
  • Product identity pinned to v1.2.0 / 90a8b212... / a0c54297....
  • Calibration corpus reduced to 47 known-label controls after four disputed Good controls were excluded pre-treatment and documented.
  • Forty-seven opaque calibration packets generated and blinding-audited.
  • First judge candidate completed all 47 cases:
    • judge: codex/gpt-5.6-terra
    • accuracy: 91.49%
    • violation recall: 88.24%
    • compliant recall: 93.33%
    • malformed: 0
    • individual threshold: PASS

Remaining preregistered work

  • Run at least two more fixed judge candidates over the same 47 packets.
  • Select and freeze the three-judge panel mechanically.
  • Complete 17 manipulation preflights and synthetic coding-agent smoke.
  • Freeze coding runtime/model and exact 340-row schedule.
  • Run readiness red-team.
  • Only after PR-A readiness passes: execute 340 coding episodes, create 1,020 blind judgements, seal/reveal/analyse, publish and terminalize.

Safety

  • No benchmark measured episode has run.
  • measured_run_allowed=false.
  • product_effect_rows=0.
  • The nine v7 boundary-unresolved tasks remain in the fixed population.
  • No automatic v9.

Current execution dependency

The repository and orchestration artifacts are available, but two additional independent judge runs and the later coding-agent episodes require an authenticated model runtime. This PR will keep that dependency explicit rather than fabricate judge or episode outputs.

Closes #853 only after final v8 publication and terminalization.

…r cases shorter than specified

v7 established that 8 of the fixed 17 decisions yield a deterministic final-tree
predicate and 9 do not, and ended without running an episode because its
population was fixed and unresolved ambiguity was terminal there. That is a
result about instruments.

Reading a decision and judging whether one finished implementation clearly takes
the ruled-out approach is a different question from writing a predicate that
covers every implementation. v8 measures the first, so its primary instrument is
a blinded three-judge panel and an unresolved machine boundary is neither an
exclusion nor a hold reason. All 17 stay.

    17 tasks x 2 arms x 10 repetitions = 340 episodes
    340 episodes x 3 blind judges      = 1,020 primary judgements

## The calibration key is 47, not 51

The SSOT builds it from 17 Good A, 17 Good B and 17 Bad A, calling them v6 frozen
controls with known labels. Two things about this repository make that wrong as
written, and both are checkable:

The Good controls are not v6 artifacts. v6 kept no control bytes -- all 89 of its
control records carry the same seven prose keys, none carries a diff, and no file
under the v6 study contains patch text. These are the 34 v7 rebuilt, each verified
against both acceptances and read by two blind sessions.

Four of those 34 have no agreed label. Their judges split, and one drew a
VIOLATION_CONFIRMED from a session reading a builder that had never been told the
decision. Scoring them COMPLIANT would penalise a judge for reading them the way
a blind session already did and select for judges that agree with a disputed key
-- and that judge becomes the primary instrument for 340 episodes.

    COMPLIANT   30    v7 rebuilds, both blind judges agreed
    VIOLATION   17    16 v6 imports, 1 v7 rebuild
    excluded     4    retained as boundary-disputed controls, not deleted

Recorded as v8-d001 with the three options considered and what each costs.

## What the corpus cannot separate, measured rather than assumed

Labels and origins are nearly confounded: every COMPLIANT case is a v7 rebuild,
16 of 17 VIOLATION cases are v6 imports, and violation patches run about 2.4 times
larger by bytes. A judge could score well by reading size instead of the decision.

So the best surface-only classifier was measured against the same thresholds:

    patch bytes      81% accuracy, 71% violation recall
    files touched    79% accuracy, 53% violation recall
    added lines      70% accuracy, 35% violation recall

None clears the individual judge threshold of 85% with 80% recall both ways. That
does not remove the cue, it bounds it: a judge clearing the 92% panel threshold is
using more than size.

## Why the v7 specifications are withheld from judges

They are archived as metadata. A judge handed a boundary specification would apply
that specification rather than read the decision, which would make the panel a
proxy for the oracle v7 could not build for nine of these -- and v8 would be
measuring v7 again under a different name.

Record-Id: r-v8opens
Provenance: authored
Certainty: firm
Blast: system
Undo: easy
Ruled-out: calibrating on all 51 as the SSOT text specifies | four of the entries are disputed, and selecting judges against a disputed key hands the primary instrument to whichever judge agrees with it
Ruled-out: scoring the four disputed controls as INDETERMINATE | that invents a label neither blind session gave, and in every one of the four a judge said NOT_A_VIOLATION
Ruled-out: deleting the four from the corpus | they are the corpus's own evidence that a compliant-by-construction control can read as contested, and that is worth keeping where the calibration lives
Ruled-out: showing judges the v7 boundary specifications | it would replace the judgement being measured with the application of a specification, and nine of the seventeen have none to apply
Limit: the calibration sample is 47 and the SSOT sized its thresholds for 51. The thresholds are unchanged, so each is now met on slightly less evidence
Limit: 30 of 30 COMPLIANT cases and 1 of 17 VIOLATION cases are v7 artifacts. Origin and label remain confounded and the surface control bounds that rather than removing it
Limit: no judge has been scored, no panel selected, and no episode assigned. Everything here is a key and a rule, not a measurement
Verified: all 17 imported inputs present with digests matching; boundary status 8 settled and 9 unresolved matches the SSOT front matter; resolveActiveStudyRoot resolves to cdeb-fresh-v8 through the real loader; no personal path in any v8 file; tsc clean on both tsconfigs; vitest 3747 passed, 13 skipped, 0 failed
…t is still in them

Each packet is a judge's whole world: the decision as recorded, the task, the
diff, and the finished tree. What is left out of it is the blinding, so this
commits the audit alongside the packets rather than asserting they are blind.

Removed before a judge sees anything: `.git`, which carries the arm in branch
names and messages, and any file or directory whose name mentions commitlore,
cdeb, arm, assignment or delivery. Zero packets retain a `.git`.

Deliberately not removed: agent-written source and its comments. Redacting those
would mean the judge reads something other than what was produced, and a comment
containing a product word is the code rather than a cue.

Scanning all 47 for five cue classes finds four at zero — control variant labels,
arm labels, candidate ids, study ids. The fifth needs explaining rather than
fixing.

`Record-Id` appears in every one of the 25 gitseed packets, in two files:
`AGENTS.md` and `docs/adr/ADR-0008`. Both document the trailer format, and the
ids they carry are `r-<6+`, `r-gsf501` and `r-enadr17` — a format placeholder and
two records that belong to no benchmark candidate. It is the base repository,
byte-identical across every gitseed packet, and stripping it would edit the tree
the judge is being asked to read. Recorded, not removed.

Packet ids are a salted hash of candidate and variant, truncated. The key mapping
them back lives outside the packet directories, so a judge reading its own
directory cannot find the answer next to the question.

Record-Id: r-v8packets
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: stripping Record-Id from AGENTS.md and the ADR | they are the base repository rather than anything an agent wrote, identical in all 25 gitseed packets, and removing them makes the judge read a tree that never existed
Ruled-out: asserting the packets are blind without scanning them | blinding is the whole instrument here, and an instrument nobody checked is the thing this study line keeps finding at the bottom of its failures
Ruled-out: putting the calibration key inside the packet directories | a judge reading its own directory would find the expected label beside the question
Limit: this scan finds textual cues. A stylistic cue would pass it -- if v6 and v7 builders write recognisably differently a judge could learn that instead of the decision, and the surface-only control bounds how far size alone gets rather than proving no such cue exists
Limit: 25 of 47 packets are gitseed and 22 are agent-operator-score, so a cue present in only one repository shows up in only part of the corpus
Verified: 47 packets built with zero patch-application failures; zero .git directories remain; four of five cue classes scan clean; the Record-Id occurrences traced to two base-repository files with their exact values recorded; the key written outside the packet tree
… where v7 said the rules run out

Forty-seven calibration packets, zero malformed outputs.

    accuracy            91.5%   threshold 85%
    VIOLATION recall    88.2%   threshold 80%
    COMPLIANT recall    93.3%   threshold 80%

    COMPLIANT -> COMPLIANT  28      COMPLIANT -> VIOLATION  2
    VIOLATION -> COMPLIANT   2      VIOLATION -> VIOLATION 15

The single trial judgement that opened this question turned out not to be a
corpus defect. Errors run two in each direction, and 91.5% sits well clear of the
81% a size-only classifier reaches here — the margin is what says this judge reads
more than patch size.

Where the errors fall is the finding:

    v7 SETTLED      21/22   95%
    v7 UNRESOLVED   22/25   88%

Three of the four errors are on decisions v7 reported as having no boundary a
program could apply. Two instruments sharing no machinery point the same way:
where a rule's own words do not settle it, a judge reading one implementation is
also less likely to match the key. That is a property of the decisions, not a
defect in either instrument.

All four errors carry confidence=high. The PRD treats confidence as descriptive
and gives it no weight in the vote; this is the evidence for that choice rather
than an assumption behind it.

Record-Id: r-v8candcodex
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: scoring partway through the batch | reading accuracy while the corpus is still running is choosing a judge during calibration, which is the one thing calibration cannot survive
Ruled-out: treating the trial disagreement as proof the key was wrong | one case cannot separate a contested candidate from a bad key, and the full corpus does
Ruled-out: weighting the vote by confidence | every error here is high-confidence, so confidence carries no information about correctness on this corpus
Limit: one candidate of up to five, and a panel needs three. Nothing is selected yet
Limit: 47 cases, so each percentage point is worth about half a case and the recall figures rest on 17 and 30 cases respectively
Limit: the accuracy split by v7 boundary status is 22 against 25 cases. It is consistent with the reading given and is not powered to establish it
Verified: 47 of 47 packets judged with zero malformed outputs; scores computed once, after the batch completed, against the key committed before any judging; the four errors listed by candidate, variant, expected label, returned label and confidence
@github-actions

github-actions Bot commented Aug 26, 2026

Copy link
Copy Markdown

CommitLore — record lint

Trailers: clean — 6 commits in origin/main..d0e69699aef6fcc0d84936b590e03d2420362152
Active constraints: 29 limits · 43 ruled-out · 0 warnings — from 14 records over 64 changed paths

Active constraints for the paths this PR touches

Limits (29)

  • r-v8calibrationtransition d0e6969 — this records progress and recovery only; it creates no semantic judgement for a measured episode
  • r-v8statuscalibration 3d5126e — two additional fresh candidate runs and panel-level scoring remain
  • r-v8operatorrecovery 6977df6 — this is an operational recovery artifact, not a study outcome or a substitute for the missing inference runtime
  • r-v8candcodex 379d931 — one candidate of up to five, and a panel needs three. Nothing is selected yet
  • r-v8candcodex 379d931 — 47 cases, so each percentage point is worth about half a case and the recall figures rest on 17 and 30 cases respectively
  • r-v8candcodex 379d931 — the accuracy split by v7 boundary status is 22 against 25 cases. It is consistent with the reading given and is not powered to establish it
  • r-v8packets ad51aa8 — this scan finds textual cues. A stylistic cue would pass it -- if v6 and v7 builders write recognisably differently a judge could learn that instead of the decision, and the surface-only control bounds how far size alone gets rather than proving no such cue exists
  • r-v8packets ad51aa8 — 25 of 47 packets are gitseed and 22 are agent-operator-score, so a cue present in only one repository shows up in only part of the corpus
  • r-v8opens 4ed43c4 — the calibration sample is 47 and the SSOT sized its thresholds for 51. The thresholds are unchanged, so each is now met on slightly less evidence
  • r-v8opens 4ed43c4 — 30 of 30 COMPLIANT cases and 1 of 17 VIOLATION cases are v7 artifacts. Origin and label remain confounded and the surface control bounds that rather than removing it
  • r-v8opens 4ed43c4 — no judge has been scored, no panel selected, and no episode assigned. Everything here is a key and a rule, not a measurement
  • r-v7terminal a980313 — every session in the boundary work is one model family. Independent of each other, not of what that family finds hard to pin down, and a different family might settle more or fewer than eight
  • r-v7terminal a980313 — three of the eight rest on one comparator session failing to construct a separating tree, which is weaker than a proof that none exists
  • r-v7terminal a980313 — eight settled counts boundaries, not oracles. None was implemented, run against a control, or attacked
  • r-v7terminal a980313 — the rebuilt Good controls are v7 artifacts occupying v6 slots, and four of the thirty-four were not read as cleanly compliant by their blind judges
  • r-v7opens f900bdd — this imports and locks. No oracle exists yet for any of the seventeen, so no episode may run and STATUS says measured_run_allowed false
  • r-v7opens f900bdd — the bundles are gitignored by a recorded decision, so this lock proves the bytes on disk match what v6 sealed and cannot restore them if they are lost
  • r-v7opens f900bdd — the claim gate still reads a primary interval that measures execution variability with the tasks fixed. If the pinned agent is near-deterministic that interval narrows toward zero width, and the preregistration says so rather than changing the gate the owner registered
  • r-v6resultpublished 6e69a58 — the finding is about two repositories at one frozen commit, and both were selected in v5 for deterministic test suites -- correlated with thorough ones. A less tested codebase would leave more wrong paths open and this study cannot say how many
  • r-v6resultpublished 6e69a58 — no episode ran, so nothing here bears on whether automatic delivery helps. The study measured what it could ask, not what it set out to answer
  • r-v6sourcelock a25e516 — the source pool inherits v5's classification of these 34 as functionally violable. That classification was made against v5 tasks, and v6 builds fresh tasks, so a decision violable under one task need not be violable under another. The task-buildability pipeline is what tests that, per candidate
  • r-v6sourcelock a25e516 — 30 of the 34 carry a Record-Id and 4 do not. Delivery is judged on decision content rather than identifier, so this does not gate them, but it does mean the pool is not homogeneous in how the record is stored
  • r-v5terminalphase 58d31c8 — the terminal-phase list still has to learn each new stopping word. A study that ends at a stage nobody has reached yet will resolve as active until someone notices, and the only defence is that the phase is written by the study itself
  • r-v5authority 1012c88 — reusing the frozen bundles bounds the corpus to what existed on 2026-08-20. If v5 holds, the 14 excluded decisions are reported by name and count so the owner can order a deliberate re-snapshot rather than have one smuggled in here
  • r-v5authority 1012c88 — this commit registers the model and the thresholds. It produces no v5 count, and the preregistration is deliberately written before the census so the counts cannot choose the rules
  • r-v4stage0id 6b427af — this proves a terminal study cannot be resolved as active through the declaration. It does not stop a caller that names a study root directly, which is why the measured-run gate is separate and still shut
  • r-v3terminalseal 7754f1a — the placeholder row remains in the ledger and always will. This makes it legible, not absent, and a reader who takes digests on faith rather than reading the deviation is still misled
  • r-v3terminalseal 7754f1a — the canonical digest binds the artifact list it is given. A transition that names too few artifacts is bound to a partial set, and nothing here decides what the right set is for a future study
  • r-v3terminalseal 7754f1a — guard coverage is unchanged -- thirteen exclusion kinds remain uncovered and one scan inert, recorded in the mutation baseline

Ruled out (43)

  • r-v8calibrationtransition d0e6969 — jumping directly to PANEL_FROZEN | only one candidate has completed and the panel-level thresholds cannot yet be evaluated
  • r-v8statuscalibration 3d5126e — marking calibration complete after one candidate | the preregistration requires a fixed three-judge panel selected from calibrated candidates
  • r-v8operatorrecovery 6977df6 — fabricating two calibration candidates to keep the study moving | judge outputs are the primary measurement instrument and invented rows would invalidate every downstream estimate
  • r-v8operatorrecovery 6977df6 — treating this operator's already-informed context as a fresh blind judge | this context has seen the calibration key and first-candidate errors
  • r-v8operatorrecovery 6977df6 — starting measured episodes before panel/runtime/schedule freeze | the preregistration explicitly forbids it
  • r-v8candcodex 379d931 — scoring partway through the batch | reading accuracy while the corpus is still running is choosing a judge during calibration, which is the one thing calibration cannot survive
  • r-v8candcodex 379d931 — treating the trial disagreement as proof the key was wrong | one case cannot separate a contested candidate from a bad key, and the full corpus does
  • r-v8candcodex 379d931 — weighting the vote by confidence | every error here is high-confidence, so confidence carries no information about correctness on this corpus
  • r-v8packets ad51aa8 — stripping Record-Id from AGENTS.md and the ADR | they are the base repository rather than anything an agent wrote, identical in all 25 gitseed packets, and removing them makes the judge read a tree that never existed
  • r-v8packets ad51aa8 — asserting the packets are blind without scanning them | blinding is the whole instrument here, and an instrument nobody checked is the thing this study line keeps finding at the bottom of its failures
  • r-v8packets ad51aa8 — putting the calibration key inside the packet directories | a judge reading its own directory would find the expected label beside the question
  • r-v8opens 4ed43c4 — calibrating on all 51 as the SSOT text specifies | four of the entries are disputed, and selecting judges against a disputed key hands the primary instrument to whichever judge agrees with it
  • r-v8opens 4ed43c4 — scoring the four disputed controls as INDETERMINATE | that invents a label neither blind session gave, and in every one of the four a judge said NOT_A_VIOLATION
  • r-v8opens 4ed43c4 — deleting the four from the corpus | they are the corpus's own evidence that a compliant-by-construction control can read as contested, and that is worth keeping where the calibration lives
  • r-v8opens 4ed43c4 — showing judges the v7 boundary specifications | it would replace the judgement being measured with the application of a specification, and nine of the seventeen have none to apply
  • r-v7terminal a980313 — running v7 on the eight settled candidates | the preregistration fixed the population at seventeen before any task existed, and reducing it after seeing which ones resolved is the discretion the floor was written to remove
  • r-v7terminal a980313 — relaxing the oracle input boundary so a provenance predicate becomes decidable | that boundary exists so an oracle cannot see the arm, and widening it to rescue one candidate reopens what it was written to close
  • r-v7terminal a980313 — reporting eight of seventeen as a result about CommitLore | no episode ran, no arm was assigned, and nothing here bears on whether automatic decision delivery helps an agent
  • r-v7opens f900bdd — importing the one v5 oracle for v4-377f04276465b59d | it is a lexical scan over six fixed paths, which the r1 priority ladder admits only where the decision is itself lexical, and reusing it would carry v5's unvalidated instrument into the endpoint
  • r-v7opens f900bdd — correcting the dist digest quietly in the lock | the declared value is what a reader of the first draft would look for, and deleting it removes the evidence that the correction was needed
  • r-v7opens f900bdd — opening v7 before v6's terminal state was re-read from the tree | the SSOT permits progression across green gates, and a gate that trusts its own summary of the predecessor is not a gate
  • r-v7opens f900bdd — leaving the v6 active-study assertion and pinning it to null | it would fail again at the next successor, and a test that has to be edited on every transition is not recording a durable fact
  • r-v6resultpublished 6e69a58 — leaving STATUS.json for a later commit since the artifacts already carried the verdict | a study whose own record does not say it ended can still be resolved as active, which is the failure the terminal-phase list exists to prevent
  • r-v6resultpublished 6e69a58 — asserting the most recent terminal study id in a test | it changes whenever a study ends, and a test that fails on a normal transition teaches nothing when it does
  • r-v6resultpublished 6e69a58 — publishing only the floors and not the obstacles | the eight failures are the finding, and a result that reported the shortfall without the file-and-line evidence would read as the corpus being too small rather than as its wrong paths being closed
  • r-v6sourcelock a25e516 — accepting the SSOT dist digest because three predecessors carried it | provenance through three studies is not verification, and the reason it survived that long is that nobody compared it to a file
  • r-v6sourcelock a25e516 — holding the study on the digest mismatch | the release is unambiguous by tag and commit, and stopping a study over a number that pins nothing would attribute a bookkeeping defect to the science
  • r-v6sourcelock a25e516 — substituting a newer release now that the pinned digest is unverifiable | the SSOT forbids it, and the commit verifies, so nothing about which build to run is actually unknown
  • r-v6sourcelock a25e516 — asserting the current occupant of the active-study slot in tests | those tests failed at the last two transitions and would fail at every future one, teaching nothing each time; they now assert what is durable -- v5 is the last terminal study and never the active one
  • r-v5terminalphase 58d31c8 — leaving v5 as the active study since no successor exists | it holds a published TERMINAL_HOLD, and a study that can still be resolved as active is one that can still have work attributed to it
  • r-v5terminalphase 58d31c8 — inventing a v6 id to keep active_study_id populated | the resolver would point at a directory that does not exist, and a preregistration written to fill a field is not a preregistration
  • r-v5terminalphase 58d31c8 — relaxing the v4 test to assert only that no error is thrown | it would then pass whether or not v4 was active, which is the one thing it exists to check
  • r-v5authority 1012c88 — renaming v4's independent-prose gate rather than removing it | the preregistration and the authority policy both forbid any gate that requires a decision to be documented outside its record, because a gate that does is the v4 gate whatever it is called
  • r-v5authority 1012c88 — re-snapshotting to grow the corpus | the only decisions a fresh snapshot adds are the 14 authored during the study, in the repository where they would decide eligibility
  • r-v5authority 1012c88 — asking the owner to attest the 14 are natural | that is owner testimony, which is disabled, and the answer would arrive after the counts were known
  • r-v5authority 1012c88 — keeping invalidated as the only terminal phase | v4 is not invalid, it is finished, and a phase list that only knows about invalidation would let a study with a published verdict be run again
  • r-v4stage0id 6b427af — keeping a hardcoded list of terminal study ids in the resolver | the list and the studies drift apart silently, and the drift shows up as a terminated study resolving cleanly
  • r-v4stage0id 6b427af — leaving the declaration at null and passing the study root explicitly everywhere | every caller then carries the choice, and the one caller that forgets picks a default nobody reviewed
  • r-v4stage0id 6b427af — reusing the v3r1 study directory under a new name | §4.3 requires a new study id, and a renamed directory keeps the qualification verdicts this estimand discards
  • r-v3terminalseal 7754f1a — correcting the placeholder digests in place | the correction is indistinguishable from the mistake it repairs, and the row is historical evidence rather than a working value
  • r-v3terminalseal 7754f1a — recomputing digests for the historical rows from today's artifacts | the artifacts have changed since, so the result would be a number that never bound anything, wearing the authority of one that did
  • r-v3terminalseal 7754f1a — hand-maintaining evidence-matrix.md beside the JSON | two copies of the same claims disagree eventually and the disagreement is silent
  • r-v3terminalseal 7754f1a — leaving cdeb-fresh-v3r1 as the default study root | a terminated study as a fallback is how a measured run gets attempted against a study that ended

git log --follow accepts exactly one pathspec, so renames are not followed for 64 paths; query one path at a time to follow its rename chain

Trailer violations fail this check. Active constraints are informational — they are what the repository already decided, not a verdict on this PR.

The prior execution agent stopped after the first judge candidate completed. This
records the exact recoverable state, the remaining gates, and the runtime resources
the replacement operator does not possess. It does not alter the preregistration,
freeze a panel, or authorize measured runs.

Record-Id: r-v8operatorrecovery
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: fabricating two calibration candidates to keep the study moving | judge outputs are the primary measurement instrument and invented rows would invalidate every downstream estimate
Ruled-out: treating this operator's already-informed context as a fresh blind judge | this context has seen the calibration key and first-candidate errors
Ruled-out: starting measured episodes before panel/runtime/schedule freeze | the preregistration explicitly forbids it
Limit: this is an operational recovery artifact, not a study outcome or a substitute for the missing inference runtime
Verified: v7 terminal merge exists; v8 branch contains the 47-case corpus and candidate-codex result; candidate-codex passes all individual thresholds; STATUS still forbids measured runs
The first of three required judge candidates completed the frozen 47-case
calibration and passed its individual thresholds. The panel is not frozen and no
measured episode is authorized.

Record-Id: r-v8statuscalibration
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: marking calibration complete after one candidate | the preregistration requires a fixed three-judge panel selected from calibrated candidates
Limit: two additional fresh candidate runs and panel-level scoring remain
Verified: candidate-codex contains 47 results, zero malformed outputs and passes all individual thresholds; product-effect rows remain zero
This advances the truthful lifecycle from draft to judge calibration while
keeping panel_frozen and measured_run_allowed false.

Record-Id: r-v8calibrationtransition
Provenance: authored
Certainty: firm
Blast: local
Undo: easy
Ruled-out: jumping directly to PANEL_FROZEN | only one candidate has completed and the panel-level thresholds cannot yet be evaluated
Limit: this records progress and recovery only; it creates no semantic judgement for a measured episode
Verified: the committed candidate result contains 47 outputs and passes every registered individual threshold
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

CDEB-Fresh v8: blind-panel calibration and final effect trial

1 participant