Benchmark QA convergence β 2026-09-10
Historical validation snapshot. Re-review rejected commit 2310099b for new
row-contract, publication-scope and devloop defects. The follow-up fixes and
their validation are recorded in QA follow-up.
The UNMEASURED default-core outcome below was a harness defect: optional
patch proof should not have escalated the overall verdict.
This branch reapplies the retained tier-2 and tier-3 work onto main
b297e8a25d08a86d28239dce4138ee889034dffc6a85aec7aaa7c98252fccc27
and addresses the rejected stack's QA findings. It is a single review branch,
not a merge receipt or a claim of industry-leading benchmark completeness.
QA workspace: /Users/mrmrs/o/worker-mrmrs-20260910T200629-bench-qa-fixes/bench.
Branch: mrmrs-eec7ec. Generated evidence is intentionally untracked in this
worker and its parent directory.
Corrections
- Publication requires exact patch application evidence only for operations
that declare that requirement or claim successful application. Structural
diff.dirtyrows remain diagnostic. Output formats come from versioned command semantics, including native JSON probes. - Git correctness instruments resolve from explicit configuration or PATH and record binary identity. Missing, killed, timed-out or invalid instruments produce unmeasured evidence. They cannot manufacture an agent task failure. Independently observed agent failures remain failures.
- Both agent and scripted rows label committed-tree evidence. Git uses the known-good Git HEAD reader; Oak uses a fresh subject-produced history export checked by that reader. Publication refuses unequal cross-subject evidence.
- Devloop distinguishes exact pre-existing patch failures from new failures and unavailable proof. Core mutations use identical bytes across subjects; failure comparisons match repetitions and instrument identities.
- Campaign admission reserves the frozen trial timeout instead of shortening it. Reports exclude unavailable-oracle and legacy shortened-timeout trials from complete paired comparisons.
- Journals retain rejected rows and uncertain starts. Read-only status/report can inspect an intact prefix; acknowledged recovery archives the original bytes and never retries an uncertain trial. Start records precede trial directory creation. Derived rows append incrementally and rebuild on resume.
- Git environment sanitization is shared. Real agent environments use an explicit allowlist. Lock paths are canonical, recorded and checked on resume; workstation instructions require an explicit shared lock path.
- Failure rows participate in baseline population identity checks. Provenance excludes prose-only changes, malformed plans fail with the documented code, and configured capture limits also govern parsing.
Verification
OAK_BENCH_LOCK_PATH=/Users/mrmrs/o/oak-bench-measurement.lock python3 -m unittest discover -s tests completed successfully: 1,054 tests, three skipped,
110.838 seconds. The complete output is ../qa-tests-final.log.
The final live devloop used the same pinned Oak 0.103.0 binary for baseline
and candidate, two repetitions, the core lane and the shared measurement lock.
It recorded 402 core rows and zero baseline regressions or new Git guardrail
breaches. The tiny-text patch defect is PREEXISTING. The overall verdict is
UNMEASURED (exit 2), because the two large binary scenarios exceed the patch
capture limit. This is not a performance win or a complete correctness pass.
See ../qa-devloop-final.log and
results/qa-devloop-final/20260910T202819Z/ in the QA worker.
All 18 real diff.dirty rows have zero output-proof gate failures. Both Git
tiny-text full-patch rows have successful exact-tree application; all four Oak
baseline/candidate rows retain the known patch failure and fail publication
proof as intended. Earlier failing integration logs are preserved separately.
Live campaign evidence in the worker parent:
qa-campaign-eight-report.json: eight successful mock trials across two workflows, Git/Oak and two repetitions. Both evidence sources are labeled.qa-campaign-admission-report.json: a 0.15-second budget against a frozen 60-second timeout starts zero trials; all eight remain unstarted.qa-campaign-torn-report-after.json: seven completed trials and one interrupted trial. Read-only inspection preserved the damaged bytes; acknowledged recovery archived them without retrying the uncertain trial.
These reports all remain claim_eligible: false. Campaign evidence predates
the final core-only mutation fix; the final instrument suite verifies the
integrated source after that fix.
Remaining qualification limits
Oak 0.103.0's no-newline patch defect remains reproducible and is tracked as
fb-521. The harness must expose it; this branch does not fix the product.
Fresh Oak exports remain subject-produced evidence, with before/after HEAD
checks rather than an atomic pinned independent object reader. Large patch
capture limits can leave applicability unmeasured.
No paid-provider campaign, independent Oak durability certification, broad
task-distribution validation, power analysis or human study was run. The
staged requirements remain in agent-campaign-plan.md. The
oakmark-throughput-benchmark-transplant preservation branch is untouched.