Abstract
Experiments comparing large language model (LLM) interventions on sparse software-engineering tasks can spend most of their budget before establishing that a comparison is measurable. This work adapts established multi-endpoint feasibility methodology to sparse LLM author/reviewer experiments, operationalises three viability endpoints, and reports a prospectively governed case in which the screen stopped a confirmatory study before its reserved evaluation pool was consumed.
Two development pilots failed criteria fixed before model calls. The paper contributes an LLM-specific operationalisation of established feasibility methodology and a documented case in which a screen stopped a study before its reserved evidence was spent; it claims no new statistical method, causal result about critique, or external validation.
Reproducibility and artifact scope
The public page intentionally distinguishes materials that are mirrored here from materials that were supplied to the journal for peer review. The confirmatory task pool remained sealed and unobserved, so there are no confirmatory task outcomes to publish. Third-party raw artifacts are not redistributed where upstream licences do not permit redistribution.
Supplementary records submitted for peer review
- CHECKPOINT-MANIFEST-external-raw.sha256
- EXTERNAL-VALIDATION-PROTOCOL.md
- PREREGISTRATION.md
- FINAL-CLAIM-REGISTRY.md
- EXTERNAL-RAW-RESULTS.md
- EXTERNAL-OUTCOME-ACCESS-LOG.md
- EXTERNAL-VALIDATION-PROTOCOL-AMENDMENT-002.md
- EXTERNAL-QC-RESULTS.md
- DERIVATIVES.md
- STAGEB-INTEGRITY-RERUN-AMENDMENT.md
- EXTERNAL-RESULTS-VALIDITY-AUDIT.md
- PILOT_PROTOCOL.md
- EXTERNAL-ENDPOINT-DEFINITION-REGISTRY.md
- EXTERNAL-PAPER-IDENTITY-REGISTRY.csv
- EXTERNAL-VALIDATION-PROTOCOL-AMENDMENT-001.md
- MAJOR2-NOISE-NULL-PROTOCOL.md
- EXTERNAL-VALIDATION-QC-PLAN.md
- README.md
Analysis and verification code submitted for peer review
- make_figures.py
- stageB_canonical_verify.py
- make_tikz_figures.py
- major2_noise_null.py
- major2_noise_null_check.py
- stageB_canonical_report.py
- glmm_power.py
- f3_stageB.py
- f3_stageB_canonical.py
- qc_fixtures.py
- endpoints.py
Derived result and ledger files submitted for peer review
- major2_noise_null_results.json
- qc_results.json
- raw_results.json
- endpoint_e_results.json
- stageB_H1_original_process.json
- stageB_H2_resume_writer.json
- stageB_canonical_rerun.jsonl
- stageA.json
- stageB_canonical_summary.json
- SUPPLEMENT-MANIFEST.sha256
- requirements.txt
- stageB_canonical_verification.txt
Public file integrity
Use the checksum below to verify the manuscript downloaded from this mirror.
da539b5760e6c4ec9a96128ad6b83ce07d5316ce44daedab616ef7f30fe2e9f5
Also available in MANIFEST.sha256.