Koch Laboratory

Two studies on machine-generated statements: criteria before results

Two studies from the forecasting programme, each settled by its own measurement: whether a system running without supervision learns rather than rationalises — tested on questions resolved after its model’s training cut-off, with safeguards that stop output instead of attaching a warning — and how lowering a model’s numerical precision affects the factual fidelity of statements derived from source documents. Neither has a result yet; the criteria are published before any result, and the entry corrects the earlier statement that implementation had not started.

Koch Laboratory — two studies on machine-generated statements: criteria before results

This entry continues the entry of 2026-08-10 on forecasts with measured own fallibility. It covers only two studies, each settled by a measurement of its own — Study A develops direction 4 of that entry, Study B has not been described before — and corrects one statement from that entry.

Method. Problem → state of the art → falsifiable criterion → method → status and next step. Withheld are the solution construction, the product form, the application domain, and the models and parameters.

Criteria before results. The criteria of both studies are described below before any result exists, and the date of this entry is stamped. Thresholds, models and data will be fixed before the first measurement and published together with the results.

1. Study A — learning or rationalisation: a control set after the training cut-off and fail-closed safeguards

Research question. Can a system running without supervision show that it is learning rather than merely getting better at justifying — on questions whose outcome became known only after its model’s training-data cut-off — and, when that control fails, does it stop publishing by itself instead of carrying on with a warning? Why it is hard. The measurement loop of such a system measures agreement with its own corpus, not with reality; a system that has learnt to justify rather than predict shows rising metrics and quiet operation — a silent sham success. A question whose outcome may have reached the training data does not separate prediction from recall, and questions beyond the cut-off become fewer with every change of model. With few resolutions per day, a statistical test detects deterioration late; a stop threshold that is too sensitive halts the system for no reason, one that is too sluggish acts after the fact. And a warning that someone has to read is not a safeguard. State of the art (published). Contamination — evaluation data leaking into training data — is documented (e.g. Sainz et al. 2023; Golchin and Surdeanu, ICLR 2024); evaluations designed to avoid it use questions newer than the model or about future events (e.g. LiveBench, ForecastBench). A model can give a persuasive explanation that does not reflect the actual basis of its answer (Turpin et al., NeurIPS 2023). Abstaining under uncertainty is known for single answers (Chow’s reject option, 1970), and moving to a safe state after a detected fault comes from functional safety (IEC 61508). Benchmarks, however, assess the model from outside; here the control is meant to work inside the running system and decide whether it may go on publishing. Success criterion. Performance on the control set is distinguishable from chance by a test fixed before the first measurement; stopping publication is demonstrated in operation — automatically, without human involvement; the stop thresholds are measured, not assumed in advance. An openly admitted negative outcome: if performance on the control set remains permanently indistinguishable from chance, the study ends with a solid negative finding — and the system, by its own rule, does not publish. Method. Control questions whose outcome is already settled but came after the training-data cut-off of the model under study; the set is renewed with every change of model and evaluated separately from all other measurements. The safeguards act by themselves; the test for the criterion and the stop conditions will be fixed before the first measurement. Status: no result on the criterion — the control set does not exist yet. Work on the safeguards began with the implementation in August 2026 and is not presented here as a result. Next step: building the control set, including an assessment of existing public question sets (licence, fit, size).

2. Study B — factual fidelity under reduced numerical precision

Research question. How does lowering a model’s numerical precision (quantisation, compression of the weights) affect the factual fidelity of statements derived from source documents — and can the measurement yield a rule for rejecting a statement? Why it is hard. The compute budget forces lower precision, whose effects are mostly assessed with aggregate measures. A loss of fidelity need not show in fluency: a model can write smoothly and add a fact that was not in the source — and is then worse than a stiff one that adds nothing. The line between generalisation and addition is blurred, because a correct generalisation formally resembles a claim from outside the source. The reference set is necessarily small, which limits the power of the resulting rule, and a rule that is too strict can reject so much that generating the statements loses its point. State of the art (published). Quantisation methods (e.g. GPTQ, Frantar et al., ICLR 2023; AWQ, Lin et al., MLSys 2024) mostly report aggregate measures: perplexity and accuracy on standard task suites. Fidelity to a source is measured by a separate line of methods, developed independently of model precision — checking the consistency of facts and atomic claims with the source (e.g. TRUE, Honovich et al. 2022; AlignScore, Zha et al. 2023; FActScore, Min et al. 2023); the survey by Ji et al. (2023) divides such errors into those contradicting the source and those unverifiable against it. A measurement that ties the precision level to the frequency of specific error types and ends in a rejection rule is not known to us. Success criterion. Fidelity before fluency: a variant (model and precision level) below the fidelity threshold fixed before the first measurement is out, regardless of the quality of its language. Errors are counted separately in four categories: an added claim (absent from the source), an altered figure or unit, shifted attribution (a statement attributed to someone else) and a changed degree of certainty (a possibility presented as a fact, or the reverse). The outcome is one of two, both admitted in advance: a measured rejection rule, or a documented finding that no model within the fixed compute budget is faithful enough. Method. A reference set of source documents and statements derived from them, in which the verdict on every statement is given or confirmed by a person. Several openly licensed models at several precision levels within the same compute budget, fixed in advance; fidelity first, fluency second. The measurement is repeated with every change of model. Status: the error categories and the measurement procedure are defined; the reference set does not exist yet, and nothing has been measured. Next step: building the reference set.

Corrections to the entry of 2026-08-10

The entry of 2026-08-10 presented the programme as being in the concept phase and stated that implementation had not started. That was already untrue on the day of publication: implementation work had begun in early August 2026. That entry remains unchanged because its content is stamped; the correction is dated by this entry.

Status of the entry. Neither study has a result on its criterion yet. The results — positive or negative — will appear in a further dated entry, together with the thresholds, models and data fixed before the first measurement.

Proof of time

This text was hashed and stamped on publication. Verifying it takes both files: the stamp proves that one exact byte string existed at that moment, and only the source below is that byte string. Any later edit changes the hash and breaks the match — which is the point.

· @529.58 · 2026-W40

SHA-256 of the stamped text: 897118bad9e4ede411841ff12ea0bd585ee4101763fa341bb464e47ee9a8950d

Verify with: ots verify <proof> --file <source>