Forecasts with measured own fallibility
A concept-phase program on machine forecasts whose first measured object is their own fallibility: calibration without model self-assessment, resolvability under not-at-random missing ground truth, mechanisms over events and fail-closed self-verification — with the entire construction withheld.
Koch Laboratory — forecasts with measured own fallibility
This program asks whether a machine-generated statement about the future can carry an honest, calibrated probability — and whether the system producing such statements can itself demonstrate that it is learning rather than merely getting better at justifying. The first-order object of measurement here is the system’s own fallibility: every forecast is to be machine-resolvable, every resolution counted, and the result compared against simple baselines before anything is called skill.
Method. As in the rest of the lab: problem → hypothesis → falsifiable criterion → method → result with boundary conditions. The status is stated openly: the concept phase is closed after external review of the foundational document, and implementation has not started — below there is not a single measured result, only formulated problems and the criteria by which the program will allow itself to be judged.
Publication note (IP, especially reduced version). We publish only the problem, the state of the art and the success criterion. Withheld are the entire solution construction, as well as the product form, the application domain and the structure of the work packages.
1. Direction 1 — calibration without model self-assessment
Research question. Can the probability attached to a forecast come exclusively from measurement — never from a language model’s own declaration — and, after a purely computational correction against measured outcomes, achieve an edge over simple baselines? Why it is hard. Agreement among samples of the same model is not accuracy: the samples share the same training priors and can be unanimously wrong; a computational correction repairs systematic overconfidence, not correlated blindness. With few events per day, calibration data are also structurally sparse. State of the art (published). Proper scoring rules, calibration curves, isotonic regression; the research record on human forecasting (including the Good Judgment Project). Success criterion. A positive skill score against three concurrently computed baselines (base rate, persistence, source consensus); the raw Brier score does not suffice, as it mostly measures how easy the chosen events were. Construction: withheld in its entirety. Status: design work.
2. Direction 2 — resolvability of forecasts under not-at-random missing ground truth
Research question. Can a forecast be fixed at creation time so that its outcome can later be determined by machine, without a fresh judgment by the model — given that a return to normal is usually not news, and absent reports do not mean absent events? Why it is hard. A model grading its own forecast after the fact is worthless — it is better at justifying than at predicting. Worse: missing data correlate with the direction of the thesis (negatively phrased theses appear to “hit”), and no downstream calibration repairs this flaw. A system that picks its own resolution date also has two escape routes: trivial proximity, or a date so distant that the resolution drowns in noise. State of the art (published). Prediction markets and superforecasting — with questions formulated and resolved by humans; model evaluation suites resolved against curated answers. Success criterion. A measurable share of forecasts resolved by machine, with undecidable cases reported openly rather than counted as hits; estimation and bounding of the bias from not-at-random missingness. Construction: withheld in its entirety. Documented correction: one early probabilistic construction proved wrong under external review and was redesigned — the error and the fix are recorded and dated; a refuted hypothesis is a document here, not an embarrassment. Status: design work.
3. Direction 3 — mechanisms, not events
Research question. Can relations between types of events — not between individual, unrepeatable events — be measured so that every relation carries its own honestly counted error rate? Why it is hard. The number of possible type pairs exceeds available observations by orders of magnitude — an ordinary significance threshold will manufacture thousands of spurious relations out of pure noise. The honest denominator (how often the effect could have been observed, not merely how often it was noticed) is hard to operationalize, and a relation without counterexamples is suspect by definition: measurement happened only where the outcome was already known. State of the art (published). Causal discovery methods, knowledge graphs, event co-occurrence databases — none of them maintains error rates or negative evidence per relation. Success criterion. A register of relations in which every entry carries its count of confirmations and refutations and is rejectable; rejected relations are preserved as a result, not deleted. Construction: withheld in its entirety. Status: design work.
4. Direction 4 — learning or rationalization: self-verification and fail-closed
Research question. How can a system working without supervision continuously tell that it is learning, rather than merely justifying ever more skilfully — when its measurement loop measures agreement with its own corpus, not with reality? Why it is hard. The most dangerous failure is not a visible error but a silent sham success: rising metrics in a system that is internally consistent yet detached from the world. The only known test that separates learning from post-hoc justification is a set of tasks with known outcomes the model cannot know from training (lying past its knowledge cutoff) — and such a resource is consumed with every model update and must be continually replenished. State of the art (published). Training-data contamination is a known problem; evaluation sets with a correct temporal cutoff exist — but they serve one-off model assessment, not the continuous self-checking of a running system. Success criterion. The fail-closed principle demonstrated in operation: a system that cannot verify itself stops publishing instead of publishing with a caveat; performance on the control tasks is distinguishable from chance. An openly admitted negative outcome: if the control remains permanently indistinguishable from chance, the program ends with a solid negative finding — “a noise generator with a well-kept surface” is also a verdict. Construction: withheld in its entirety. Status: design work.
Program status: concept phase closed (foundational document after external review), implementation not started; the next stage is a manual feasibility check on a small sample of real cases, with a negative outcome openly admitted as a possible result.