Coherence training for forecasts · Kelvin, MATS

Now: 32B baseline → 0.6B

Qwen3-32B first: same family and training cutoff as the 0.6B student.

Target

0.50.550.60.650.7Constant 0.50.693Form-only prior0.607Base-rate table0.578Qwen3-32B zero-shotrunning nowQwen3-0.6B + consistencynexttest log loss, lower is better
Test set: 1,066 pairs of real Manifold questions (May–Dec 2025), 12 logical forms each. A model must beat the 12-number base-rate table.

Next

Baseline result

Pending

Running now.

Table appears here.
Calibration plot appears here.
Reliability of Qwen3-32B verbalized probabilities on the test set.

Frontier traces remember, not forecast

00.050.10.15Polymarket price0.128Frontier model (Sol)0.111Brier, pre-2025 questions
On pre-2025 questions a frontier model beats the market price: it recalls outcomes. So its traces are not used for training, and all evaluation is post-cutoff.