Qwen3-32B first: same family and training cutoff as the 0.6B student.
Target
Test set: 1,066 pairs of real Manifold questions (May–Dec 2025), 12 logical forms each. A model must beat the 12-number base-rate table.
Next
1Qwen3-32B zero-shot score (below).
2Train Qwen3-0.6B, with and without consistency, to beat it.
3Check the gain is coherence, not shrinkage to the prior.
Baseline result
Pending
Running now.
Table appears here.
Calibration plot appears here.
Reliability of Qwen3-32B verbalized probabilities on the test set.
Frontier traces remember, not forecast
On pre-2025 questions a frontier model beats the market price: it recalls outcomes. So its traces are not used for training, and all evaluation is post-cutoff.