Coherence training for forecasts · Kelvin, MATS

First training run

Qwen3-0.6B on 12 Boolean forms per question pair: with vs without a consistency penalty, one label vs all labels.

Void. A prompt bug dropped the multiple-choice option from 79% of prompts (fixed 23 Sep). Numbers kept as a record.

Test log loss by epoch

0.58 0.62 0.66 0.70 0.74 epoch 1 epoch 2 epoch 3 epoch 4 test log loss (lower is better) base-rate table 0.578 Boolean prior 0.607 A2 epoch 1: 0.598 A2 epoch 2: 0.598 A2 epoch 3: 0.610 A3 epoch 1: 0.595 A3 epoch 2: 0.598 A3 epoch 3: 0.605 A4 epoch 1: 0.726 A4 epoch 2: 0.665 A4 epoch 3: 0.634 A4 epoch 4: 0.733 A5 epoch 1: 0.701 A5 epoch 2: 0.642 A5 epoch 3: 0.614
1,066 test pairs. No arm beats the 12-number base-rate table. Penalty λ = 1 on the one-label arm.

Shrinkage, not coherence

0 0.03 0.06 0.09 0.12 λ = 00.097 λ = 0.30.081 λ = 10.064 λ = 30.046 spread (SD) of p(AND) across pairs, one-label arms
The penalty pulls p(AND) toward one value as λ grows. p(P) and p(¬P) still correlate +0.92 to +0.96: no negation learned.