Asked only "A and B (and C)", Qwen3-32B says rare combinations are too likely. Fitting those answers is worse than just multiplying its singles (0.340 vs 0.298) and hurts the singles. Asked a sign-balanced set, the fit beats independence on joints (0.393 vs 0.480) and repairs the singles (0.668 → 0.602).
839triangles
6weeks
16questions each
$4.63cost
16 questions per triangle
One real triangle, week of 23 June 2025. Triangles are three events GPT-6 Luna linked to each other (it picks up to 5 related events per event). Each event is randomly asked as YES or NO, so no answer type always favours "too likely".
The bias follows rarity, not wording
Joint and count questions are mostly rare, so they are over-stated; the OR mirrors are common, so under-stated.
no NOT in the statementone or more NOTsdot size = number of statements
Every question type, split by how many NOTs it has. Rare statements are said too likely, common ones too unlikely, whatever the words. A YES single (true 22%) is +0.20; a NO single (true 80%) is −0.19.
The raw answers break the rules
1.55the 4 counts add to (should be 1)
18%order checks fail (triple ≤ pair ≤ single)
79%triangles with at least one failure
average error, 95% CItypical size of the error (mean |error|)
NO and YES singles agree on average but are 0.22 off in a typical triangle. The fit below turns these into one coherent set of probabilities per triangle.
Result: the balanced design lets the fit help
raw answersfitted to all answersmodel’s singles multiplied (independence)
AND-only: held-out joints and all singles over 35 weeks (weekly world model). Sign-balanced: the same model, 839 triangles in 6 of those weeks, joints scored on the balanced forms, YES singles only. Compare within a design; the question sets differ. Fitted = the closest coherent probabilities to all answers. Correcting each type for its bias adds nothing more.