RESEARCH
MAHJ-Bench.
A controlled benchmark of model reasoning under hidden state, uncertainty, and strict numerical constraints.
MAHJ-Bench separates whether a model responds, chooses the correct action, and completes the full numerical decision contract.
- Model configurations
- 57
- Task IDs
- 100
- Capability categories
- 10
- Terminal outcomes
- 4,813
93 unique prompts · 5,700 planned model–task slots · Completion-adjusted analysis
SELECTED FINDINGS
Correct action is not a complete answer.
Across 5,700 planned model–task evaluations, 62.8% select the exact action, but only 44.0% satisfy the full decision contract.
EXACT ACTION → STRICT COMPLETE
- Exact action
- 62.8%
- Strict complete
- 44.0%
18.8 percentage-point completion gap.
1,072 otherwise-correct responses fail on schema, probability, or value requirements.
§5.1 · Action and complete correctness4,813 / 5,700 planned slots
3,579 / 5,700 planned slots
2,507 / 5,700 planned slots
CONFIDENCE / ACCURACY
Confidence exceeds accuracy.
A 14.8 percentage-point gap on the same 4,423 eligible predictions. Of 845 wrong actions with usable confidence, 697 report at least 90% confidence.
§6.2 · Confidence calibrationCAPABILITY CONCENTRATION
Strict-complete performance remains constrained across capabilities.
- Reliability / consistency
- 29.3%
- Policy adaptation
- 32.1%
- Long-horizon planning
- 32.3%
- Strategic plasticity
- 32.5%
No capability exceeds 61.2% strict completion. Four of ten remain below one-third strict-complete performance.
Table 4 · Strict completeness by capabilityCAPABILITY PROFILES
Performance diverges sharply by capability.
Strict-complete performance spans 31.9 percentage points across the ten benchmark capabilities—from 29.3% to 61.2%. Reliability, adaptation, long-horizon planning, and strategic plasticity remain particularly difficult.
Correct action, valid full schema, and correct probability and value contracts.
- Hidden-state reasoning42.1%240 / 570
- Uncertainty intelligence50.7%289 / 570
- Policy adaptation32.1%183 / 570
- Opponent modeling42.1%240 / 570
- Long-horizon planning32.3%184 / 570
- Reliability / consistency29.3%167 / 570
- Recovery intelligence61.2%349 / 570
- Information acquisition57.2%326 / 570
- Counterfactual reasoning60.4%344 / 570
- Strategic plasticity32.5%185 / 570
Categories describe this finite task suite. Reliability / consistency has three unique prompts; one is repeated eight times. Scores also reflect missing outcomes and response compliance.
Table 4 · Original capability slicesMETHOD & SCOPE
Bounded tasks. Explicit state. Reproducible scoring.
Controlled decision problems
Each task specifies hidden state, observable evidence, admissible actions, and an explicit utility model. Reference probabilities and values are calculated from the supplied model.
A four-part decision contract
Each response must contain an action, probability estimates, value estimate, and confidence. Scoring separates response delivery, exact action correctness, and strict numerical completeness.
Coverage is part of the result
4,813 of 5,700 planned evaluations produced terminal outcomes. Missing responses are reported separately from reasoning errors.
FINITE EXPECTED-UTILITY SELECTION
The study evaluates static numerical microdecisions rather than full-game or live-agent interaction. Results reflect the archived execution configurations and controls. Aggregate tables and figures are reproducible from the release; full scoring replay requires the underlying execution artifacts. Read the limitations ↗
RESEARCH ARTIFACTS
Aggregate download includes source tables and supporting figures.
EVIDENCE RELEASE / Request execution files