Human preference leaderboard of the 14 evaluated models on the 300-question WROP exam. Twenty raters produced 361 blind pairwise judgments between the fourteen models (50–52 per model; ties count 0.5 for each side); strengths are Bradley–Terry maximum-likelihood estimates, regularised by one virtual draw per model, rescaled to an Elo scale with mean 1500. 95% confidence intervals come from 1,000 rater-clustered bootstrap resamples; overlapping intervals should not be read as significant rank differences. Score rate is the raw win rate. Wan 3.0 Prime and MiniMax H3 have identical records and tie for first.
| Rank | Model | Class | Elo | Score rate |
|---|---|---|---|---|
| 1 | Wan 3.0 Prime | 1723.6 | 77.9% | |
| 2 | MiniMax H3 | 1723.6 | 77.9% | |
| 3 | PWM-WROP (ours) | 1679.5 | 73.1% | |
| 4 | Seedance 2.5 | 1649.6 | 69.6% | |
| 5 | Runway Aleph 2 | 1518.3 | 52.9% | |
| 6 | Wan-VACE 14B | 1506.7 | 51.0% | |
| 7 | Gemini Omni Flash 1.1 | 1492.5 | 49.0% | |
| 8 | Kling O3 Pro | 1471.3 | 46.2% | |
| 9 | Grok Imagine (video extend) | 1457.0 | 44.2% | |
| 10 | LTX-2.3 Extend | 1453.4 | 43.1% | |
| 11 | Cosmos3 Super | 1409.2 | 37.3% | |
| 12 | LTX-2.3 Dev | 1398.9 | 36.5% | |
| 13 | HY-OmniWeaving | 1268.5 | 21.0% | |
| 14 | MAGI-1 24B | 1248.0 | 19.2% |
Answers of every model to every question: the benchmark dataset. Model weights: PWM-WROP. Method and per-family results: the paper.