Training Object Permanence in World Models

Human preference leaderboard of the 14 evaluated models on the 300-question WROP exam. Twenty raters produced 361 blind pairwise judgments between the fourteen models (50–52 per model; ties count 0.5 for each side); strengths are Bradley–Terry maximum-likelihood estimates, regularised by one virtual draw per model, rescaled to an Elo scale with mean 1500. 95% confidence intervals come from 1,000 rater-clustered bootstrap resamples; overlapping intervals should not be read as significant rank differences. Score rate is the raw win rate. Wan 3.0 Prime and MiniMax H3 have identical records and tie for first.

RankModelClassElo95% CIScore rateGames
1Wan 3.0 PrimeReference-to-video1723.6[1629.2, 1864.9]77.9%52
2MiniMax H3Reference-to-video1723.6[1644.5, 1837.6]77.9%52
3PWM-WROP (ours)True continuation1679.5[1603.5, 1781.5]73.1%52
4Seedance 2.5Reference-to-video1649.6[1554.2, 1751.5]69.6%51
5Runway Aleph 2Edit / transfer1518.3[1425.1, 1621.6]52.9%51
6Wan-VACE 14BEdit / transfer1506.7[1429.9, 1590.8]51.0%52
7Gemini Omni Flash 1.1Edit / transfer1492.5[1404.1, 1572.9]49.0%52
8Kling O3 ProEdit / transfer1471.3[1404.0, 1534.0]46.2%52
9Grok Imagine (video extend)True continuation1457.0[1362.7, 1555.1]44.2%52
10LTX-2.3 ExtendTrue continuation1453.4[1363.6, 1545.8]43.1%51
11Cosmos3 SuperEdit / transfer1409.2[1318.2, 1488.7]37.3%51
12LTX-2.3 DevEdit / transfer1398.9[1297.7, 1482.3]36.5%52
13HY-OmniWeavingEdit / transfer1268.5[1159.4, 1344.2]21.0%50
14MAGI-1 24BTrue continuation1248.0[1137.2, 1315.6]19.2%52

Answers of every model to every question: the benchmark dataset. Model weights: PWM-WROP. Method and per-family results: the paper.