What Is Elo Scoring?
Elo originates from chess and was adapted by Artificial Analysis for AI model evaluation:
- Two model outputs are blindly compared
- Judges (human or AI) pick the better output
- Elo ratings adjust based on wins/losses
Elo reflects relative strength, not an absolute score. 1269 means Seedance 2.0 wins significantly more head-to-head matchups against other models.
Current Rankings (February 2026)
| Rank | Model | Elo |
|---|---|---|
| 1 | Seedance 2.0 | 1269 |
| 2 | Google Veo 3 | 1241 |
| 3 | OpenAI Sora 2 | 1228 |
| 4 | Runway Gen-4.5 | 1205 |
| 5 | Kling 2.0 | 1198 |
Where Seedance 2.0 Wins
1. AV Sync Is a Step Change
In blind Elo tests, videos with native audio almost always beat silent video. Seedance 2.0 is among the few mainstream models consistently outputting high-quality synchronized AV.
2. Long-Horizon Consistency
In 15-second multishot output, character and scene consistency clearly beats competitors limited to 4–8 second clips. Judges ask «which feels like a complete piece?»
3. Physical Realism
In prompts with motion, collision, and fluids, Seedance 2.0 wins most often—judges say it «looks actually filmed.»
4. Instruction Precision
On complex prompts, Seedance 2.0 less often «ignores half the instructions.»
Elo Limitations
Elo is not perfect:
- Subjectivity: Blind judges have aesthetic bias
- Prompt bias: Test sets may favor certain models
- Not speed/cost: Elo scores quality only
Industry Meaning
Seedance 2.0 topping Elo signals a new phase: competition shifts from «who generates video» to «who generates complete audiovisual works.» For developers and creators, multimodal joint generation is the direction forward.