Latest news
Announcements, benchmark releases, and the work behind both.
Research & Engineering•9 min read
FrontierSWE
Our ultra-long-horizon coding benchmark. Frontier models clear only a fraction of its tasks.
Read more →
Leaderboard
| # | Model | Harness | Avg Rank¹ | Dominance² | Implementation | Performance | Research |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5Claude Code | Claude Code | 2.47 | 89% | 1.80 | 1.67 | 6.00 |
| 2 | Grok 4.5Grok CLI | Grok CLI | 4.09 | 78% | 5.70 | 3.89 | 2.00 |
| 3 | Claude Opus 4.8Claude Code | Claude Code | 4.82 | 73% | 3.60 | 5.33 | 5.33 |
| 4 | GLM-5.2Claude Code | Claude Code | 4.85 | 72% | 4.70 | 5.78 | 2.33 |
| 5 | GPT-5.5Codex | Codex | 5.21 | 70% | 7.00 | 4.00 | 5.83 |
| 6 | Claude Opus 4.7Claude Code | Claude Code | 6.47 | 61% | 5.40 | 7.00 | 6.67 |
| 7 | Claude Opus 4.6Claude Code | Claude Code | 7.59 | 53% | 7.40 | 7.78 | 7.33 |
| 8 | GPT-5.4Codex | Codex | 7.88 | 51% | 6.90 | 9.39 | 5.00 |
| 9 | Composer 2.5Cursor CLI | Cursor CLI | 9.65 | 38% | 8.00 | 11.61 | 6.50 |
| 10 | Gemini 3.1 ProGemini CLI | Gemini CLI | 9.79 | 37% | 11.80 | 7.61 | 13.00 |
| 11 | GLM-5.1Claude Code | Claude Code | 11.00 | 29% | 10.90 | 11.17 | 10.67 |
| 12 | DeepSeek V4 ProClaude Code | Claude Code | 11.18 | 27% | 11.00 | 11.33 | 11.00 |
| 13 | Kimi K2.6Kimi CLI | Kimi CLI | 11.44 | 25% | 9.50 | 12.56 | 11.33 |
| 14 | Kimi K2.5Kimi CLI | Kimi CLI | 11.50 | 25% | 12.50 | 10.22 | 13.67 |
| 15 | Qwen3.6-PlusQwen Code | Qwen Code | 12.06 | 21% | 13.80 | 10.67 | 13.33 |
¹Rank: avg position across tasks (lower = better)²Dominance: win rate vs random opponent on task
Research & Engineering•7 min read
Our Problems
An overview of the problems we’re working on at Proximal.
Read more →Company•4 min read
Announcing Proximal
We believe data is becoming one of the central research problems in AI, and no one is working on it the right way. Proximal is a research lab for data.
Read more →