2024-10 · o1 reasoning models
Launch: Claude 3.5 Sonnet (new)
o1 reasoning models · Launch: Claude 3.5 Sonnet (new) · Scores are era-relative (month top ≈95–100) for landscape review — not official historical Artificial Analysis replays. From 2026-08, the latest month aligns with the live AA-style snapshot.
Score method:Scores are era-relative (month top ≈95–100) for landscape review — not official historical Artificial Analysis replays. From 2026-08, the latest month aligns with the live AA-style snapshot.
Month landscape board
| # | Model | Vendor | Era score | Elo≈ | Pricing | Highlight |
|---|---|---|---|---|---|---|
| 🥇 | o1-preview | OpenAI | 99 | 1330 | 推理溢价 | 数学/竞赛跃迁 |
| 🥈 | Claude 3.5 SonnetNEW | Anthropic | 95 | 1280 | 中 | 日常编码仍强 |
| 🥉 | GPT-4o | OpenAI | 93 | 1270 | 中 | 默认体验 |
| 4 | o1-mini | OpenAI | 90 | 1290 | 中低 | 便宜推理 |
Elo≈ is a LMSYS-style discussion-scale estimate for recap — not an official historical snapshot.
Launches this month
- Claude 3.5 Sonnet (new)Anthropicapi
Computer use 等
Pricing changes
- Claude 3.5 Sonnet · 新版同价档 · approx. $3/$15 /1M
能力提升、价稳
Benchmark context
Benchmark context: public discussion of Arena Elo, AA Intelligence, and open suites (GPQA / SWE-bench, etc.). This page is a narrative archive, not a live probe.
Some historical narratives remain bilingual; model names stay in original form.
Related tools
Clavue CLI / imux IDE / clavue-2.1
Clavue · CLI 与 Agent 平台
CLI / GUI Agent 运行时、设备登录、会员额度与 OpenAI 兼容 API(api.clavue.com)。开发脚本、CI 与本地工具一条链路。
打开 Clavue ↗imux · 原生 macOS AI IDE
Clavue 平台的原生 IDE:多 Agent 分屏、Ghostty 级终端、Agent Chat、浏览器自动化与 Supervisor——不是又一个 Electron 壳。
了解 imux ↗Clavue 2.1 · 产品级大模型
旗舰模型 clavue-2.1(及 fast / pro / rev):适合复杂推理与长程 Agent 循环。Chat、imux、API 共用会员额度。
查看模型 ↗