CITATION — REFERENCE ENTRY
benchmarks-v1 · wang2025sdk
- Citation
- wang2025sdk
- Claim ID
- benchmarks-v1
- Assertion
- Self-reported results for the V1 SDK: on SWE-Bench Verified, 72.8% with Claude Sonnet 4.5, 68.8% with GPT-5 at high reasoning effort, 68.0% with Claude Sonnet 4, and 65.2% with Qwen3 Coder 480B. On the GAIA validation set, 67.9% with Claude Sonnet 4.5, 62.4% with GPT-5, 57.6% with Claude Sonnet 4, and 41.2% with Qwen3 Coder 480B. The authors state evaluations were run at a single named commit of the SDK and of the benchmark harness.
- Locator
- table: Table 2 and Section 5.2
Available in