CITATION — REFERENCE ENTRY

benchmarks-v1 · wang2025sdk

Revision 7a33edc9-8749-4e9f-ac1e-6d5312aa3082 · 8/29/2026, 12:12:18 PM UTC
Citation
wang2025sdk
Claim ID
benchmarks-v1
Assertion
Self-reported results for the V1 SDK: on SWE-Bench Verified, 72.8% with Claude Sonnet 4.5, 68.8% with GPT-5 at high reasoning effort, 68.0% with Claude Sonnet 4, and 65.2% with Qwen3 Coder 480B. On the GAIA validation set, 67.9% with Claude Sonnet 4.5, 62.4% with GPT-5, 57.6% with Claude Sonnet 4, and 41.2% with Qwen3 Coder 480B. The authors state evaluations were run at a single named commit of the SDK and of the benchmark harness.
Quote
SDK achieves 72% resolution rate using Claude Sonnet 4.5 with extended thinking ... SDK achieves 67.9% accuracy with Claude Sonnet 4.5 ... Evaluations were performed at commit 54c5858 of the SDK and commit 88f1d80 of the benchmarks
Quote language
en
Locator
table: Table 2 and Section 5.2
Available in