CITATION — REFERENCE ENTRY

benchmarks-v1 · wang2025sdk

Revision ae06e5b2-6c40-4953-a4bf-c70a24479a97 · 8/29/2026, 11:17:22 AM UTC
Citation
wang2025sdk
Claim ID
benchmarks-v1
Assertion
Self-reported results for the V1 SDK: on SWE-Bench Verified, 72.8% with Claude Sonnet 4.5, 68.8% with GPT-5 at high reasoning effort, 68.0% with Claude Sonnet 4, and 65.2% with Qwen3 Coder 480B. On the GAIA validation set, 67.9% with Claude Sonnet 4.5, 62.4% with GPT-5, 57.6% with Claude Sonnet 4, and 41.2% with Qwen3 Coder 480B. The authors state evaluations were run at a single named commit of the SDK and of the benchmark harness.
Locator
table: Table 2 and Section 5.2
Available in