CITATION — REFERENCE ENTRY
benchmarks-v1 · wang2025sdk
- Citation
- wang2025sdk
- Claim ID
- benchmarks-v1
- Assertion
- Self-reported results for the V1 SDK: on SWE-Bench Verified, 72.8% with Claude Sonnet 4.5, 68.8% with GPT-5 at high reasoning effort, 68.0% with Claude Sonnet 4, and 65.2% with Qwen3 Coder 480B. On the GAIA validation set, 67.9% with Claude Sonnet 4.5, 62.4% with GPT-5, 57.6% with Claude Sonnet 4, and 41.2% with Qwen3 Coder 480B. The authors state evaluations were run at a single named commit of the SDK and of the benchmark harness.
- Quote
SDK achieves 72% resolution rate using Claude Sonnet 4.5 with extended thinking ... SDK achieves 67.9% accuracy with Claude Sonnet 4.5 ... Evaluations were performed at commit 54c5858 of the SDK and commit 88f1d80 of the benchmarks
- Quote language
- en
- Locator
- table: Table 2 and Section 5.2
Available in