CITATION — REFERENCE ENTRY

swebench-lite-2024 · wang2025openhands

Revision e97e2ee7-2f7c-4814-8cef-3c40ec478fa1 · 8/29/2026, 12:11:53 PM UTC
Claim ID
swebench-lite-2024
Assertion
On SWE-Bench Lite, a 300-instance subset, CodeActAgent v1.8 resolved 26.0% of issues using claude-3-5-sonnet at an average cost of $1.10 per instance, 22.0% using gpt-4o at $1.72, and 7.0% using gpt-4o-mini at $0.01. Results were obtained without the benchmark's optional hint text. The paper notes that running the full 2,294-instance SWE-Bench set was estimated to cost about $6,900.
Quote
CodeActAgent v1.8, using claude-3.5-sonnet, achieves a competitive resolve rate of 26%
Quote language
en
Locator
table: Table 4 and Section 4.2
Available in