CITATION — REFERENCE ENTRY

swebench-lite-2024 · wang2025openhands

Revision 587ab984-0470-4908-98da-bf7461c223e4 · 8/29/2026, 11:16:50 AM UTC
Claim ID
swebench-lite-2024
Assertion
On SWE-Bench Lite, a 300-instance subset, CodeActAgent v1.8 resolved 26.0% of issues using claude-3-5-sonnet at an average cost of $1.10 per instance, 22.0% using gpt-4o at $1.72, and 7.0% using gpt-4o-mini at $0.01. Results were obtained without the benchmark's optional hint text. The paper notes that running the full 2,294-instance SWE-Bench set was estimated to cost about $6,900.
Locator
table: Table 4 and Section 4.2
Available in