SWE-ABS / ICML 2026
What does a passing benchmark really prove?
- Approach
- SWE-ABS first uses program slicing to guide tests toward untested code, then uses plausible incorrect patches to generate adversarial tests that expose semantic blind spots.
- Findings & contribution
- The study strengthened tests for 50.2% of the 500 benchmark instances. The leading evaluated agent's success rate fell from 78.80% to 62.20%, and the previously top-ranked agent moved to fifth place. These figures describe the systems and benchmark evaluated in the paper.