SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Coverage- and mutation-driven test augmentation rejected 19.71% of previously passing patches in our evaluation of 30 agents on SWE-Bench Verified.
Senior Research Fellow · Lero, Ireland
I study how to test and evaluate software and AI systems, with a focus on coding agents, trustworthy AI, and automated testing.
I received my Ph.D. from The Chinese University of Hong Kong, Shenzhen in 2025, advised by Prof. Pinjia He.

Strengthening tests and examining the reliability of benchmark results.
Using relations between programs and executions to detect inconsistencies.
Controlled transformations reveal errors in vision and language systems.
LightAD, OpenRCA 2.0, and CLEANet examine model efficiency, causal diagnosis, and robustness to contaminated data.
New on the site: recent preprints on RAG hallucination detection, conversational agent testing, code specification alignment, and causal root-cause analysis. See the updated publication list for all recent work.
Our ICML and ICML Position papers were accepted: "SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark" and "How Should We Build A Benchmark? Revisiting 274 Code-Related Benchmarks For LLMs".
Our paper "UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench" was accepted by ACL'2025.
Coverage- and mutation-driven test augmentation rejected 19.71% of previously passing patches in our evaluation of 30 agents on SWE-Bench Verified.
Generated unit tests exposed 345 erroneous patches that had originally passed SWE-Bench evaluation, showing how test quality changes conclusions about coding agents.
Retromorphic testing with claim-level, local-to-global verification identifies unsupported or contradictory claims in RAG answers and links them to answer spans and context-side evidence.
A dual-program approach to constructing test oracles: transform an input, map the output back to the input domain, and check the resulting relation.
Object removal and inpainting create natural-looking image pairs for testing the consistency of image captions.
Metamorphic relations across similar contexts and entities support automated testing and improvement of named entity recognition systems.
An empirical comparison of classical and deep learning approaches examines accuracy and computational cost across five log datasets.
I contribute to research through program committees, peer review, artifact evaluation, student supervision, and teaching.
SANER 2025 program committee and ICSE 2025 shadow research track committee; artifact evaluation for OSDI, USENIX ATC, OOPSLA, and ISSTA.
Full service recordI co-supervise PhD student Zhiqing Zhong(September 2023-present). My teaching experience includes software engineering, algorithms, and introductory programming at CUHK-Shenzhen.