News & updates
Announcements from the research archive.
New on the site: recent preprints on RAG hallucination detection, conversational agent testing, code specification alignment, and causal root-cause analysis. See the updated publication list for all recent work.
Our ICML and ICML Position papers were accepted: "SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark" and "How Should We Build A Benchmark? Revisiting 274 Code-Related Benchmarks For LLMs".
Our paper "UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench" was accepted by ACL'2025.
Our extended abstract "DSPy Guardrails: Building Safe LLM Applications via Self-Refining Language Model Pipelines" was accepted by Compound AI Systems Workshop (June 13th, 2024 in San Francisco at Data + AI Summit).
Our paper "Testing Graph Database Systems via Equivalent Query Rewriting" was accepted by ICSE'2024.
We introduce “Retromorphic Testing,” a new, general methodology to the test oracle problem. It is a black-box technique, which constructs a dual program architecture to test the target software, inspired by the concept of inverse function. Read the paper
Our paper "Deep Learning or Classical Machine Learning? An Empirical Study on Log-Based Anomaly Detection" was accepted by ICSE'2024.
Our paper "Automated Testing and Improvement of Named Entity Recognition Systems" was accepted by ESEC/FSE'2023.
Our paper "ROME: Testing Image Captioning Systems via Recursive Object Melting" was accepted by ISSTA'2023.