5 papers
Escaping the Self-Repair Trap: Improving Test Oracle Generation via Dual-Context Awareness
Kefan Li, Hongyue Yu, Yuan Yuan
Large Language Models (LLMs) have shown strong potential for regression-oracle completion, where a test prefix is given and the current program version is treated as expected behav…
TARSE: Test-Time Adaptation via Retrieval of Skills and Experience for Reasoning Agents
Junda Wang, Zonghai Tao, Hansi Zeng +3
Complex clinical decision making often fails not because a model lacks facts, but because it cannot reliably select and apply the right procedural knowledge and the right prior exa…
CoCoEvo: Co-Evolution of Programs and Test Cases to Enhance Code Generation
Kefan Li, Yuan Yuan, Hongyue Yu +2
Large Language Models (LLMs) have shown remarkable performance in automated code generation. However, existing approaches often rely heavily on pre-defined test cases, which become…
Data Extraction Attacks in Retrieval-Augmented Generation via Backdoors
Yuefeng Peng, Junda Wang, Hong Yu +1
Despite significant advancements, large language models (LLMs) still struggle with providing accurate answers when lacking domain-specific or up-to-date knowledge. Retrieval-Augmen…
CRTRE: Causal Rule Generation with Target Trial Emulation Framework
Junda Wang, Weijian Li, Han Wang +5
Causal inference and model interpretability are gaining increasing attention, particularly in the biomedical domain. Despite recent advance, decorrelating features in nonlinear env…