6 papers
Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents
Liting Lin, Boxi Yu, Yuzhong Zhang +3
Conversational LLM agents can cause real-world harm when their internal workflows fail, such as completing a transaction without confirmation. Testing these state-dependent failure…
OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models
Shijie Cao, Qingyu Zhang, Boxi Yu +6
Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate large token sequences under audio-…
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
Aoyang Fang, Yifan Yang, Jin'ao Shang +7
Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suff…
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
Jiasheng Zheng, Boxi Cao, Boxi Yu +6
Reinforcement Learning with Verifiable Rewards (RLVR) has recently emerged as the cornerstone for shaping the remarkable coding abilities of Large Language Models (LLMs). However,…
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
Boxi Yu, Yuzhong Zhang, Liting Lin +2
Large language models can still hallucinate in retrieval-augmented generation (RAG), producing claims that are unsupported by or conflict with the retrieved context. Detecting such…
SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark
Boxi Yu, Yang Cao, Yuzhong Zhang +9
The SWE-Bench Verified leaderboard is approaching saturation, with the top system achieving 78.80%. However, we show that this performance is inflated. Our re-evaluation reveals th…