22 papers
LegalWorld: A Life-Cycle Interactive Environment for Legal Agents
Songhan Zuo, Shengbin Yue, Tao Chiang +4
Civil litigation is inherently a life-cycle process: what a lawyer drafts on day one constrains what unfolds at trial months later. Yet existing legal benchmarks evaluate isolated…
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng +6
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
Sizhe Wang, Feiyu Duan, Juelin Wang +2
Large Language Models (LLMs) have reshaped user profiling, yet current evaluations mainly focus on static data snapshots. This paradigm overlooks the reality of personalized system…
HyLaT: Efficient Multi-Agent Communication via Hybrid Latent-Text Protocol
Xinyi Mou, Siyuan Wang, Zejun Li +2
Communication protocol design is a central challenge in large language model-based multi-agent systems. Existing single-channel approaches face an inherent communication trilemma:…
ACE: Self-Evolving LLM Coding Framework via Adversarial Unit Test Generation and Preference Optimization
Yixu Huang, Xinglei Yu, Zhongyu Wei
Large Language Models (LLMs) excel at code generation but remain heavily reliant on large-scale annotated solutions and verification-based supervision, which constrains scalability…
Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation
Zilin Lu, Ruifeng Yuan, Weiwei Cao +6
Radiologists highly desire fully automated AI for radiology report generation (R2G), yet existing approaches fall short in clinical utility. Reinforcement learning (RL) holds poten…