Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
From Knowing to Acting: Benchmarking Self-Awareness Capability of LLM Agents
Yifan Li, Shengbin Yue, Boyu Feng +6
The integration of external tools has transitioned LLM agents from passive responders to autonomous systems. However, current benchmarks prioritize execution success, neglecting se…
cs.AI2026
PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning
Xiaoyou Liu, Xinyi Mou, Shengbin Yue +5
As users increasingly expect LLMs to align with their preferences, personalized information becomes valuable. However, personalized information can be a double-edged sword: it can…
cs.AI2026
Ready Jurist One: Benchmarking Language Agents for Legal Intelligence in Dynamic Environments
Zheng Jia, Shengbin Yue, Wei Chen +5
The gap between static benchmarks and the dynamic nature of real-world legal practice poses a key barrier to advancing legal intelligence. To this end, we introduce J1-ENVS, the fi…