5 papers
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
Aoyang Fang, Yifan Yang, Jin'ao Shang +7
Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and tool use. However, existing datasets suff…
Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration
Junjie Xu, Xingjiao Wu, Zihao Zhang +6
AI agents that plan, retain memory across sessions, invoke external tools and act with partial autonomy are transforming human--AI collaboration. Research on affective computing, s…
VoiceAgentEval: A Dual-Dimensional Benchmark for Expert-Level Intelligent Voice-Agent Evaluation of Xbench's Professional-Aligned Series
Pengyu Xu, Shijia Li, Ao Sun +15
We propose OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarios. Unlike existing methods that…
Introducing LongCat-Flash-Thinking: A Technical Report
Meituan LongCat Team, Anchun Gui, Bei Li +122
We present LongCat-Flash-Thinking, an efficient 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model. Its advanced capabilities are cultivated through a metic…
Configurable multi-agent framework for scalable and realistic testing of llm-based agents
Sai Wang, Senthilnathan Subramanian, Mudit Sahni +6
Large-language-model (LLM) agents exhibit complex, context-sensitive behaviour that quickly renders static benchmarks and ad-hoc manual testing obsolete. We present Neo, a configur…