10 papers
DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning
Yiming Yang, Zhuoyuan Li, Fanxiang Zeng +2
Multi-agent LLM systems consistently outperform single-agent baselines, yet practitioners still cannot predict which design works for a new task or diagnose why one fails. We argue…
Macaron-A2UI: A Model for Generative UI in Personal Agents
Fancy Kong, Congjie Zheng, Murphy Zhuang +8
As personal agents evolve to handle complex, user-centric tasks, static plain-text chat is rapidly becoming a bottleneck. Generative UI emerges as the necessary new interface layer…
Benchmarking Real-Time Question Answering via Executable Code Workflows
Wenjie Zhou, Yuan Gao, Xin Zhou +5
Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and ther…
Sparse Personalized Text Generation with Multi-Trajectory Reasoning
Bo Ni, Haowei Fu, Qinwen Ge +10
As Large Language Models (LLMs) advance, personalization has become a key mechanism for tailoring outputs to individual user needs. However, most existing methods rely heavily on d…
SPAN: Benchmarking and Improving Cross-Calendar Temporal Reasoning of Large Language Models
Zhongjian Miao, Hao Fu, Chen Wei
We introduce SPAN, a cross-calendar temporal reasoning benchmark, which requires LLMs to perform intra-calendar temporal reasoning and inter-calendar temporal conversion. SPAN feat…
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
Haowei Fu, Bo Ni, Han Xu +3
Retrieval-Augmented Generation (RAG) and Supervised Finetuning (SFT) have become the predominant paradigms for equipping Large Language Models (LLMs) with external knowledge for di…