9 papers
No Action Without a NOD: A Heterogeneous Multi-Agent Architecture for Reliable Service Agents
Zixu Yang, Hang Zheng, Nan Jiang +5
Large language model (LLM) agents have increasingly advanced service applications, such as booking flight tickets. However, these service agents suffer from unreliability in long-h…
Internalizing Curriculum Judgment for LLM Reinforcement Fine-Tuning
Han Zheng, Yining Ma, Karthick Gunasekaran +4
In LLM Reinforcement Fine-Tuning (RFT), curriculum learning drives both efficiency and performance. Yet, current methods externalize curriculum judgment via handcrafted heuristics…
Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents
Rong Shan, Te Gao, Hang Zheng +6
The implicit policy of maintaining relatively stable acceptance rates at top AI conferences, despite exponentially growing submissions, introduces a critical structural vulnerabili…
CapsID: Soft-Routed Variable-Length Semantic IDs for Generative Recommendation
Wenzhuo Cheng, Menghang Gong, Qixin Guo +4
Generative recommendation maps each item to a sequence of Semantic IDs (SIDs) and recasts retrieval as autoregressive token generation. In this paradigm the main bottleneck is the…
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
Tiancheng Huang, Ruisheng Cao, Yuxin Zhang +8
The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are…
DiSRouter: Distributed Self-Routing for LLM Selections
Hang Zheng, Hongshen Xu, Yongkai Lin +3
The proliferation of Large Language Models (LLMs) has created a diverse ecosystem of models with highly varying performance and costs, necessitating effective query routing to bala…