7 papers
APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL
Bowen Cao, Weibin Liao, Yushi Sun +3
Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The primary limitation lies in their reli…
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
Zhiqin Yang, Yonggang Zhang, Wei Xue +3
Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implem…
LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment
Shipeng Li, Zhiqin Yang, Shikun Li +7
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing LLMs' reasoning abilities, yet its data inefficiency remains a major bottleneck. To a…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
Binyan Xu, Dong Fang, Haitao Li +1
Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through workflows spanning stages, tools, shared state, verification, and repair. Distill…
FELA: A Multi-Agent Evolutionary System for Feature Engineering of Industrial Event Log Data
Kun Ouyang, Haoyu Wang, Dong Fang
Event log data, recording fine-grained user actions and system events, represent one of the most valuable assets for modern digital services. However, the complexity and heterogene…
Beyond Playtesting: A Generative Multi-Agent Simulation System for Massively Multiplayer Online Games
Ran Zhang, Kun Ouyang, Tiancheng Ma +2
Optimizing numerical systems and mechanism design is crucial for enhancing player experience in Massively Multiplayer Online (MMO) games. Traditional optimization approaches rely o…