12 papers
DPEPO: Diverse Parallel Exploration Policy Optimization for LLM-based Agents
Junshuo Zhang, Chengrui Huang, Feng Guo +6
Large language model (LLM) agents that follow the sequential "reason-then-act" paradigm have achieved superior performance in many complex tasks.However, these methods suffer from…
HTAA: Enhancing LLM Planning via Hybrid Toolset Agentization & Adaptation
Chengrui Huang, Junshuo Zhang, Zhiyuan Ma +7
Enabling large language models to scale and reliably use hundreds of tools is critical for real-world applications, yet challenging due to the inefficiency and error accumulation i…
FAVE: Flow-based Average Velocity Establishment for Sequential Recommendation
Ke Shi, Yao Zhang, Feng Guo +4
Generative recommendation has emerged as a transformative paradigm for capturing the dynamic evolution of user intents in sequential recommendation. While flow-based methods improv…
PACE: Prefix-Protected and Difficulty-Aware Compression for Efficient Reasoning
Ruixiang Feng, Yuntao Wen, Silin Zhou +14
Language Reasoning Models (LRMs) achieve strong performance by scaling test-time computation but often suffer from ``overthinking'', producing excessively long reasoning traces tha…
TOOL4POI: A Tool-Augmented LLM Framework for Next POI Recommendation
Dongsheng Wang, Shen Gao, Chengrui Huang +3
Next Point-of-Interest (POI) recommendation is a fundamental task in location-based services. While recent advances leverage Large Language Model (LLM) for sequential modeling, exi…
CoSineVerifier: Tool-Augmented Answer Verification for Computation-Oriented Scientific Questions
Ruixiang Feng, Zhenwei An, Yuntao Wen +9
Answer verification methods are widely employed in language model training pipelines spanning data curation, evaluation, and reinforcement learning with verifiable rewards (RLVR).…