agentic reinforcement learning 1hierarchical reward modeling 1sandbox simulation 1tool-use agents 1travel planning 1
From the 1 of 4 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
The Hrunting of AI: Where and How to Improve English Dialectal Fairness
Wei Li, Adrian de Wynter
It is known that large language models (LLMs) underperform in English dialects, and that improving them is difficult due to data scarcity. In this work we investigate how quality a…
cs.CL2025
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
Yansong Ning, Wei Li, Jun Fang +2
Compressing long chain-of-thought (CoT) from large language models (LLMs) is an emerging strategy to improve the reasoning efficiency of LLMs. Despite its promising benefits, exist…
cs.CL2025
O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
Haotian Luo, Li Shen, Haiying He +6
Recently, long-thought reasoning LLMs, such as OpenAI's O1, adopt extended reasoning processes similar to how humans ponder over complex problems. This reasoning paradigm significa…