From the 1 of 19 linked papers with an AI index.
19 papers
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Runyang You, Zhiyuan Liu, Yongqi Li +1
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remai…
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Hongru Cai, Yongqi Li, Ran Wei +1
PalmClaw is an open‑source framework that runs large language model agents directly on mobile phones, exposing device capabilities as explicit tools to improve task success and spe…
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
Qiancheng Xu, Yongqi Li, Fan Liu +3
Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools during reasoning. Existing TIR meth…
Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Prompting
Heming Xia, Cunxiao Du, Rui Li +3
Large reasoning models (LRMs) have demonstrated remarkable proficiency in tackling complex tasks through step-by-step thinking. However, this lengthy reasoning process incurs subst…
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
Dongding Lin, Jian Wang, Yongqi Li +1
Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recom…
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
Hongru Cai, Yongqi Li, Tiezheng Yu +4
Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individual users. This relies on persona…