From the 1 of 18 linked papers with an AI index.
18 papers
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Runyang You, Zhiyuan Liu, Yongqi Li +1
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remai…
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Hongru Cai, Yongqi Li, Ran Wei +1
PalmClaw is an open‑source framework that runs large language model agents directly on mobile phones, exposing device capabilities as explicit tools to improve task success and spe…
ToolSpec: Accelerating Tool Calling via Schema-Aware and Retrieval-Augmented Speculative Decoding
Heming Xia, Yongqi Li, Cunxiao Du +2
Tool calling has greatly expanded the practical utility of large language models (LLMs) by enabling them to interact with external applications. As LLM capabilities advance, effect…
TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
Zixu Li, Yupeng Hu, Zhiheng Fu +3
Composed Image Retrieval (CIR) is an important image retrieval paradigm that enables users to retrieve a target image using a multimodal query that consists of a reference image an…
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
Dongding Lin, Jian Wang, Yongqi Li +1
Situated conversational recommendation (SCR), which utilizes visual scenes grounded in specific environments and natural language dialogue to deliver contextually appropriate recom…
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
Hongru Cai, Yongqi Li, Tiezheng Yu +4
Alignment of Large Language Models (LLMs) aims to align outputs with human preferences, and personalized alignment further adapts models to individual users. This relies on persona…