Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure
Yueyang Wang, Baolong Bi, Shuo Lu +2
Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain behavior at the cost of deg…
cs.CL2026
How to Train Your Deep Research Agent? Prompt, Reward, and Policy Optimization in Search-R1
Yinuo Xu, Shuo Lu, Jianjie Cheng +5
Deep Research agents tackle knowledge-intensive tasks through multi-round retrieval and decision-oriented generation. While reinforcement learning (RL) has been shown to improve pe…
cs.CL2025
DeepResearch-Slice: Bridging the Retrieval-Utilization Gap via Explicit Text Slicing
Shuo Lu, Yinuo Xu, Jianjie Cheng +3
Deep Research agents predominantly optimize search policies to maximize retrieval probability. However, we identify a critical bottleneck: the retrieval-utilization gap, where mode…