3 papers
cs.AI2026
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
Xiao Ma, Zhiquan Hu, Yi Wei +6
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups provide no comparative signal and may hi…
cs.CV2026
Thought Graph Traversal for Test-time Scaling in Chest X-ray VLLMs
Yue Yao, Zelin Wen, Yan Tong +5
Test-time scaling offers a promising way to improve the reasoning performance of vision-language large models (VLLMs) without additional training. In this paper, we explore a simpl…
cs.CL2026
PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents
Minjia Wang, Yunfeng Wang, Xiao Ma +9
Digital footprints (records of individuals' interactions with digital systems) are essential for studying behavior, developing personalized applications, and training machine learn…