15 papers
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents
Qi Liu, Yiqun Chen, Zidan Chen +6
Search agents now answer questions that take dozens of searches to settle, yet how such an agent reads a page has drawn far less attention than how it finds one. Nearly all of them…
UniNote: A Unified Embedding Model for Multimodal Representation and Ranking
Jinghan Zhao, Wenwei Jin, Anqi Li +5
Item-to-Item (I2I) retrieval is a fundamental part of modern content platforms, supporting critical industrial workflows from recommendation engines to content auditing. While mult…
Deep Research as Rubric for Reinforcement Learning
Wangyi Mei, Zhouhong Gu, Zhenhan Bai +9
Open-ended reasoning and long-form generation tasks lack reliable automatic verification signals for reward-based policy optimization. Rubrics offer a promising alternative, but ex…
Preference-Aware Rubric Learning for Personalized Evaluation
Yilun Qiu, Xiaoyan Zhao, Yang Zhang +7
As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model behavior with individual prefere…
AgentCVR: Active Multi-Agent Cross-Video Reasoning via Script-Simulated Reinforcement Learning
Yilun Qiu, Jiahe Wang, Cilin Yan +4
Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple v…
Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling
Xinglin Wang, Hao Lin, Shaoxiong Feng +9
Test-Time Scaling (TTS) enhances the reasoning capabilities of large language models by allocating additional inference compute to explore the solution space. However, existing par…