7 papers
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…
DOPD: Dual On-policy Distillation
Xinlei Yu, Gen Li, Qingyi Si +13
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…
Vero: An Open RL Recipe for General Visual Reasoning
Gabriel Sarch, Linrong Cai, Qunzhong Wang +3
What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest vision-language models (VLMs) suggest tha…
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
Qunzhong Wang, Jie Liu, Jiajun Liang +7
Recent advancements in multimodal reward models (RMs) have substantially improved post-training for visual generative models. However, current RMs face inherent limitations: (1) vi…
QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
Yiliu Yang, Yilei Jiang, Qunzhong Wang +5
Safety risks arise as large language model-based agents solve complex tasks with tools, multi-step plans, and inter-agent messages. However, deployer-written policies in natural la…
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
Yilei Jiang, Yaozhi Zheng, Yuxuan Wan +4
Automating the transformation of user interface (UI) designs into front-end code holds significant promise for accelerating software development and democratizing design workflows.…