22 papers
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
Zhenkun Gao, Yicheng Bao, Jinlong Peng +13
Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimo…
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
Tingyang Chen, Shuo Lu, Kang Zhao +11
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet to…
GSCompleter: A Distillation-Free Plugin for Metric-Aware 3D Gaussian Splatting Completion in Seconds
Ao Gao, Jingyu Gong, Xin Tan +3
3D Gaussian Splatting (3DGS) has revolutionized high-fidelity neural rendering with its explicit representation and efficiency. However, reconstructing scenes from sparse viewpoint…
NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
Zhihao Luo, Wentao Yan, Jingyu Gong +5
Recent advances in Graphical User Interface (GUI) and embodied navigation have driven progress, yet these domains have largely evolved in isolation, with disparate datasets and tra…
World2Minecraft: Occupancy-Driven Simulated Scenes Construction
Lechao Zhang, Haoran Xu, Jingyu Gong +3
Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limi…
ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards
Wentao Yan, Shengqin Wang, Huichi Zhou +4
Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outcome-based supervision and th…