25 papers
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
Mingyang Wu, Kaituo Feng, Bohao Li +3
Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarci…
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Yunlong Lin, Zixu Lin, Zhaohu Xing +23
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, aud…
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…
UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization
Yaozhi Zheng, Yilei Jiang, Manyuan Zhang +5
Visual-to-Code generation, which transforms scientific plots, vector graphics, and webpages into executable scripts, demands a level of pixel-precise alignment that standard Multim…
DOPD: Dual On-policy Distillation
Xinlei Yu, Gen Li, Qingyi Si +13
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…
InterleaveThinker: Reinforcing Agentic Interleaved Generation
Dian Zheng, Harry Lee, Manyuan Zhang +4
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their archi…