collaborators

25 papers

cs.CV2026

AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward

Mingyang Wu, Kaituo Feng, Bohao Li +3

Detailed audio-video joint captioning is essential for multimodal video understanding and generation. However, prior works are constrained by three main limitations: (1) the scarci…

cs.CV2026

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Yunlong Lin, Zixu Lin, Zhaohu Xing +23

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, aud…

cs.AI2026

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15

Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…

cs.CV2026

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

Yaozhi Zheng, Yilei Jiang, Manyuan Zhang +5

Visual-to-Code generation, which transforms scientific plots, vector graphics, and webpages into executable scripts, demands a level of pixel-precise alignment that standard Multim…

cs.AI2026

DOPD: Dual On-policy Distillation

Xinlei Yu, Gen Li, Qingyi Si +13

On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sourc…

cs.CV2026

InterleaveThinker: Reinforcing Agentic Interleaved Generation

Dian Zheng, Harry Lee, Manyuan Zhang +4

Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their archi…