collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE

Yangming Shi, Shixiang Zhu, Tao Shen +14

We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the m…

cs.CV2026

TIR-Flow: Active Video Search and Reasoning with Frozen VLMs

Hongbo Jin, Siyi Xie, Jiayu Ding +2

While Large Video-Language Models (Video-LLMs) have achieved remarkable progress in perception, their reasoning capabilities remain a bottleneck. Existing solutions typically resor…

cs.CV2025

VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition

Hongbo Jin, Kuanwei Lin, Wenhao Zhang +2

Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or…

cs.CV2025

MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation

Gengliang Li, Rongyu Chen, Bin Li +2

Ensuring factual consistency and reliable reasoning remains a critical challenge for medical vision-language models. We introduce MEDFACT-R1, a two-stage framework that integrates…

cs.CV2024

AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production

Jiuniu Wang, Zehua Du, Yuyuan Zhao +9

The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on S…