3 papers
cs.CV2026
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
Bo Zhang, Wenxin Wang, Feng Chen +4
Recent advancements in MLLM-based long-form video understanding have mitigated inference-time computational cost and limited context lengths by selecting query-relevant frames. How…
cs.CV2026
Agentic Collaborative Cognition for Zero-Shot 3D Understanding
Wenxin Wang, Bo Zhang, Feng Chen +4
Recent advancements have explored agentic zero-shot 3D understanding by reformulating it as video keyframe understanding with Multimodal Large Language Models (MLLMs). However, exi…
cs.CV2025
OmniGen2: Towards Instruction-Aligned Multimodal Generation
Chenyuan Wu, Pengfei Zheng, Ruiran Yan +19
In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, imag…