3 papers
cs.CV2026
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
Peng Sun, Jun Xie, Tao Lin
Unified Multimodal Models (UMMs) are often constrained by the pre-training of their , which typically relies on inefficient paradigms and sca…
cs.CV2025
Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
Kuo Wang, Quanlong Zheng, Junlin Xie +6
Video Multimodal Large Language Models~(Video-MLLM) have achieved remarkable advancements in video understanding tasks. However, constrained by the context length limitation in the…
cs.CV2025
H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding
Qi Wu, Quanlong Zheng, Yanhao Zhang +8
With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating…