3 papers
cs.CV2026
Efficient Multimodal Large Language Models: A Survey
Yizhang Jin, Jian Li, Yexin Liu +10
In the past year, Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance in tasks such as visual question answering, visual understanding and reasoning.…
cs.CV2025
PointSeg: A Training-Free Paradigm for 3D Scene Segmentation via Foundation Models
Qingdong He, Jinlong Peng, Zhengkai Jiang +2
Recent success of vision foundation models have shown promising performance for the 2D perception tasks. However, it is difficult to train a 3D foundation network directly due to t…
cs.CV2025
OSV: One Step is Enough for High-Quality Image to Video Generation
Xiaofeng Mao, Zhengkai Jiang, Fu-Yun Wang +5
Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to su…