12 papers
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Jinbo Yan, Limeng Qiao, Jie Qin +3
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation. However, their final tokens discard fine…
HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Jinhong Deng, Limeng Qiao, Guanglu Wan
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Languag…
X2SAM: Any Segmentation in Images and Videos
Hao Wang, Limeng Qiao, Chi Zhang +4
Multimodal Large Language Models (MLLMs) have demonstrated strong image-level visual understanding and reasoning, yet their pixel-level perception across both images and videos rem…
GeoAlign: Geometric Feature Realignment for MLLM Spatial Reasoning
Zhaochen Liu, Limeng Qiao, Guanglu Wan +1
Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by i…
UniComp: Rethinking Video Compression Through Informational Uniqueness
Chao Yuan, Shimin Chen, Minliang Lin +3
Distinct from attention-based compression methods, this paper presents an information uniqueness driven video compression framework, termed UniComp, which aims to maximize the info…
X-SAM: From Segment Anything to Any Segmentation
Hao Wang, Limeng Qiao, Zequn Jie +6
Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although…