activity
20222026
most citedOpen-vocabulary Semantic Segmentation with Frozen Vision-Language Models

10 citations · 10 across the 4 of their papers we have counts for

collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification

Wujian Peng, Lingchen Meng, Yuxuan Cai +7

Unified Multimodal Modeling aims to integrate visual understanding and generation within a single system. However, existing approaches typically rely on two disparate visual tokeni…

cs.CV2026

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning

Chaofan Ma, Zhenjie Mao, Yuhuan Yang +5

Spatial reasoning from egocentric videos is inherently challenging because the observable evidence is constrained by the camera trajectory. Existing methods rely on single-turn inf…

cs.CV2026

GenMask: Adapting DiT for Segmentation via Direct Mask Generation

Yuhuan Yang, Xianwei Zhuang, Yuxuan Cai +6

Recent approaches for segmentation have leveraged pretrained generative models as feature extractors, treating segmentation as a downstream adaptation task via indirect feature ret…

cs.CV2025

SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation

Zhenjie Mao, Yuhuan Yang, Chaofan Ma +4

Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and…

cs.CV2025

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

Yuhuan Yang, Chaofan Ma, Zhenjie Mao +3

Video understanding is a complex challenge that requires effective modeling of spatial-temporal dynamics. With the success of image foundation models (IFMs) in image understanding,…

cs.CV202210 cited

Open-vocabulary Semantic Segmentation with Frozen Vision-Language Models

Chaofan Ma, Yuhuan Yang, Yanfeng Wang +2

When trained at a sufficient scale, self-supervised learning has exhibited a notable ability to solve a wide range of visual or language understanding tasks. In this paper, we inve…