collaborators

5 papers

cs.CV2026

Local Margin Restoration for Test-Time Adaptation of Vision-Language Models

Yan Huang, Guowei Wang, Xu Wang +2

Vision-language models (VLMs) such as CLIP exhibit remarkable zero-shot capabilities, yet their performance frequently degrades sharply under unexpected test-time distribution shif…

cs.CV2026

PanopticQuery: Unified Query-Time Reasoning for 4D Scenes

Ruilin Tang, Yang Zhou, Zhong Ye +3

Understanding dynamic 4D environments through natural language queries requires not only accurate scene reconstruction but also robust semantic grounding across space, time, and vi…

cs.LG2026

Subspace Control: Turning Constrained Model Steering into Controllable Spectral Optimization

Yancheng Huang, Changsheng Wang, Chongyu Fan +7

Foundation models, such as large language models (LLMs), are powerful but often require customization before deployment to satisfy practical constraints such as safety, privacy, an…

cs.CV2026

VideoCoF: Unified Video Editing with Temporal Reasoner

Xiangpeng Yang, Ji Xie, Yiyuan Yang +4

Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temp…

cs.CV2025

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

Yan Huang, Yongyi Su, Xin Lin +2

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, gi…