activity
20232026
most citedGraph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language Models

1 citations · 1 across the 38 of their papers we have counts for

collaborators
Showing cs.CVShow all

36 papers · 1 filter

cs.CV2026

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Liang Xu, Chengqun Yang, Zili Lin +6

The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approache…

cs.CV2026

MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models

Wenjie Zhu, Yabin Zhang, Wenjun Zeng +1

Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable M…

cs.CV2026

Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images

Hu Zhu, Bohan Li, Xianda Guo +5

Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…

cs.CV2026

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?

Yuyang Zhang, Wenyao Zhang, Zekun Qi +7

World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense mult…

cs.CV2026

Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs

Wenjie Zhu, Yabin Zhang, Liang Xu +3

While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in…

cs.CV2026

An Efficient Streaming Video Understanding Framework with Agentic Control

Jinming Liu, Jianguo Huang, Zhaoyang Jia +7

Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, such as fixed memory compressi…