1 citations · 1 across the 38 of their papers we have counts for
36 papers · 1 filter
Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis
Liang Xu, Chengqun Yang, Zili Lin +6
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approache…
MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
Wenjie Zhu, Yabin Zhang, Wenjun Zeng +1
Multimodal Large Language Models (MLLMs) have achieved strong performance on a wide range of vision-language tasks, but often fail under imperfect or shifted contexts. A reliable M…
Bridging 3D Gaussians and Semantic Occupancy for Comprehensive Open-Vocabulary Scene Understanding from Unposed Images
Hu Zhu, Bohan Li, Xianda Guo +5
Comprehensive 3D scene understanding from sparse, unposed images requires a model to recover renderable geometry, open-vocabulary semantics, and free/occupied 3D space without rely…
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
Yuyang Zhang, Wenyao Zhang, Zekun Qi +7
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense mult…
Dual Distribution Estimation for Zero-shot Noisy Test-Time Adaptation with VLMs
Wenjie Zhu, Yabin Zhang, Liang Xu +3
While test-time adaptation (TTA) empowers vision-language models to adapt without costly retraining, it remains highly vulnerable to out-of-distribution (OOD) outliers prevalent in…
An Efficient Streaming Video Understanding Framework with Agentic Control
Jinming Liu, Jianguo Huang, Zhaoyang Jia +7
Streaming video requires handling dynamic information density under strict latency budgets. Yet, existing methods typically employ static strategies, such as fixed memory compressi…