activity
20242026
collaborators

7 papers

cs.CV2026

HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models

Jiajun Wu, Haoyu Kang, Yining Sun +13

Large vision-language models (LVLMs) have recently shown immense potential in automated content moderation, sparking growing interest in developing harmful-video benchmarks. Howeve…

cs.CV2026

Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models

Yuxiao Chen, Jue Wang, Zhikang Zhang +8

With recent advancements in video backbone architectures, combined with the remarkable achievements of large language models (LLMs), the analysis of long-form videos spanning tens…

cs.CV2026

LAB-Det: Language as a Domain-Invariant Bridge for Training-Free One-Shot Domain Generalization in Object Detection

Xu Zhang, Zhe Chen, Jing Zhang +1

Foundation object detectors such as GLIP and Grounding DINO excel on general-domain data but often degrade in specialized and data-scarce settings like underwater imagery or indust…

cs.CV2026

DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding

Moulik Choraria, Xinbo Wu, Akhil Bhimaraju +5

Hyperscaling of data and parameter count in LLMs is yielding diminishing improvement when weighed against training costs, underlining a growing need for more efficient finetuning a…

cs.MM2025

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

Liuyue Xie, Avik Kuthiala, George Z. Wei +12

We introduce MAVERIX (Multimodal audiovisual Evaluation and Recognition IndeX), a unified benchmark to probe the video understanding in multimodal LLMs, encompassing video, audio,…

cs.CV2025

Achieving More with Less: Additive Prompt Tuning for Rehearsal-Free Class-Incremental Learning

Haoran Chen, Ping Wang, Zihan Zhou +3

Class-incremental learning (CIL) enables models to learn new classes progressively while preserving knowledge of previously learned ones. Recent advances in this field have shifted…