4 papers
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
Yi Li, Hongze Shen, Lexiang Tang +6
Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending these models to fine-grained dense predic…
LensWalk: Agentic Video Understanding by Planning How You See in Videos
Keliang Li, Yansong Li, Hongze Shen +3
The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understan…
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
Keliang Li, Hongze Shen, Hao Shi +9
The aspiration for artificial general intelligence, fueled by the rapid progress of multimodal models, demands human-comparable performance across diverse environments. We propose…
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
Keliang Li, Zaifei Yang, Jiahe Zhao +5
The significant advancements in visual understanding and instruction following from Multimodal Large Language Models (MLLMs) have opened up more possibilities for broader applicati…