6 papers
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
Haozhe Qi, Kevin Qu, Mahdi Rad +3
Long video understanding remains challenging for Multi-modal Large Language Models (MLLMs) due to high memory costs and context-length limits. Prior approaches mitigate this by sco…
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
Kevin Qu, Haozhe Qi, Mihai Dusmanu +3
Multimodal Large Language Models (MLLMs) have made impressive progress in connecting vision and language, but they still struggle with spatial understanding and viewpoint-aware rea…
LLaVAction: evaluating and training multi-modal large language models for action understanding
Haozhe Qi, Shaokai Ye, Alexander Mathis +1
Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multim…
Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
Shengkai Hu, Haozhe Qi, Jun Wan +4
Recent advances in deep learning have significantly improved facial landmark detection. However, existing facial landmark detection datasets often define different numbers of landm…
EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models
Andy Bonnetto, Haozhe Qi, Franklin Leong +7
Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive funct…
MammAlps: A multi-view video behavior monitoring dataset of wild mammals in the Swiss Alps
Valentin Gabeff, Haozhe Qi, Brendan Flaherty +3
Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors en…