collaborators

6 papers

cs.CV2026

Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

Valentin Gabeff, Baptiste Maquignaz, Jennifer Shan +5

Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large video-language models (VLMs) have potentia…

cs.CV2026

AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding

Haozhe Qi, Kevin Qu, Mahdi Rad +3

Long video understanding remains challenging for Multi-modal Large Language Models (MLLMs) due to high memory costs and context-length limits. Prior approaches mitigate this by sco…

cs.CV2026

LLaVAction: evaluating and training multi-modal large language models for action understanding

Haozhe Qi, Shaokai Ye, Alexander Mathis +1

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multim…

cs.CV2025

EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

Andy Bonnetto, Haozhe Qi, Franklin Leong +7

Understanding behavior requires datasets that capture humans while carrying out complex tasks. The kitchen is an excellent environment for assessing human motor and cognitive funct…

cs.CV2025

MammAlps: A multi-view video behavior monitoring dataset of wild mammals in the Swiss Alps

Valentin Gabeff, Haozhe Qi, Brendan Flaherty +3

Monitoring wildlife is essential for ecology and ethology, especially in light of the increasing human impact on ecosystems. Camera traps have emerged as habitat-centric sensors en…

cs.CL2025

PICLe: Pseudo-Annotations for In-Context Learning in Low-Resource Named Entity Detection

Sepideh Mamooler, Syrielle Montariol, Alexander Mathis +1

In-context learning (ICL) enables Large Language Models (LLMs) to perform tasks using few demonstrations, facilitating task adaptation when labeled examples are hard to obtain. How…