9 papers · 1 filter
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
Alexey Gavryushin, Dingxi Zhang, Zhao Huang +8
Human-human collaboration is a fundamental aspect of everyday life, essential to success in a wide range of goal-directed activities from household tasks to professional teamwork.…
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
Jiuming Liu, Chaojun Ni, Mengmeng Liu +7
With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream do…
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
Marcel Gröpl, Jaewoo Jung, Seungryong Kim +2
Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread across multiple regions, as in docu…
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
Haozhe Qi, Kevin Qu, Mahdi Rad +3
Long video understanding remains challenging for Multi-modal Large Language Models (MLLMs) due to high memory costs and context-length limits. Prior approaches mitigate this by sco…
CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modeling
Sayan Deb Sarkar, Rémi Pautrat, Ondrej Miksik +4
Video Language Models (VideoLMs) enable AI systems to understand temporal dynamics in videos. To fit within the maximum context window constraint, current methods use keyframe samp…
Training-free Detection and 6D Pose Estimation of Unseen Surgical Instruments
Jonas Hein, Lilian Calvet, Matthias Seibold +3
Purpose: Accurate detection and 6D pose estimation of surgical instruments are crucial for many computer-assisted interventions. However, supervised methods lack flexibility for ne…