1 citations · 1 across the 11 of their papers we have counts for
4 papers · 1 filter
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
Martin Q. Ma, Willis Guo, Aditya Agrawal +4
Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…
Act2See: Emergent Active Visual Perception for Video Reasoning
Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4
Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning p…
PoseBench3D: A Cross-Dataset Analysis Framework for 3D Human Pose Estimation via Pose Lifting Networks
Saad Manzur, Bryan Vela, Brandon Vela +4
Reliable three-dimensional human pose estimation (3D HPE) remains challenging due to the differences in viewpoints, environments, and camera conventions among datasets. As a result…
Latent Representation Matters: Human-like Sketches in One-shot Drawing Tasks
Victor Boutin, Rishav Mukherji, Aditya Agrawal +4
Humans can effortlessly draw new categories from a single exemplar, a feat that has long posed a challenge for generative models. However, this gap has started to close with recent…