activity
20242026
collaborators

11 papers

cs.CV2026

Human detectors are surprisingly powerful reward models

Kumar Ashutosh, XuDong Wang, Xi Yin +4

Video generation models have recently achieved impressive visual fidelity and temporal coherence. Yet, they continue to struggle with complex, non-rigid motions, especially when sy…

cs.CV2025

Learning Skill-Attributes for Transferable Assessment in Video

Kumar Ashutosh, Kristen Grauman

Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual spor…

cs.CV2025

When Thinking Drifts: Evidential Grounding for Robust Video Reasoning

Mi Luo, Zihui Xue, Alex Dimakis +1

Video reasoning, the task of enabling machines to infer from dynamic visual content through multi-step logic, is crucial for advanced AI. While the Chain-of-Thought (CoT) mechanism…

cs.HC2025

Vid2Coach: Transforming How-To Videos into Task Assistants

Mina Huh, Zihui Xue, Ujjaini Das +3

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our o…

cs.CV2025

Seeing the Arrow of Time in Large Multimodal Models

Zihui Xue, Mi Luo, Kristen Grauman

The Arrow of Time (AoT)-time's irreversible flow shaping physical events-is fundamental to video comprehension, yet remains a significant challenge for modern large multimodal mode…

cs.CV2025

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi +26

Vision-language models are integral to computer vision research, yet many high-performing models remain closed-source, obscuring their data, design and training recipe. The researc…