From the 1 of 6 linked papers with an AI index.
6 papers
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin
StoryTeller is a training‑free framework that adds a persistent narrative memory to video‑language models so they can generate coherent, story‑aware audio descriptions for long vid…
Learning Video Dynamics with Predictive Differentiable Rendering
Yujin Tang, Tian Zhou, Xin Lin +5
How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video prediction models operate in discrete pixel sp…
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
Wayner Barrios, Andrés Villa, Juan León Alcázar +2
Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs)…
A Foundation Model for Wearable Movement Data in Mental Health Research
Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh +2
Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavi…
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
Wayner Barrios, SouYoung Jin
We introduce CRYSTAL (Clear Reasoning via Yielded Steps, Traceability, and Logic), a diagnostic benchmark with 6,372 instances that evaluates multimodal reasoning through verifiabl…
FT2TF: First-Person Statement Text-To-Talking Face Generation
Xingjian Diao, Ming Cheng, Wayner Barrios +1
Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatar…