works on

From the 1 of 6 linked papers with an AI index.

collaborators

6 papers

cs.CV2026

StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description

Seung Hyun Hahm, Minh T. Dinh, SouYoung Jin

StoryTeller is a training‑free framework that adds a persistent narrative memory to video‑language models so they can generate coherent, story‑aware audio descriptions for long vid…

cs.CV2026

Learning Video Dynamics with Predictive Differentiable Rendering

Yujin Tang, Tian Zhou, Xin Lin +5

How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video prediction models operate in discrete pixel sp…

cs.CV2026

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs

Wayner Barrios, Andrés Villa, Juan León Alcázar +2

Multimodal Large Language Models (MLLMs) have achieved remarkable success in instruction-following tasks by integrating pretrained visual encoders with large language models (LLMs)…

cs.LG2026

A Foundation Model for Wearable Movement Data in Mental Health Research

Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh +2

Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavi…

cs.AI2026

Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation

Wayner Barrios, SouYoung Jin

We introduce CRYSTAL (Clear Reasoning via Yielded Steps, Traceability, and Logic), a diagnostic benchmark with 6,372 instances that evaluates multimodal reasoning through verifiabl…

cs.CV2024

FT2TF: First-Person Statement Text-To-Talking Face Generation

Xingjian Diao, Ming Cheng, Wayner Barrios +1

Talking face generation has gained immense popularity in the computer vision community, with various applications including AR, VR, teleconferencing, digital assistants, and avatar…