From the 1 of 7 linked papers with an AI index.
7 papers
Context-Informed Ship Trajectory Prediction via Conditional Attention
Yuan Guan, Chandler Squires, Timothy Hu +1
The paper introduces the Conditional Informer, a Transformer-based model that predicts ship trajectories by explicitly conditioning vessel states on environmental contexts using a…
OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing
Zebang Cheng, Shuimu Chen, Boxue Yang +7
Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounce…
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Shuimu Chen, Yuteng Chen, Yuanshen Guan +7
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evide…
Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate
Duncan Soiffer, Chandler Squires, Yuan Guan +2
The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-def…
High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer
Shen Zheng, Jiaran Cai, Yuansheng Guan +7
Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or re…
Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback
Xingpei Ma, Shenneng Huang, Jiaran Cai +5
Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, ex…