works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.LG2026

Context-Informed Ship Trajectory Prediction via Conditional Attention

Yuan Guan, Chandler Squires, Timothy Hu +1

The paper introduces the Conditional Informer, a Transformer-based model that predicts ship trajectories by explicitly conditioning vessel states on environmental contexts using a…

cs.CV2026

OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing

Zebang Cheng, Shuimu Chen, Boxue Yang +7

Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounce…

cs.CV2026

Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding

Shuimu Chen, Yuteng Chen, Yuanshen Guan +7

Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal parameters. Lacking objective external evide…

cs.LG2026

Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate

Duncan Soiffer, Chandler Squires, Yuan Guan +2

The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to produce samples from compositionally-def…

cs.CV2025

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

Shen Zheng, Jiaran Cai, Yuansheng Guan +7

Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or re…

cs.CV2025

Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback

Xingpei Ma, Shenneng Huang, Jiaran Cai +5

Recent advances in diffusion models have significantly improved audio-driven human video generation, surpassing traditional methods in both quality and controllability. However, ex…