3 papers
cs.CV2026
SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation
Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong +1
Long-horizon video generation is evaluated with whole-frame metrics that reward motion and temporal consistency. For fixed-camera nature scenes this creates an ambiguity: motion of…
cs.CV2026
Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion
Matiur Rahman Minar, Seunghun Oh, GangHyeon Jeong +1
Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability…
cs.SD2024
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
Yunji Chu, Yunseob Shim, Unsang Park
We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensit…