7 papers
TIE: Time Interval Encoding for Video Generation over Events
Zhilei Shu, Shangwen Zhu, Zihang Liang +10
Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…
BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition
Yize Cai, Rui Feng, Anlan Yu +2
Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the scarcity of labeled data lim…
Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction
Han Zhang, Jinghong Mao, Shangwen Zhu +6
Diffusion reconstruction plays a critical role in various applications such as image editing, restoration, and style transfer. In theory, the reconstruction should be simple - it j…
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs
Zhantao Yang, Ruili Feng, Keyu Yan +13
Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…
RAIN: Real-time Animation of Infinite Video Stream
Zhilei Shu, Ruili Feng, Yang Cao +1
Live animation has gained immense popularity for enhancing online engagement, yet achieving high-quality, real-time, and stable animation with diffusion models remains challenging,…
Lipschitz Singularities in Diffusion Models
Zhantao Yang, Ruili Feng, Han Zhang +8
Diffusion models, which employ stochastic differential equations to sample images through integrals, have emerged as a dominant class of generative models. However, the rationality…