activity
20242026
collaborators

7 papers

cs.CV2026

TIE: Time Interval Encoding for Video Generation over Events

Zhilei Shu, Shangwen Zhu, Zihang Liang +10

Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…

cs.CV2026

BenchHAR: Benchmarking Self-Supervised Learning for Generalizable Sensor-based Activity Recognition

Yize Cai, Rui Feng, Anlan Yu +2

Human Activity Recognition (HAR) from wearable sensors supports broad healthcare and behavior science applications. However, data heterogeneity and the scarcity of labeled data lim…

cs.LG2025

Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction

Han Zhang, Jinghong Mao, Shangwen Zhu +6

Diffusion reconstruction plays a critical role in various applications such as image editing, restoration, and style transfer. In theory, the reconstruction should be simple - it j…

cs.CV2025

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Zhantao Yang, Ruili Feng, Keyu Yan +13

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions…

cs.CV2024

RAIN: Real-time Animation of Infinite Video Stream

Zhilei Shu, Ruili Feng, Yang Cao +1

Live animation has gained immense popularity for enhancing online engagement, yet achieving high-quality, real-time, and stable animation with diffusion models remains challenging,…

cs.CV2024

Lipschitz Singularities in Diffusion Models

Zhantao Yang, Ruili Feng, Han Zhang +8

Diffusion models, which employ stochastic differential equations to sample images through integrals, have emerged as a dominant class of generative models. However, the rationality…