4 papers · 1 filter
NoisEasier: Test-Time Noise Optimization for Text-to-Video Generation
Yujiang Pu, Yu Kong
Diffusion models have recently advanced text-to-video (T2V) generation, yet they still struggle with fine-grained compositional alignment, such as attribute binding, spatial relati…
Procedural Mistake Detection via Action Effect Modeling
Wenliang Guo, Yujiang Pu, Yu Kong
Mistake detection in procedural tasks is essential for building intelligent systems that support learning and task execution. Existing approaches primarily analyze how an action is…
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
Yujiang Pu, Zhanbo Huang, Vishnu Boddeti +1
Generating visual instructions in a given context is essential for developing interactive world simulators. While prior works address this problem through either text-guided image…
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding
Zixu Cheng, Yujiang Pu, Shaogang Gong +2
Temporal grounding, also known as video moment retrieval, aims at locating video segments corresponding to a given query sentence. The compositional nature of natural language enab…