46 citations · 59 across the 10 of their papers we have counts for
10 papers
Sora Generates Videos with Stunning Geometrical Consistency
Xuanyi Li, Daquan Zhou, Chenxu Zhang +3
The recently developed Sora model [1] has exhibited remarkable capabilities in video generation, sparking intense discussions regarding its ability to simulate real-world phenomena…
Fast Window-Based Event Denoising with Spatiotemporal Correlation Enhancement
Huachen Fang, Jinjian Wu, Qibin Hou +2
Previous deep learning-based event denoising methods mostly suffer from poor interpretability and difficulty in real-time processing due to their complex architecture designs. In t…
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
Senmao Li, Joost van de Weijer, Taihang Hu +4
The success of recent text-to-image diffusion models is largely due to their capacity to be guided by a complex text prompt, which enables users to precisely describe the desired c…
ChatAnything: Facetime Chat with LLM-Enhanced Personas
Yilin Zhao, Xinbin Yuan, Shanghua Gao +4
In this technical report, we target generating anthropomorphized personas for LLM-based characters in an online manner, including visual appearance, personality and tones, with onl…
MaskDiffusion: Boosting Text-to-Image Consistency with Conditional Mask
Yupeng Zhou, Daquan Zhou, Zuo-Liang Zhu +3
Recent advancements in diffusion models have showcased their impressive capacity to generate visually striking images. Nevertheless, ensuring a close match between the generated im…
Delving Deeper into Data Scaling in Masked Image Modeling
Cheng-Ze Lu, Xiaojie Jin, Qibin Hou +3
Understanding whether self-supervised learning methods can scale with unlimited data is crucial for training large-scale models. In this work, we conduct an empirical study on the…