most citedKangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

1 citations · 1 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CV2025

MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation

Jia Wang, Jie Hu, Xiaoqi Ma +3

Text-to-image (T2I) generation has achieved remarkable progress in instruction following and aesthetics. However, a persistent challenge is the prevalence of physical artifacts, su…

cs.CV2025

Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes

Huijie Liu, Bingcan Wang, Jie Hu +2

Dish images play a crucial role in the digital era, with the demand for culturally distinctive dish images continuously increasing due to the digitization of the food industry and…

cs.GR2025

Image Editing with Diffusion Models: A Survey

Jia Wang, Jie Hu, Xiaoqi Ma +3

With deeper exploration of diffusion model, developments in the field of image generation have triggered a boom in image creation. As the quality of base-model generated images con…

cs.CV2025

Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

Huijie Liu, Jingyun Wang, Shuai Ma +3

Motion customization aims to adapt the diffusion model (DM) to generate videos with the motion specified by a set of video clips with the same motion concept. To realize this goal,…

cs.CV2024

RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction

Xiaoping Wu, Jie Hu, Xiaoming Wei

Diffusion Probabilistic Models (DPMs) have emerged as the de facto approach for high-fidelity image synthesis, operating diffusion processes on continuous VAE latent, which signifi…

cs.CV2024

CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

Lichen Ma, Tiezhu Yue, Pei Fu +4

Recently, significant advancements have been made in diffusion-based visual text generation models. Although the effectiveness of these methods in visual text rendering is rapidly…