11 papers
SAM2Matting: Generalized Image and Video Matting
Ruiqi Shen, Guangquan Jie, Chang Liu +1
Despite impressive advances in image matting, video matting remains challenging due to the inherent gap between high-level tracking, which requires frame-wise understanding, and lo…
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
Chang Liu, Mingwen Shao, Xiang Lv +5
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing methods remain two practical challenges: (1) Th…
Masked Next-Scale Prediction for Self-supervised Scene Text Recognition
Zhuohao Chen, Zeng Li, Yifei Zhang +2
Scene Text Recognition requires modeling visual structures that evolve from coarse layouts to fine-grained character strokes. Training such models relies on large amounts of annota…
Count Anything at Any Granularity
Chang Liu, Haoning Wu, Weidi Xie
Open-world object counting remains brittle: despite rapid advances in vision-language models (VLMs), reliably counting the objects a user intends is far from solved. We argue that…
Improving Human Image Animation via Semantic Representation Alignment
Chang Liu, Mengting Chen, Yixuan Huang +5
The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long…
Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE
Yangming Shi, Shixiang Zhu, Tao Shen +14
We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the m…