Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Visual Language Hypothesis
Xiu Li
We study visual representation learning from a structural and topological perspective. We begin from a single hypothesis: that visual understanding presupposes a semantic language…
cs.CV2025
HERO: Hierarchical Extrapolation and Refresh for Efficient World Models
Quanjian Song, Xinyu Wang, Donghao Zhou +3
Generation-driven world models create immersive virtual environments but suffer slow inference due to the iterative nature of diffusion models. While recent advances have improved…
cs.CV2024
MultiBooth: Towards Generating All Your Concepts in an Image from Text
Chenyang Zhu, Kai Li, Yue Ma +2
This paper introduces MultiBooth, a novel and efficient technique for multi-concept customization in image generation from text. Despite the significant advancements in customized…