2 papers
cs.CV2025
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
Yuxuan Wu, Le Wang, Sanping Zhou +3
Controllable layout generation aims to create plausible visual arrangements of element bounding boxes within a graphic design according to certain optional constraints, such as the…
cs.CV2025
Moment Quantization for Video Temporal Grounding
Xiaolong Sun, Le Wang, Sanping Zhou +5
Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishin…