Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
Yiming Ren, Zhiqiang Lin, Yu Li +8
Controllable captioning is essential for precise multimodal alignment and instruction following, yet existing models often lack fine-grained control and reliable evaluation protoco…
cs.CV2024
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
Yu Li, Yifan Chen, Gongye Liu +6
Layout generation is a foundation task of graphic design, which requires the integration of visual aesthetics and harmonious expression of content delivery. However, existing metho…