3 papers
cs.CV2026
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models
Xiaojun Shan, Haoyu Shen, Yucheng Mao +5
We present CyCLeGen, a unified vision-language foundation model capable of both image understanding and image generation within a single autoregressive framework. Unlike existing v…
cs.LG2025
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
Xiaojun Shan, Qi Cao, Xing Han +2
Recent advances in multimodal foundation models have achieved state-of-the-art performance across a range of tasks. These breakthroughs are largely driven by new pre-training parad…
cs.CV2025
Visual Acoustic Fields
Yuelei Li, Hyunjin Kim, Fangneng Zhan +7
Objects produce different sounds when hit, and humans can intuitively infer how an object might sound based on its appearance and material properties. Inspired by this intuition, w…