7 papers
AffineTok: Semantic Affine Consistency for Diffusion-Friendly Visual Tokenizer
Junqiu Yu, Pandeng Li, Yikai Wang +11
Visual tokenizers increasingly inject semantic supervision into latent spaces to make downstream diffusion easier. Yet how these semantics should be organized to facilitate denoisi…
Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
Shuochen Chang, Tong Bai, Xiaofeng Zhang +5
Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over explicit Chain-of-Thought (CoT…
Bottleneck Tokens for Unified Multimodal Retrieval
Siyu Sun, Jing Ren, Zhaohe Liao +8
Adapting decoder-only multimodal large language models (MLLMs) for unified multimodal retrieval faces two structural gaps. First, existing methods rely on implicit pooling, which o…
AIBench: Evaluating Visual-Logical Consistency in Academic Illustration Generation
Zhaohe Liao, Kaixun Jiang, Zhihang Liu +11
Although image generation has boosted various applications via its rapid evolution, whether the state-of-the-art models are able to produce ready-to-use academic illustrations for…
ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
Zhihang Liu, Xiaoyi Bao, Pandeng Li +7
While existing generation and unified models excel at general image generation, they struggle with tasks requiring deep reasoning, planning, and precise data-to-visual mapping abil…
AnimateScene: Camera-controllable Animation in Any Scene
Qingyang Liu, Bingjie Gao, Weiheng Huang +10
Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plaus…