6 papers
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
Zhiheng Wang, Bo Peng, Lai Wei +1
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal o…
CauTion: Knowing When to Trust LLMs for Ensemble Causal Discovery
Bo Peng, Kaiwen Wu, Sirui Chen +3
Causal discovery from observational data remains challenging due to the fundamental limitations of purely statistical methods, such as statistical distinguishability within equival…
Visual-Advantage On-Policy Distillation for Vision-Language Models
Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-p…
Turning Internal Gap into Self-Improvement: Promoting the Generation-Understanding Unification in MLLMs
Yujin Han, Hao Chen, Andi Han +5
Although unified MLLMs aim to unify generation and understanding, they are considered to exhibit an internal gap, with understanding outperforming generation. Through large-scale e…
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
Bo Peng, Zhiheng Wang, Heyang Gong +1
In modern dialogue systems, the ability to implicitly infer user backgrounds from conversations and leverage this information for personalized assistance is crucial. However, the s…
One-Shot Multilingual Font Generation Via ViT
Zhiheng Wang, Jiarui Liu
Font design poses unique challenges for logographic languages like Chinese, Japanese, and Korean (CJK), where thousands of unique characters must be individually crafted. This pape…