4 papers
Latent Action Control for Reasoning-Guided Unified Image Generation
Fuxiang Zhai, Sixiang Chen, Yingjin Li +4
Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may…
Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning
Rui Song, Lida Shi, Ruihua Qi +2
In recent years, rapid advances in Multimodal Large Language Models (MLLMs) have increasingly stimulated research on ancient Chinese scripts. As the evolution of written characters…
Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction
Juncheng Hu, Zijian Zhang, Zeyu Wang +3
Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on im…
Shortcut Learning in In-Context Learning: A Survey
Rui Song, Yingji Li, Lida Shi +2
Shortcut learning refers to the phenomenon where models employ simple, non-robust decision rules in practical tasks, which hinders their generalization and robustness. With the rap…