activity
20242026
collaborators

5 papers

cs.CV2026

Latent Action Control for Reasoning-Guided Unified Image Generation

Fuxiang Zhai, Sixiang Chen, Yingjin Li +4

Unified multimodal models can encode visual understanding and image generation within a shared backbone, yet understanding does not automatically translate into control: models may…

cs.CL2026

Enhancing Multimodal Large Language Models for Ancient Chinese Character Evolution Analysis via Glyph-Driven Fine-Tuning

Rui Song, Lida Shi, Ruihua Qi +2

In recent years, rapid advances in Multimodal Large Language Models (MLLMs) have increasingly stimulated research on ancient Chinese scripts. As the evolution of written characters…

cs.CV2025

Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction

Juncheng Hu, Zijian Zhang, Zeyu Wang +3

Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on im…

cs.CL2024

Shortcut Learning in In-Context Learning: A Survey

Rui Song, Yingji Li, Lida Shi +2

Shortcut learning refers to the phenomenon where models employ simple, non-robust decision rules in practical tasks, which hinders their generalization and robustness. With the rap…

cs.LG2024

Fusion Matrix Prompt Enhanced Self-Attention Spatial-Temporal Interactive Traffic Forecasting Framework

Mu Liu, MingChen Sun YingJi Li, Ying Wang

Recently, spatial-temporal forecasting technology has been rapidly developed due to the increasing demand for traffic management and travel planning. However, existing traffic fore…