47 papers
StructRL: Structured Action-Space Exploration for Flow-Based VLAs
Jiarui Yang, Bin Zhu, Jingjing Chen +4
Flow-based Vision-Language-Action (VLA) models are now widely used for continuous robotic manipulation, and online reinforcement learning (RL) is emerging as a key technique for ad…
ReGraph: Learning to Generate Recipe Graphs from Food Images
Guoshan Liu, Bin Zhu, Pengkun Jiao +3
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which in…
Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation
Yaole Wang, Xiaoyu Chen, Xin Ma +5
Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing me…
SphereVideo: Prototype-anchored Hyperspherical Boundary for Continual AI-generated Video Detection
Fei Li, Yue Yu, Yuran Wang +3
AI-generated video (AIGV) detection aims to distinguish real videos from AI-generated ones. In practice, detectors trained on existing data often fail to generalize to newly emergi…
DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection
Zihao Cai, Xinghan Li, Ruiyan Yang +3
The paper introduces DECODE, a framework that addresses both representation and decision-level forgetting in continual learning for AI-generated image detection, using subspace div…
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
Pengkun Jiao, Bin Zhu, Jingjing Chen +1
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…