6 papers · 1 filter
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Aniri, Jinhe Bi, Peng Liao +5
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language models (MLLMs). Existing methods draw pr…
Edit the Bits, Diff the Codes: Bitwise Residual Editing for Visual Autoregressive Models
Shengqiang Zhang, Ruotong Liao, Volker Tresp +2
Text-guided image editing with visual autoregressive (VAR) generators requires controlling both what the model samples and where the sampled change is written back into the image c…
TunerDiT: Training-free Progressive Steering of Diffusion Transformer for Multi-Event Video Generation
Ruotong Liao, Guowen Huang, Qing Cheng +6
Text-to-video (T2V) generation faces challenging questions when generating videos with long horizons containing multiple events. Inspired by the intrinsics of the diffusion process…
ReEXplore: Improving MLLMs for Embodied Exploration with Contextualized Retrospective Experience Replay
Gengyuan Zhang, Mingcong Ding, Jingpei Wu +2
Embodied exploration is a target-driven process that requires embodied agents to possess fine-grained perception and knowledge-enhanced decision making. While recent attempts lever…
When and Where do Events Switch in Multi-Event Video Generation?
Ruotong Liao, Guowen Huang, Qing Cheng +3
Text-to-video (T2V) generation has surged in response to challenging questions, especially when a long video must depict multiple sequential events with temporal coherence and cont…
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
Gengyuan Zhang, Tanveer Hannan, Hermine Kleiner +6
An ideal vision-language agent serves as a bridge between the human users and their surrounding physical world in real-world applications like autonomous driving and embodied agent…