9 papers
ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
Muzhi Zhu, Hao Zhong, Canyu Zhao +9
Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Zheng Huang, Mingyu Liu, Xiaoyi Lin +9
Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic fo…
Exploring Spatial Intelligence from a Generative Perspective
Muzhi Zhu, Shunyao Jiang, Huanyi Zheng +9
Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern gener…
Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
Zekai Luo, Zongze Du, Zhouhang Zhu +7
Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a signific…
LLaDA2.1: Speeding Up Text Diffusion via Token Editing
Tiwei Bie, Maosong Cao, Xiang Cao +47
While LLaDA2.0 showcased the scaling potential of 100B-level block-diffusion models and their inherent parallelization, the delicate equilibrium between decoding speed and generati…
SE360: Semantic Edit in 360 Panoramas via Hierarchical Data Construction
Haoyi Zhong, Fang-Lue Zhang, Andrew Chalmers +1
While instruction-based image editing is emerging, extending it to 360 panoramas introduces additional challenges. Existing methods often produce implausible results in bot…