5 papers
Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan +11
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to acti…
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
Hongzhu Yi, Yujia Yang, Yuanxiang Wang +18
In recent years, image editing models have made significant progress, enabling users to manipulate visual content in a flexible and interactive manner through natural language inst…
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
Yujia Yang, Yuanxiang Wang, Zhenyu Guan +11
While Instruction-based Image Editing (IIE) has achieved significant progress, existing benchmarks pursue task breadth via mixed evaluations. This paradigm obscures a critical fail…
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
Chen Bao, Chuanbing Huo, Qinyu Chen +1
This paper proposes AS-ASR, a lightweight aphasia-specific speech recognition framework based on Whisper-tiny, tailored for low-resource deployment on edge devices. Our approach in…
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
Chen Bao, Jiarui Xu, Xiaolong Wang +2
How can we predict future interaction trajectories of human hands in a scene given high-level colloquial task specifications in the form of natural language? In this paper, we exte…