collaborators

6 papers

cs.RO2026

Tactile-WAM: Touch-Aware World Action Model with Tactile Asymmetric Attention

Siyu Wu, Linjing You, Junjie Zhu +10

World Action Models (WAMs) jointly predict future visual observations and actions, but visual futures alone often miss slip, jamming, contact-direction changes, and subtle misalign…

cs.CV2026

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

Keda Tao, Yuhua Zheng, Jia Xu +13

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily fo…

cs.CV2025

OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM

Jinhong Wang, Shuo Tong, Jian liu +7

Despite the remarkable progress of multimodal large language models (MLLMs), they continue to face challenges in achieving competitive performance on ordinal regression (OR; a.k.a.…

cs.AI2025

ST-GDance: Long-Term and Collision-Free Group Choreography from Music

Jing Xu, Weiqiang Wang, Cunjian Chen +2

Group dance generation from music has broad applications in film, gaming, and animation production. However, it requires synchronizing multiple dancers while maintaining spatial co…

cs.CV2025

Scalable Autoregressive Monocular Depth Estimation

Jinhong Wang, Jian Liu, Dongqi Tang +5

This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with…

cs.CV2025

TextSleuth: Towards Explainable Tampered Text Detection

Chenfan Qu, Jian Liu, Haoxing Chen +4

Recently, tampered text detection has attracted increasing attention due to its essential role in information security. Although existing methods can detect the tampered text regio…