4 papers
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
Yu Shang, Zhuohang Li, Yiding Ma +18
While world models have emerged as a cornerstone of embodied intelligence by enabling agents to reason about environmental dynamics through action-conditioned prediction, their eva…
GR-Dexter Technical Report
Ruoshi Wen, Guangzeng Chen, Zhongren Cui +23
Vision-language-action (VLA) models have enabled language-conditioned, long-horizon robot manipulation, but most existing systems are limited to grippers. Scaling VLA policies to b…
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
Hanlei Zhang, Zhuohang Li, Yeshuang Zhu +5
Multimodal language analysis is a rapidly evolving field that leverages multiple modalities to enhance the understanding of high-level semantics underlying human conversational utt…
Transferable Learned Image Compression-Resistant Adversarial Perturbations
Yang Sui, Zhuohang Li, Ding Ding +4
Adversarial attacks can readily disrupt the image classification system, revealing the vulnerability of DNN-based recognition tasks. While existing adversarial perturbations are pr…