12 papers · 1 filter
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
Linpeng Huang, Weixing Chen, Zexin Chen +2
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are predomin…
PhyScene3D: Physically Consistent Interactive 3D Tabletop Scene Generation
Weixing Chen, Zhuoqian Feng, Yang Liu +6
Generating physically consistent 3D tabletop scenes is a fundamental yet underexplored problem for interactive and generalist robotic learning. The challenge stems from dense objec…
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
Jinzhou Tang, Sidi Liu, Waikit Xiu +2
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critica…
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
Shicheng Yin, Kaixuan Yin, Weixing Chen +3
World models are essential for autonomous robotic planning. However, the substantial computational overhead of existing dense Transformerbased models significantly hinders real-tim…
DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
Shicheng Yin, Kaixuan Yin, Yang Liu +2
The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
Yang Liu, Weixing Chen, Yongjie Bai +4
Embodied Artificial Intelligence (Embodied AI) is crucial for achieving Artificial General Intelligence (AGI) and serves as a foundation for various applications (e.g., intelligent…