Publications (31)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
Weichen Zhang, Zile Zhou, Xin Zeng +7
Spatial reasoning is a fundamental capability of multimodal large language models (MLLMs), yet their performance in open aerial environments remains underexplored. In this work, we…
iPanda: An LLM-based Agent for Automated Conformance Testing of Communication Protocols
Xikai Sun, Fan Dang, Shiqi Jiang +8
Conformance testing is essential for ensuring that protocol implementations comply with their specifications. However, traditional testing approaches involve manually creating nume…
Self-Paced Collaborative and Adversarial Network for Unsupervised Domain Adaptation
Weichen Zhang, Dong Xu, Wanli Ouyang +1
This paper proposes a new unsupervised domain adaptation approach called Collaborative and Adversarial Network (CAN), which uses the domain-collaborative and domain-adversarial lea…
Maximum-Margin Structured Learning with Deep Networks for 3D Human Pose Estimation
Sijin Li, Weichen Zhang, Antoni B. Chan
This paper focuses on structured-output learning using deep neural networks for 3D human pose estimation from monocular images. Our network takes an image and 3D pose as inputs and…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
AirScape: An Aerial Generative World Model with Motion Controllability
Baining Zhao, Rongze Tang, Mingyuan Jia +9
How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general s…
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
Weichen Zhang, Ruiying Peng, Xin Zeng +9
3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some promising results, the advantages of p…
Towards Arbitrary Text-driven Image Manipulation via Space Alignment
Yunpeng Bai, Zihan Zhong, Chao Dong +3
The recent GAN inversion methods have been able to successfully invert the real image input to the corresponding editable latent code in StyleGAN. By combining with the language-vi…
iWorld-Bench: A Benchmark for Interactive World Models with a Unified Action Generation Framework
Jianjie Fang, Yingshan Lei, Qin Wan +8
Achieving Artificial General Intelligence (AGI) requires agents that learn and interact adaptively, with interactive world models providing scalable environments for perception, re…
Understanding and Evaluating Hallucinations in 3D Visual Language Models
Ruiying Peng, Kaiyuan Li, Weichen Zhang +3
Recently, 3D-LLMs, which combine point-cloud encoders with large models, have been proposed to tackle complex tasks in embodied intelligence and scene understanding. In addition to…
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
Weichen Zhang, Yiyou Sun, Pohao Huang +3
Hallucinations pose critical risks for large language model (LLM)-based agents, often manifesting as hallucinative actions resulting from fabricated or misinterpreted information w…
Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space
Weichen Zhang, Peizhi Tang, Xin Zeng +12
Unmanned aerial vehicles (UAVs) have emerged as powerful embodied agents. One of the core abilities is autonomous navigation in large-scale three-dimensional environments. Existing…
GUI Action Narrator: Where and When Did That Action Take Place?
Qinchen Wu, Difei Gao, Kevin Qinghong Lin +6
The advent of Multimodal LLMs has significantly enhanced image OCR recognition capabilities, making GUI automation a viable reality for increasing efficiency in digital tasks. One…
ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation
Difei Gao, Lei Ji, Zechen Bai +10
Graphical User Interface (GUI) automation holds significant promise for assisting users with complex tasks, thereby boosting human productivity. Existing works leveraging Large Lan…
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
Jianjie Fang, Yongyan Xu, Ziyou Wang +13
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, fo…
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
Baining Zhao, Jianjie Fang, Zichao Dai +8
Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a ben…
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
Weichen Zhang, Shiquan Yu, Yinan Zhu +9
We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. The benchmark decomposes active percept…
AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
Weichen Zhang, Zhui Zhu, Ningbo Li +3
Vision-language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference…
IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
Xiaoya Lu, Zeren Chen, Xuhao Hu +5
Flawed planning from VLM-driven embodied agents poses significant safety hazards, hindering their deployment in real-world household tasks. However, existing static, non-interactiv…
Model Compression using Progressive Channel Pruning
Jinyang Guo, Weichen Zhang, Wanli Ouyang +1
In this work, we propose a simple but effective channel pruning framework called Progressive Channel Pruning (PCP) to accelerate Convolutional Neural Networks (CNNs). In contrast t…
Progressive Modality Cooperation for Multi-Modality Domain Adaptation
Weichen Zhang, Dong Xu, Jing Zhang +1
In this work, we propose a new generic multi-modality domain adaptation framework called Progressive Modality Cooperation (PMC) to transfer the knowledge learned from the source do…
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
Chen Gao, Baining Zhao, Weichen Zhang +9
Embodied artificial intelligence emphasizes the role of an agent's body in generating human-like behaviors. The recent efforts on EmbodiedAI pay a lot of attention to building up m…
WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation
Baining Zhao, Jiacheng Xu, Weicheng Feng +13
Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial V…
MA-NeRF: Motion-Assisted Neural Radiance Fields for Face Synthesis from Sparse Images
Weichen Zhang, Xiang Zhou, Yukang Cao +2
We address the problem of photorealistic 3D face avatar synthesis from sparse images. Existing Parametric models for face avatar reconstruction struggle to generate details that or…
A Case for Agentic Tuning: From Documentation to Action in PostgreSQL
Hongyu Lin, Mingyu Li, Weichen Zhang +4
Documentation has long guided computer system tuning by distilling expert knowledge into per-parameter recommendations. Yet such guides capture only what experts conclude, discardi…
CityNavAgent: Aerial Vision-and-Language Navigation with Hierarchical Semantic Planning and Global Memory
Weichen Zhang, Chen Gao, Shiquan Yu +6
Aerial vision-and-language navigation (VLN), requiring drones to interpret natural language instructions and navigate complex urban environments, emerges as a critical embodied AI…
WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation
Shengtao Zheng, Kai Li, Weichen Zhang +5
End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observations to directly predict acti…
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
Shanghai AI Lab, :, Xiaoyang Chen +35
To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier…
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
Weichen Zhang, Kebin Liu, Fan Dang +3
Semantic segmentation in open-vocabulary scenarios presents significant challenges due to the wide range and granularity of semantic categories. Existing weakly-supervised methods…
Neural Latent Arbitrary Lagrangian-Eulerian Grids for Fluid-Solid Interaction
Shilong Tao, Zhe Feng, Shaohan Chen +3
Fluid-solid interaction (FSI) problems are fundamental in many scientific and engineering applications, yet effectively capturing the highly nonlinear two-way interactions remains…
How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
Baining Zhao, Ziyou Wang, Jianjie Fang +8
Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate wheth…