7 papers
HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks
Jingyu Guo, Ziye Chen, Ziwen Li +7
Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descriptions with goal-centric evaluat…
AnisoLift: Anisotropic Latent Representations for Coarse Particle Liquid Enhancement
Zhengqing Gao, Huaxi Huang, Runqi Lin +6
Particle-based liquid simulation is widely used in graphics and physical modeling, but high-resolution rollouts remain computationally expensive. Consequently, many methods aim to…
FLUIDSPLAT: Reconstructing Physical Fields from Sparse Sensors via Gaussian Primitives
Huaxi Huang, Meng Li, Zhengqing Gao +3
Reconstructing continuous flow fields from sparse surface-mounted sensors is central to aerodynamic design, flow control, and digital-twin instrumentation. Existing neural methods…
KineVLA: Towards Kinematics-Aware Vision-Language-Action Models with Bi-Level Action Decomposition
Gaoge Han, Zhengqing Gao, Ziwen Li +5
In this paper, we introduce a novel kinematics-rich vision-language-action (VLA) task, in which language commands densely encode diverse kinematic attributes (such as direction, tr…
Mirage2Matter: A Physically Grounded Gaussian World Model from Video
Zhengqing Gao, Ziwen Li, Xin Wang +12
The scalability of embodied intelligence is fundamentally constrained by the scarcity of real-world interaction data. While simulation platforms provide a promising alternative, ex…
MLLM-For3D: Adapting Multimodal Large Language Model for 3D Reasoning Segmentation
Jiaxin Huang, Runnan Chen, Ziwen Li +5
Reasoning segmentation aims to segment target objects in complex scenes based on human intent and spatial reasoning. While recent multimodal large language models (MLLMs) have demo…