8 papers
PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets
Jie Huang, Xiaohe Li, Jiahao Li +6
Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset f…
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Xiaohe Li, Yiru Wang, Junhao Fan +6
Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a…
Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models
Xiaohe Li, Jiahao Li, Kaixin Zhang +5
While Multimodal Large Language Models (MLLMs) excel in general vision-language tasks, their application to remote sensing change understanding is hindered by a fundamental "tempor…
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Kaixin zhang, Xiaohe Li, Jiahao Li +4
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal cau…
FAST: Flexible and Adaptive Semantic Transmission for Resource-constrained Multi-user Generative Semantic Communication
Yiru Wang, Wanting Yang, Fangli Mou +4
The rapid advancement of generative artificial intelligence has spurred innovative approaches to semantic communication, giving rise to a new paradigm known as generative semantic…
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
Xiaohe Li, Pengfei Li, Zide Fan +4
Multi-view multi-object tracking (MVMOT) has found widespread applications in intelligent transportation, surveillance systems, and urban management. However, existing studies rare…