4 papers
PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets
Jie Huang, Xiaohe Li, Jiahao Li +6
Simulation-ready 3D assets are central to robotics and embodied AI. Generating them from a single image is usually framed as a vision-language model that emits a serialized asset f…
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
Kaixin zhang, Xiaohe Li, Jiahao Li +4
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal cau…
FusionTrack: End-to-End Multi-Object Tracking in Arbitrary Multi-View Environment
Xiaohe Li, Pengfei Li, Zide Fan +4
Multi-view multi-object tracking (MVMOT) has found widespread applications in intelligent transportation, surveillance systems, and urban management. However, existing studies rare…
SAFE: Self-Adjustment Federated Learning Framework for Remote Sensing Collaborative Perception
Xiaohe Li, Haohua Wu, Jiahao Li +5
The rapid increase in remote sensing satellites has led to the emergence of distributed space-based observation systems. However, existing distributed remote sensing models often r…