Publications (25)
From Observations to Events: Event-Aware World Model for Reinforcement Learning
Zhao-Han Peng, Shaohui Li, Zhi Li +3
While model-based reinforcement learning (MBRL) improves sample efficiency by learning world models from raw observations, existing methods struggle to generalize across structural…
Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling
Xiyan Feng, Wenbo Zhang, Lu Zhang +3
This technical report represents the award-winning solution to the Cross-platform 3D Object Detection task in the RoboSense2025 Challenge. Our approach is built upon PVRCNN++, an e…
GLRT-Based Metric Learning for Remote Sensing Object Retrieval
Linping Zhang, Yu Liu, Xueqian Wang +2
With the improvement in the quantity and quality of remote sensing images, content-based remote sensing object retrieval (CBRSOR) has become an increasingly important topic. Howeve…
SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis
Chen-Chen Fan, Peiyao Guo, Linping Zhang +7
Given the limitations of satellite orbits and imaging conditions, multi-modal remote sensing (RS) data is crucial in enabling long-term earth observation. However, maritime surveil…
A Marginal Distributionally Robust Kalman Filter for Centralized Fusion
Weizhi Chen, Yaowen Li, Yu Liu +1
State estimation is a fundamental problem for multi-sensor information fusion, essential in applications such as target tracking, power systems, and control automation. Previous re…
ROI Pooled Correlation Filters for Visual Tracking
Yuxuan Sun, Chong Sun, Dong Wang +2
The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection met…
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
Jiazuo Yu, Haomiao Xiong, Lu Zhang +7
Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily…
The RoboSense Challenge: Sense Anything, Navigate Anywhere, Adapt Across Platforms
Lingdong Kong, Shaoyuan Xie, Zeying Gong +135
Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under…
Distributed Detection of Sparse Stochastic Signals via Fusion of 1-bit Local Likelihood Ratios
Chengxi Li, You He, Xueqian Wang +2
In this letter, we consider the detection of sparse stochastic signals with sensor networks (SNs), where the fusion center (FC) collects 1-bit data from the local sensors and then…
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models
Xiaomin Li, Xu Jia, Qinghe Wang +5
Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exh…
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
Sicheng Yang, Shulan Ruan, Shiwei Wu +4
While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing b…
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
Sicheng Yang, Yukai Huang, Weitong Cai +8
What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overl…
Magicremover: Tuning-free Text-guided Image inpainting with Diffusion Models
Siyuan Yang, Lu Zhang, Liqian Ma +3
Image inpainting aims to fill in the missing pixels with visually coherent and semantically plausible content. Despite the great progress brought from deep generative models, this…
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
Jiazuo Yu, Yunzhi Zhuge, Lu Zhang +4
Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the p…
ReNeg: Learning Negative Embedding with Reward Guidance
Xiaomin Li, Yixuan Liu, Takashi Isobe +8
In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative em…
Change Detection in Heterogeneous Optical and SAR Remote Sensing Images via Deep Homogeneous Feature Fusion
Xiao Jiang, Gang Li, Yu Liu +2
Change detection in heterogeneous remote sensing images is crucial for disaster damage assessment. Recent methods use homogenous transformation, which transforms the heterogeneous…
A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects
Shulan Ruan, Rongwei Wang, Xuchen Shen +8
Multi-sensor fusion perception (MSFP) is a key technology for embodied AI, which can serve a variety of downstream tasks (e.g., 3D object detection and semantic segmentation) and a…
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
Sicheng Yang, Yukai Huang, Weitong Cai +6
The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visu…
GaLore: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection
Xutao Liao, Shaohui Li, Yuhui Xu +3
Recent low-rank training methods, such as GaLore, have significantly reduced the memory required to optimize large language models (LLMs). However, these methods often suffer from…
FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
Lu Zhang, Jiazuo Yu, Haomiao Xiong +4
Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs f…
EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively
Bingyang Wang, Kaer Huang, Bin Li +4
Open-World Tracking (OWT) aims to track every object of any category, which requires the model to have strong generalization capabilities. Trackers can improve their generalization…
Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
Xiaomin Li, Tala Wang, Zichen Zhong +7
Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning.…
RadarPLM: Adapting Pre-trained Language Models for Marine Radar Target Detection by Selective Fine-tuning
Qiying Hu, Yaowen Li, Shengyi Zhang +3
Recent advances in pre-trained language models (PLMs) have demonstrated their capabilities in capturing universal knowledge, making them promising for radar signal processing appli…
Video Diffusion Models with Local-Global Context Guidance
Siyuan Yang, Lu Zhang, Yu Liu +2
Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget…
V2X-DSC: Multi-Agent Collaborative Perception with Distributed Source Coding Guided Communication
Yuankun Zeng, Shaohui Li, Zhi Li +3
Collaborative perception improves 3D understanding by fusing multi-agent observations, yet intermediate-feature sharing faces strict bandwidth constraints as dense BEV features sat…