papers

Publications (25)

cs.LG2026

From Observations to Events: Event-Aware World Model for Reinforcement Learning

Zhao-Han Peng, Shaohui Li, Zhi Li +3

While model-based reinforcement learning (MBRL) improves sample efficiency by learning world models from raw observations, existing methods struggle to generalize across structural…

cs.CV2026

Towards Cross-Platform Generalization: Domain Adaptive 3D Detection with Augmentation and Pseudo-Labeling

Xiyan Feng, Wenbo Zhang, Lu Zhang +3

This technical report represents the award-winning solution to the Cross-platform 3D Object Detection task in the RoboSense2025 Challenge. Our approach is built upon PVRCNN++, an e…

cs.CV2024

GLRT-Based Metric Learning for Remote Sensing Object Retrieval

Linping Zhang, Yu Liu, Xueqian Wang +2

With the improvement in the quantity and quality of remote sensing images, content-based remote sensing object retrieval (CBRSOR) has become an increasingly important topic. Howeve…

cs.CV2026

SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis

Chen-Chen Fan, Peiyao Guo, Linping Zhang +7

Given the limitations of satellite orbits and imaging conditions, multi-modal remote sensing (RS) data is crucial in enabling long-term earth observation. However, maritime surveil…

eess.SP2025

A Marginal Distributionally Robust Kalman Filter for Centralized Fusion

Weizhi Chen, Yaowen Li, Yu Liu +1

State estimation is a fundamental problem for multi-sensor information fusion, essential in applications such as target tracking, power systems, and control automation. Previous re…

cs.CV2019

ROI Pooled Correlation Filters for Visual Tracking

Yuxuan Sun, Chong Sun, Dong Wang +2

The ROI (region-of-interest) based pooling method performs pooling operations on the cropped ROI regions for various samples and has shown great success in the object detection met…

cs.AI2024

LLMs Can Evolve Continually on Modality for X-Modal Reasoning

Jiazuo Yu, Haomiao Xiong, Lu Zhang +7

Multimodal Large Language Models (MLLMs) have gained significant attention due to their impressive capabilities in multimodal understanding. However, existing methods rely heavily…

cs.RO2026

The RoboSense Challenge: Sense Anything, Navigate Anywhere, Adapt Across Platforms

Lingdong Kong, Shaoyuan Xie, Zeying Gong +135

Autonomous systems are increasingly deployed in open and dynamic environments -- from city streets to aerial and indoor spaces -- where perception models must remain reliable under…

eess.SP2019

Distributed Detection of Sparse Stochastic Signals via Fusion of 1-bit Local Likelihood Ratios

Chengxi Li, You He, Xueqian Wang +2

In this letter, we consider the detection of sparse stochastic signals with sensor networks (SNs), where the fusion center (FC) collects 1-bit data from the local sensors and then…

cs.CV2024

MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

Xiaomin Li, Xu Jia, Qinghe Wang +5

Existing pretrained text-to-video (T2V) models have demonstrated impressive abilities in generating realistic videos with basic motion or camera movement. However, these models exh…

cs.CL2026

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects

Sicheng Yang, Shulan Ruan, Shiwei Wu +4

While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription. Existing b…

cs.HC2026

Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI

Sicheng Yang, Yukai Huang, Weitong Cai +8

What if accessing the web did not require a screen, a stable desk, or even free hands? For people navigating crowded cities, living with low vision, or experiencing cognitive overl…

cs.CV2023

Magicremover: Tuning-free Text-guided Image inpainting with Diffusion Models

Siyuan Yang, Lu Zhang, Liqian Ma +3

Image inpainting aims to fill in the missing pixels with visually coherent and semantically plausible content. Despite the great progress brought from deep generative models, this…

cs.CV2024

Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters

Jiazuo Yu, Yunzhi Zhuge, Lu Zhang +4

Continual learning can empower vision-language models to continuously acquire new knowledge, without the need for access to the entire historical dataset. However, mitigating the p…

cs.CV2025

ReNeg: Learning Negative Embedding with Reward Guidance

Xiaomin Li, Yixuan Liu, Takashi Isobe +8

In text-to-image (T2I) generation applications, negative embeddings have proven to be a simple yet effective approach for enhancing generation quality. Typically, these negative em…

cs.CV2020

Change Detection in Heterogeneous Optical and SAR Remote Sensing Images via Deep Homogeneous Feature Fusion

Xiao Jiang, Gang Li, Yu Liu +2

Change detection in heterogeneous remote sensing images is crucial for disaster damage assessment. Recent methods use homogenous transformation, which transforms the heterogeneous…

cs.MM2025

A Survey of Multi-sensor Fusion Perception for Embodied AI: Background, Methods, Challenges and Prospects

Shulan Ruan, Rongwei Wang, Xuchen Shen +8

Multi-sensor fusion perception (MSFP) is a key technology for embodied AI, which can serve a variety of downstream tasks (e.g., 3D object detection and semantic segmentation) and a…

cs.HC2025

Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation

Sicheng Yang, Yukai Huang, Weitong Cai +6

The performance of egocentric AI agents is fundamentally limited by multimodal intent ambiguity. This challenge arises from a combination of underspecified language, imperfect visu…

cs.CL2024

GaLore: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection

Xutao Liao, Shaohui Li, Yuhui Xu +3

Recent low-rank training methods, such as GaLore, have significantly reduced the memory required to optimize large language models (LLMs). However, these methods often suffer from…

cs.CV2025

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

Lu Zhang, Jiazuo Yu, Haomiao Xiong +4

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs f…

cs.CV2025

EffOWT: Transfer Visual Language Models to Open-World Tracking Efficiently and Effectively

Bingyang Wang, Kaer Huang, Bin Li +4

Open-World Tracking (OWT) aims to track every object of any category, which requires the model to have strong generalization capabilities. Trackers can improve their generalization…

cs.CV2026

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

Xiaomin Li, Tala Wang, Zichen Zhong +7

Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning.…

eess.SP2026

RadarPLM: Adapting Pre-trained Language Models for Marine Radar Target Detection by Selective Fine-tuning

Qiying Hu, Yaowen Li, Shengyi Zhang +3

Recent advances in pre-trained language models (PLMs) have demonstrated their capabilities in capturing universal knowledge, making them promising for radar signal processing appli…

cs.CV2023

Video Diffusion Models with Local-Global Context Guidance

Siyuan Yang, Lu Zhang, Yu Liu +2

Diffusion models have emerged as a powerful paradigm in video synthesis tasks including prediction, generation, and interpolation. Due to the limitation of the computational budget…

cs.CV2026

V2X-DSC: Multi-Agent Collaborative Perception with Distributed Source Coding Guided Communication

Yuankun Zeng, Shaohui Li, Zhi Li +3

Collaborative perception improves 3D understanding by fusing multi-agent observations, yet intermediate-feature sharing faces strict bandwidth constraints as dense BEV features sat…