collaborators

10 papers

cs.RO2026

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Hongyu Qu, Jianzhe Gao, Xiaobin Hu +6

Mainstream Vision-Language-Action (VLA) models predict actions primarily from the current observation under a Markovian assumption, thus struggling with long-horizon, temporally de…

cs.CV2026

SinkTrack: Attention Sink based Context Anchoring for Large Language Models

Xu Liu, Guikun Chen, Wenguan Wang

Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these problems, where LLMs' focus sh…

cs.CV2026

Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

Bo Zhou, Qiuxia Lai, Zeren Sun +3

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such…

cs.CV2026

Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation

Xinhao Cai, Gensheng Pei, Zeren Sun +3

In this paper, we propose \textbf{Iris}, a deterministic framework for Monocular Depth Estimation (MDE) that integrates real-world priors into the diffusion model. Conventional fee…

cs.CV2026

PKINet-v2: Towards Powerful and Efficient Poly-Kernel Remote Sensing Object Detection

Xinhao Cai, Liulei Li, Gensheng Pei +3

Object detection in remote sensing images (RSIs) is challenged by the coexistence of geometric and spatial complexity: targets may appear with diverse aspect ratios, while spanning…

cs.CV2026

Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis

Xinhao Cai, Liulei Li, Gensheng Pei +4

This paper presents a generation-based debiasing framework for object detection. Prior debiasing methods are often limited by the representation diversity of samples, while naive g…