collaborators

7 papers

cs.RO2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

Jingkai Wang, Zihan Tang, Gu Zhang +7

Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this c…

cs.CV2026

UGround: Towards Unified Visual Grounding with Unrolled Transformers

Rui Qian, Xin Yin, Chuanhang Deng +4

We present UGround, a \textbf{U}nified visual \textbf{Ground}ing paradigm that dynamically selects intermediate layers across \textbf{U}nrolled transformers as ``mask as prompt,''…

cs.LG2026

PAMNet: Cycle-aware Phase-Amplitude Modulation Network for Multivariate Time Series Forecasting

Yingbo Zhou, Yutong Ye, Zhiwei Ling +5

Reliable periodic patterns serve as a fundamental basis for accurate multivariate time series forecasting. However, existing methods either implicitly extract periodicity through c…

cs.LG2026

PAMod: Modeling Cyclical Shifts via Phase-Amplitude Modulation for Non-stationary Time Series Forecasting

Yingbo Zhou, Yutong Ye, Shuhao Li +5

Real-world time series forecasting faces the fundamental challenge of non-stationary statistical properties, including shifts in mean and variance over time. While reversible insta…

cs.CV2026

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation

Rui Qian, Chuanhang Deng, Qiang Huang +6

Reasoning segmentation requires models to ground complex, implicit textual queries into precise pixel-level masks. Existing approaches rely on a single segmentation token $\texttt{…

cs.CV2026

Reasoning to Attend: Try to Understand How <SEG> Token Works

Rui Qian, Xin Yin, Dejing Dou

Current Large Multimodal Models (LMMs) empowered visual grounding typically rely on tokens as a text prompt to jointly optimize the vision-language model (e.g., LL…