works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.RO2026

APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model

Yuanjie Lu, Beichen Wang, Zhengqi Wu +4

The paper introduces APPLV, a system that uses vision‑language models to predict parameters for classical motion planners, combining safety of traditional planners with adaptabilit…

cs.CV2026

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning

Yuan Qing, Chengzhi Mao, Boqing Gong

Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, we find a discrepancy: VIT ofte…

cs.CV2026

SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation

Shilong Xiang, Zirui Zhang, Lijun Yu +1

Autoregressive models excel in visual generation by treating images as 1D sequences of discrete tokens, mirroring language modeling. However, this flattening discards the intrinsic…

cs.CV2026

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

Chengzhi Mao, Xudong Lin, Wen-Sheng Chu

Vision foundation models are typically trained as static feature extractors, placing the burden of task adaptation onto large downstream models. We propose an alternative paradigm:…

cs.CV2026

SCOPE: Self-Supervised Concept Discovery via Preference Learning

Shilong Xiang, Zirui Zhang, Chengzhi Mao

Current representation learning paradigms force a fundamental compromise: self-supervised methods scale to massive datasets but yield opaque features, whereas interpretable models…

cs.AI2026

LACE: Lattice Attention for Cross-thread Exploration

Yang Li, Zirui Zhang, Yang Liu +1

Current large language models reason in isolation. Although it is common to sample multiple reasoning paths in parallel, these trajectories do not interact, and often fail in the s…