From the 1 of 10 linked papers with an AI index.
10 papers
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
Yuanjie Lu, Beichen Wang, Zhengqi Wu +4
The paper introduces APPLV, a system that uses vision‑language models to predict parameters for classical motion planners, combining safety of traditional planners with adaptabilit…
StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning
Yuan Qing, Chengzhi Mao, Boqing Gong
Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, we find a discrepancy: VIT ofte…
SSD: Spatially Speculative Decoding Accelerates Autoregressive Image Generation
Shilong Xiang, Zirui Zhang, Lijun Yu +1
Autoregressive models excel in visual generation by treating images as 1D sequences of discrete tokens, mirroring language modeling. However, this flattening discards the intrinsic…
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception
Chengzhi Mao, Xudong Lin, Wen-Sheng Chu
Vision foundation models are typically trained as static feature extractors, placing the burden of task adaptation onto large downstream models. We propose an alternative paradigm:…
SCOPE: Self-Supervised Concept Discovery via Preference Learning
Shilong Xiang, Zirui Zhang, Chengzhi Mao
Current representation learning paradigms force a fundamental compromise: self-supervised methods scale to massive datasets but yield opaque features, whereas interpretable models…
LACE: Lattice Attention for Cross-thread Exploration
Yang Li, Zirui Zhang, Yang Liu +1
Current large language models reason in isolation. Although it is common to sample multiple reasoning paths in parallel, these trajectories do not interact, and often fail in the s…