9 papers
MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models
Yuncheng Yang, Feiyang Ye, Shixian Luo +7
Vision-Language Models (VLMs) have achieved success using homogeneous Transformers to process multimedia data. Recent studies show that heterogeneous structures interleaving effici…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camer…
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
Guoyang Xia, Fengfa Li, Hongjin Ji +4
Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training paradigms remain difficult to compare because…
Beyond N-gram: Data-Aware X-GRAM Extraction for Efficient Embedding Parameter Scaling
Yilong Chen, Yanxi Xie, Zitian Gao +10
Large token-indexed lookup tables provide a compute-decoupled scaling path, but their practical gains are often limited by poor parameter efficiency and rapid memory growth. We att…
AFD-SLU: Adaptive Feature Distillation for Spoken Language Understanding
Yan Xie, Yibo Cui, Liang Xie +1
Spoken Language Understanding (SLU) is a core component of conversational systems, enabling machines to interpret user utterances. Despite its importance, developing effective SLU…
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
Luoyang Sun, Jiwen Jiang, Yifeng Ding +9
Vision-Language-Action Models (VLAs) have emerged as a key paradigm of Physical AI and are increasingly deployed in autonomous vehicles, robots, and smart spaces. In these resource…