collaborators

12 papers

cs.RO2026

LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Tao Lin, Yuxin Du, Yiran Mao +13

Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…

cs.CV2026

-Scene: Physically Grounded Image-to-3D Scene Reconstruction

Haodong Li, Lulu Shao, Haolin Lu +4

Reconstructing compositional 3D scenes from a single image is a fundamental challenge in 3D world modeling. Recent methods can recover high-fidelity, complete 3D objects and predic…

cs.RO2026

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Runze Wang, Yuqian Fu, Yu Li +7

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

cs.CV2026

Accelerating Vision Foundation Models with Drop-in Depthwise Convolution

Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov +5

Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs,…

cs.AI2026

VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation

Linan ZHU, Zihao Zhai, Xiao Han +4

Emotion Recognition in Conversation (ERC) is essential for effective human-machine interaction, aiming to identify speakers' emotional states in multi-turn dialogues. Early text-ba…