12 papers
LA4VLA: Learning to Act without Seeing via Language-Action Pretraining
Tao Lin, Yuxin Du, Yiran Mao +13
Vision-Language-Action (VLA) models are commonly pretrained on robot demonstrations by jointly mapping visual observations and language instructions to actions. However, dense visu…
-Scene: Physically Grounded Image-to-3D Scene Reconstruction
Haodong Li, Lulu Shao, Haolin Lu +4
Reconstructing compositional 3D scenes from a single image is a fundamental challenge in 3D world modeling. Recent methods can recover high-fidelity, complete 3D objects and predic…
Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
Runze Wang, Yuqian Fu, Yu Li +7
Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…
StepAudio 2.5 Technical Report
Bin Lin, Bo Zhao, Boyong Wu +98
Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…
Accelerating Vision Foundation Models with Drop-in Depthwise Convolution
Carmelo Scribano, Mohammad Mahdi, Nedyalko Prisadnikov +5
Pretrained vision foundation models deliver strong performance across tasks with limited fine-tuning. However, their Vision Transformer (ViT) backbones impose high inference costs,…
VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation
Linan ZHU, Zihao Zhai, Xiao Han +4
Emotion Recognition in Conversation (ERC) is essential for effective human-machine interaction, aiming to identify speakers' emotional states in multi-turn dialogues. Early text-ba…