2 papers
cs.CL2026
Real-Time Voice AI Hears but Does Not Listen
Martijn Bartelds, Federico Bianchi, James Zou
Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live…
cs.RO2026
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
Gia-Binh Nguyen, Trong-Bao Ho, Thien-Loc Ha +18
Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion parameter architectures impose pro…