Showing 2026 · cs.AIShow all
2 papers · 2 filters
cs.AI2026
SteerDuplex: Steerable Duplex Speech Dialogue Models
Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai +13
Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability…
cs.AI2026
Do Audio-Visual Large Language Models Really See and Hear?
Ramaneswaran Selvakumar, Kaousheik Jayakumar, S Sakshi +3
Audio-Visual Large Language Models (AVLLMs) are emerging as unified interfaces to multimodal perception. We present the first mechanistic interpretability study of AVLLMs, analyzin…