5 papers
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
Tuochao Chen, Bandhav Veluri, Hongyu Gong +1
Dialogue models falter in noisy, multi-speaker environments, often producing irrelevant responses and awkward turn-taking. We present AV-Dialog, the first multimodal dialog framewo…
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
Bandhav Veluri, Benjamin N Peloquin, Bokai Yu +2
Despite broad interest in modeling spoken dialogue agents, most approaches are inherently "half-duplex" -- restricted to turn-based interaction with responses requiring explicit pr…
IRIS: Wireless Ring for Vision-based Smart Home Interaction
Maruchi Kim, Antonio Glenn, Bandhav Veluri +5
Integrating cameras into wireless smart rings has been challenging due to size and power constraints. We introduce IRIS, the first wireless vision-enabled smart ring system for sma…
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
Hongyu Gong, Bandhav Veluri
Expressive speech-to-speech translation (S2ST) is a key research topic in seamless communication, which focuses on the preservation of semantics and speaker vocal style in translat…
Look Once to Hear: Target Speech Hearing with Noisy Examples
Bandhav Veluri, Malek Itani, Tuochao Chen +2
In crowded settings, the human brain can focus on speech from a target speaker, given prior knowledge of how they sound. We introduce a novel intelligent hearable system that achie…