5 papers
Realizing record-high transverse thermoelectric figure of merit at room temperature in artificially tilted multilayers based on high power factor NiFe alloy
Yebin Lee, Fuyuki Ando, Takamasa Hirai +3
Transverse thermoelectric conversion using artificially tilted multilayers (ATMLs) offers a versatile device architecture that circumvents the structural limitations of conventiona…
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
Zhen Wan, Chao-Han Huck Yang, Jinchuan Tian +15
We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is…
What do Speech Foundation Models Learn? Analysis and Applications
Ankita Pasad
Speech foundation models (SFMs) are designed to serve as general-purpose representations for a wide range of speech-processing tasks. The last five years have seen an influx of inc…
Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang +77
Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spo…
Training and Inference Efficiency of Encoder-Decoder Speech Models
Piotr Żelasko, Kunal Dhawan, Daniel Galvez +7
Attention encoder-decoder model architecture is the backbone of several recent top performing foundation speech models: Whisper, Seamless, OWSM, and Canary-1B. However, the reporte…