works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cond-mat.mtrl-sci2026

Realizing record-high transverse thermoelectric figure of merit at room temperature in artificially tilted multilayers based on high power factor NiFe alloy

Yebin Lee, Fuyuki Ando, Takamasa Hirai +3

Transverse thermoelectric conversion using artificially tilted multilayers (ATMLs) offers a versatile device architecture that circumvents the structural limitations of conventiona…

cs.MM2026

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

Thanh V. T. Tran, Ngoc-Son Nguyen, Luong Tran +4

The paper introduces Flowley, an end‑to‑end model that generates synchronized audio directly from silent video using a novel progressive soft‑masked cross‑attention mechanism, and…

cs.SD2026

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data

Subhankar Ghosh, Jason Li, Paarth Neekhara +4

Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, speaker inconsistencies and sentence bounda…

eess.AS2026

Frame-Stacked Local Transformers For Efficient Multi-Codebook Speech Generation

Roy Fejgin, Paarth Neekhara, Xuesong Yang +6

Speech generation models based on large language models (LLMs) typically operate on discrete acoustic codes, which differ fundamentally from text tokens due to their multicodebook…

cs.AI2025

Align2Speak: Improving TTS for Low Resource Languages via ASR-Guided Online Preference Optimization

Shehzeen Hussain, Paarth Neekhara, Xuesong Yang +7

Developing high-quality text-to-speech (TTS) systems for low-resource languages is challenging due to the scarcity of paired text and speech data. In contrast, automatic speech rec…

eess.AS2025

HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset

Ryan Langman, Xuesong Yang, Paarth Neekhara +4

This paper introduces HiFiTTS-2, a large-scale speech dataset designed for high-bandwidth speech synthesis. The dataset is derived from LibriVox audiobooks, and contains approximat…