3 papers
cs.LG2026
RiTTA: Modeling Event Relations in Text-to-Audio Generation
Yuhang He, Yash Jain, Xubo Liu +2
Despite significant advancements in Text-to-Audio (TTA) generation models achieving high-fidelity audio with fine-grained context understanding, they struggle to model the relation…
cs.LG2026
Bayesian Inverse Games with High-Dimensional Multi-Modal Observations
Yash Jain, Xinjie Liu, Lasse Peters +2
Many multi-agent interaction scenarios can be naturally modeled as noncooperative games, where each agent's decisions depend on others' future actions. However, deploying game-theo…
cs.LG2025
Test-time Prompt Refinement for Text-to-Image Models
Mohammad Abdul Hafeez Khan, Yash Jain, Siddhartha Bhattacharyya +1
Text-to-image (T2I) generation models have made significant strides but still struggle with prompt sensitivity: even minor changes in prompt wording can yield inconsistent or inacc…