2 papers
cs.CV2025
AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
Gengyuan Zhang, Tanveer Hannan, Hermine Kleiner +6
An ideal vision-language agent serves as a bridge between the human users and their surrounding physical world in real-world applications like autonomous driving and embodied agent…
cs.SD2025
Improving AI-generated music with user-guided training
Vishwa Mohan Singh, Sai Anirudh Aryasomayajula, Ahan Chatterjee +2
AI music generation has advanced rapidly, with models like diffusion and autoregressive algorithms enabling high-fidelity outputs. These tools can alter styles, mix instruments, or…