Showing cs.MMShow all
2 papers · 1 filter
cs.MM2026
Inference-Time Scaling for Joint Audio-Video Generation
Jaemin Jung, Kyeongha Rho, Inkyu Shin +1
Joint audio-video generation aims to synthesize realistic audio-video pairs that are both semantically aligned with text prompts and precisely synchronized. While existing joint au…
cs.MM2025
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
Kyeongha Rho, Hyeongkeun Lee, Valentio Iverson +1
Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality.…