2 papers
cs.CV2025
Moment Sampling in Video LLMs for Long-Form Video QA
Mustafa Chasmai, Gauri Jagatap, Gouthaman KV +3
Recent advancements in video large language models (Video LLMs) have significantly advanced the field of video question answering (VideoQA). While existing methods perform well on…
cs.SD2025
The iNaturalist Sounds Dataset
Mustafa Chasmai, Alexander Shepard, Subhransu Maji +1
We present the iNaturalist Sounds Dataset (iNatSounds), a collection of 230,000 audio files capturing sounds from over 5,500 species, contributed by more than 27,000 recordists wor…