Showing 2025Show all
2 papers · 1 filter
cs.SD2025
Assessing Factual Music Comprehension in Large Audio Language Models
Daniel Chenyu Lin, Michael Freeman, John Thickstun
Large audio language models (LALMs) leverage multimodal representations to generate open-ended answers to natural language queries about audio. In this paper, we (1) provide empiri…
cs.CV2025
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
Nate Gillman, Charles Herrmann, Michael Freeman +4
Recent advances in video generation models have sparked interest in world models capable of simulating realistic environments. While navigation has been well-explored, physically m…