1 paper
Daehwa Kim, Chris Harrison
We introduce and explore a new multimodal input representation for vision-language models: acoustic field video. Unlike conventional video (RGB with stereo/mono audio), our video s…