2 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Xudong Lu, Huankang Guan, Yang Bo +10
Multimodal Large Language Models excel at offline audio-visual understanding, but their ability to serve as mobile assistants in continuous real-world streams remains underexplored…