1 paper
Enjun Du, Siyi Liu, Ziyu Zheng +4
Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an…