1 paper
Xianyun Sun, Chaoyou Fu, Zhengye Zhang +6
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific…