4 papers
Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models
Yiming Zhong, Chang Nie, Caifeng Shan
Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are…
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
Chang Nie, Jiaju Wei, Junlan Feng +2
Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. However, advanced video agents o…
EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memory
Chang Nie, Chaoyou Fu, Junlan Feng +1
Existing embedding models are inherently static: they encode text segments in isolation, ignoring their surrounding context and temporal order. This paper introduces EvoEmbedding,…
PersonaVLM: Long-Term Personalized Multimodal LLMs
Chang Nie, Chaoyou Fu, Yifan Zhang +2
Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual preferences remains limited. P…