adaptive inference 1agentic visual reasoning 1diffusion models 1multimodal language models 1multimodal large language models 1multi-reference video editing 1reference tokens 1reinforcement learning 1structured instructions 1tool use adaptiveness 1
From the 2 of 42 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
Yue Ding, Yiyan Ji, Jungang Li +12
Omni-modal Large Language Models (Omni-LLMs) have demonstrated strong capabilities in audio-video understanding tasks. However, their reliance on long multimodal token sequences le…
cs.CL2026
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
Xinlong Chen, Weihong Lin, Jingyun Hua +10
Accurate dialogue description in audiovisual video captioning is crucial for downstream understanding and generation tasks. However, existing models generally struggle to produce f…