2 papers
cs.CV2026
Omni-Prune: Query-Aware Unified Token Pruning for Efficient Omnimodal Large Language Models
Yiming Zhong, Chang Nie, Caifeng Shan
Omnimodal large language models (OmniLLMs) are rapidly extending multimodal reasoning to cover synchronized audio and video. However, the resulting audio-video token sequences are…
cs.CV2026
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
Runzhi Deng, Yundi Hu, Yiming Zhong +5
Large Multimodal Models (LMMs) show strong few-shot generalization, but industrial anomaly detection remains difficult because defects are small, input resolution is limited, and t…