9 papers
ReGraph: Learning to Generate Recipe Graphs from Food Images
Guoshan Liu, Bin Zhu, Pengkun Jiao +3
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which in…
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
Pengkun Jiao, Bin Zhu, Jingjing Chen +1
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…
Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction
Ziyao Tang, Pengkun Jiao, Xinhang Chen +3
Given the quadratic complexity of attention, KV cache eviction is vital to accelerate model inference. Current KV cache eviction methods typically rely on instantaneous heuristic m…
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
Ziyao Tang, Pengkun Jiao, Bin Zhu +3
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely…
Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
Xinlan Wu, Bin Zhu, Feng Han +2
Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LM…
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
Pengkun Jiao, Bin Zhu, Jingjing Chen +2
Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task di…