activity
20242026
collaborators

9 papers

cs.AI2026

ReGraph: Learning to Generate Recipe Graphs from Food Images

Guoshan Liu, Bin Zhu, Pengkun Jiao +3

Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a structured transformation process in which in…

cs.CV2026

Disentangling Semantic Attention from Structural Bias in the Attention Manifold

Pengkun Jiao, Bin Zhu, Jingjing Chen +1

The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disprop…

cs.LG2026

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction

Ziyao Tang, Pengkun Jiao, Xinhang Chen +3

Given the quadratic complexity of attention, KV cache eviction is vital to accelerate model inference. Current KV cache eviction methods typically rely on instantaneous heuristic m…

cs.CV2026

Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models

Ziyao Tang, Pengkun Jiao, Bin Zhu +3

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely…

cs.LG2025

Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning

Xinlan Wu, Bin Zhu, Feng Han +2

Food analysis has become increasingly critical for health-related tasks such as personalized nutrition and chronic disease prevention. However, existing large multimodal models (LM…

cs.CV2025

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning

Pengkun Jiao, Bin Zhu, Jingjing Chen +2

Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task di…