5 papers
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
Xuehai Bai, Yang Shi, Yi-Fan Zhang +7
Recent image editing models have achieved remarkable progress in instruction following, multimodal understanding, and complex visual editing. However, existing benchmarks often fai…
HybridKV: Hybrid KV Cache Compression for Efficient Multimodal Large Language Model Inference
Bowen Zeng, Feiyang Ren, Jun Zhang +4
Multimodal Large Language Models (MLLMs) have advanced unified reasoning over text, images, and videos, but their inference is hindered by the rapid growth of key-value (KV) caches…
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
Xuehai Bai, Xiaoling Gu, Akide Liu +3
Recent advances in instruction-based image editing have shown remarkable progress. However, existing methods remain limited to relatively simple editing operations, hindering real-…
Distribution-Based Masked Medical Vision-Language Model Using Structured Reports
Shreyank N Gowda, Ruichi Zhang, Xiao Gu +2
Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models of…
Is Temporal Prompting All We Need For Limited Labeled Action Recognition?
Shreyank N Gowda, Boyan Gao, Xiao Gu +1
Video understanding has shown remarkable improvements in recent years, largely dependent on the availability of large scaled labeled datasets. Recent advancements in visual-languag…