6 papers
CAMEO: A Conditional and Quality-Aware Multi-Agent Image Editing Orchestrator
Yuhan Pu, Hao Zheng, Ziqian Mo +5
Conditional image editing aims to modify a source image according to textual prompts and optional reference guidance. Such editing is crucial in scenarios requiring strict structur…
Evian: Towards Explainable Visual Instruction-tuning Data Auditing
Zimu Jia, Mingjie Xu, Andrew Estornell +1
The efficacy of Large Vision-Language Models (LVLMs) is critically dependent on the quality of their training data, requiring a precise balance between visual fidelity and instruct…
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
Mingjie Xu, Jinpeng Chen, Yuzhi Zhao +12
Multimodal large language models (MLLMs) have enabled a wide range of advanced vision-language applications, including fine-grained object recognition and contextual understanding.…
Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
Mingjie Xu, Andrew Estornell, Hongzheng Yang +4
The application of visual instruction tuning and other post-training techniques has significantly enhanced the capabilities of Large Language Models (LLMs) in visual understanding,…
When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
Zirui Pang, Haosheng Tan, Yuhan Pu +4
Image classification benchmark datasets such as CIFAR, MNIST, and ImageNet serve as critical tools for model evaluation. However, despite the cleaning efforts, these datasets still…
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
Zhijie Deng, Chris Yuhao Liu, Zirui Pang +5
Large Language Models (LLMs) have demonstrated strong capabilities in memorizing vast amounts of knowledge across diverse domains. However, the ability to selectively forget specif…