activity
20232026
most citedFine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

11 citations · 16 across the 22 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

6 papers · 2 filters

cs.CV2024

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Zhiqi Ge, Juncheng Li, Xinglei Pang +7

Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based age…

cs.CV2024

STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Haiyi Qiu, Minghe Gao, Long Qian +7

Video Large Language Models (Video-LLMs) have recently shown strong performance in basic video understanding tasks, such as captioning and coarse-grained question answering, but st…

cs.CV2024★ 1 cited

AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Qifan Yu, Wei Chow, Zhongqi Yue +7

Instruction-based image editing aims to modify specific image elements with natural language instructions. However, current models in this domain often struggle to accurately execu…

cs.CV2024

Unified Generative and Discriminative Training for Multi-modal Large Language Models

Wei Chow, Juncheng Li, Qifan Yu +7

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle…

cs.CV2024★ 2 cited

Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration

Kaihang Pan, Zhaoyu Fan, Juncheng Li +6

The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and ex…

cs.CV2024

Auto-Encoding Morph-Tokens for Multimodal LLM

Kaihang Pan, Siliang Tang, Juncheng Li +6

For multimodal LLMs, the synergy of visual comprehension (textual output) and generation (visual output) presents an ongoing challenge. This is due to a conflicting objective: for…