6 papers
Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Models
Leijiang Gu, Zhen Zeng, Feng Li +2
Existing methods in Multimodal Knowledge Editing (MKE) have advanced the ability to correct outdated or inaccurate knowledge in Multimodal Large Language Models (MLLMs). However, t…
StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs
Chang Che, Ziqi Wang, Hui Ma +2
Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models to incrementally acquire new abilities. However, existing CVIT methods operate under a restricti…
LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
Chang Che, Ziqi Wang, Pengwan Yang +3
Continual Visual Instruction Tuning (CVIT) enables Multimodal Large Language Models (MLLMs) to incrementally learn new tasks over time. However, this process is challenged by catas…
CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs
Zhen Zeng, Leijiang Gu, Feng Li +2
Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned responses in cross-cultural settings…
WebUncertainty: Dual-Level Uncertainty Driven Planning and Reasoning For Autonomous Web Agent
Lingfeng Zhang, Yongan Sun, Jinpeng Hu +5
Recent advancements in large language models (LLMs) have empowered autonomous web agents to execute natural language instructions directly on real-world webpages. However, existing…
In-Context Examples Matter: Improving Emotion Recognition in Conversation with Instruction Tuning
Hui Ma, Bo Zhang, Jinpeng Hu +1
Emotion recognition in conversation (ERC) aims to identify the emotion of each utterance in a conversation, playing a vital role in empathetic artificial intelligence. With the gro…