5 papers
AerialMind: Towards Referring Multi-Object Tracking in UAV Scenarios
Chenglizhao Chen, Shaofeng Liang, Runwei Guan +6
Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intell…
ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
Hang Ding, Jiawei Zhou, Haiyun Jiang
Retrieval-Augmented Generation (RAG) has emerged as a powerful framework for knowledge-intensive tasks, yet its effectiveness in long-context scenarios is often bottlenecked by the…
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
Xinhan Zheng, Huyu Wu, Xueting Wang +2
Multimodal large language models (MLLMs) exhibit a pronounced preference for textual inputs when processing vision-language data, limiting their ability to reason effectively from…
When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models
Huyu Wu, Meng Tang, Xinhan Zheng +1
Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across a diverse range of multimodal tasks. However, these models suffer from a core problem know…
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
Bolun Sun, Yifan Zhou, Haiyun Jiang
This paper presents a novel application of large language models (LLMs) to enhance user comprehension of privacy policies through an interactive dialogue agent. We demonstrate that…