6 citations · 13 across the 7 of their papers we have counts for
7 papers
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
Yiqun Wang, Zhao Zhou, Xiangcheng Du +3
When dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In oth…
DCQA: Document-Level Chart Question Answering towards Complex Reasoning and Common-Sense Understanding
Anran Wu, Luwei Xiao, Xingjiao Wu +6
Visually-situated languages such as charts and plots are omnipresent in real-world documents. These graphical depictions are human-readable and are often analyzed in visually-rich…
Progressive Evidence Refinement for Open-domain Multimodal Retrieval Question Answering
Shuwen Yang, Anran Wu, Xingjiao Wu +4
Pre-trained multimodal models have achieved significant success in retrieval-based question answering. However, current multimodal retrieval question-answering models face two main…
CGMI: Configurable General Multi-Agent Interaction Framework
Shi Jinxin, Zhao Jiabao, Wang Yilei +3
Benefiting from the powerful capabilities of large language models (LLMs), agents based on LLMs have shown the potential to address domain-specific tasks and emulate human behavior…
FairMonitor: A Four-Stage Automatic Framework for Detecting Stereotypes and Biases in Large Language Models
Yanhong Bai, Jiabao Zhao, Jinxin Shi +3
Detecting stereotypes and biases in Large Language Models (LLMs) can enhance fairness and reduce adverse impacts on individuals or groups when these LLMs are applied. However, the…
DDT: Dual-branch Deformable Transformer for Image Denoising
Kangliang Liu, Xiangcheng Du, Sijie Liu +3
Transformer is beneficial for image denoising tasks since it can model long-range dependencies to overcome the limitations presented by inductive convolutional biases. However, dir…