5 papers
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
Zhengbo Zhang, Jinbo Su, Zhaowen Zhou +14
The rapid advancement of Multimodal Large Language Models (MLLMs) has enabled browsing agents to acquire and reason over multimodal information in the real world. But existing benc…
EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models
Yifei Zhang, Mingyang Li, Henry Gao +1
Large language models (LLMs) demonstrate strong cognitive intelligence (IQ), yet many real-world interactions also require emotional intelligence (EQ) to produce responses that are…
Cross-modal RAG: Sub-dimensional Text-to-Image Retrieval-Augmented Generation
Mengdan Zhu, Senhao Cheng, Guangji Bai +2
Text-to-image generation increasingly demands access to domain-specific, fine-grained, and rapidly evolving knowledge that pretrained models cannot fully capture, necessitating the…
MEGL: Multimodal Explanation-Guided Learning
Yifei Zhang, Tianxu Jiang, Bo Pan +3
Explaining the decision-making processes of Artificial Intelligence (AI) models is crucial for addressing their "black box" nature, particularly in tasks like image classification.…
GraphNarrator: Generating Textual Explanations for Graph Neural Networks
Bo Pan, Zhen Xiong, Guanchen Wu +3
Graph representation learning has garnered significant attention due to its broad applications in various domains, such as recommendation systems and social network analysis. Despi…