17 papers
MG-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
Sijun Dai, Qiang Huang, Xiaoxing You +1
Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle with complex cross-modal reasoning. Flat ve…
One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents
Yingchaojie Feng, Qiang Huang, Xiaoya Xie +4
Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain la…
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
Yiqun Sun, Qiang Huang, Anthony K. H. Tung +1
This position paper argues that text embedding research should move beyond surface meaning and embrace implicit semantics as a central modeling objective. Text embeddings are a fou…
Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
Haodong Chen, Qiang Huang, Jiaqi Zhao +3
Vision-Language Models (VLMs) are increasingly deployed in socially consequential settings, raising concerns about social bias driven by demographic cues. A central challenge in me…
Weight-Informed Self-Explaining Clustering for Mixed-Type Tabular Data
Lehao Li, Qiang Huang, Yihao Ang +3
Clustering mixed-type tabular data is fundamental for exploratory analysis, yet remains challenging due to misaligned numerical-categorical representations, uneven and context-depe…
Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events
Xiaoxing You, Qiang Huang, Lingyu Li +2
Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing app…