4 papers
Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG
Francielle Vargas, João Robiatti, Diego Alves +6
Ensuring factuality and interpretability in RAG remains an open and urgent problem. We introduce Contrastive Evidence Rationale Attention (CERA), the first retrieval framework to e…
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
Che Liu, Yingji Zhang, Dong Zhang +13
This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…
LangVAE and LangSpace: Building and Probing for Language Model VAEs
Danilo S. Carvalho, Yingji Zhang, Harriet Unsworth +1
We present LangVAE, a novel framework for modular construction of variational autoencoders (VAEs) on top of pre-trained large language models (LLMs). Such language model VAEs can e…
Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning
Bohao Yang, Yingji Zhang, Dong Liu +2
Recent large language models (LLMs) have advanced table understanding capabilities but rely on converting tables into text sequences. While multimodal large language models (MLLMs)…