11 papers
SHALA-LLM: Smartly Handling Ambiguous Labels in Aligning LLMs
Jingyao Wu, Ashley Wang, Keane Ong +2
Many human-centered tasks, including natural language inference (NLI) and emotion recognition (ER), have multiple plausible interpretations, leading to label ambiguity and challeng…
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
Yuriel Ryan, Hei Man Ip, Adriel Kuek +2
Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting t…
SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning
Konstantinos Kontras, Teodora Gagaleska, Thomas Strypsteen +4
A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from an…
SCATR: Simple Calibrated Test-Time Ranking
Divya Shyamal, Marta KneževiÄ, Lan Tran +3
Test-time scaling (TTS) improves large language models (LLMs) by allocating additional compute at inference time. In practice, TTS is often achieved through parallel scaling: gener…
Interleaved Head Attention
Sai Surya Duvvuri, Chanakya Ekbote, Rachit Bansal +6
Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, MHA suffers from a fundamental linear scaling limitation: $H…
Group-Adaptive Threshold Optimization for Robust AI-Generated Text Detection
Minseok Jung, Cynthia Fuertes Panizo, Liam Dugan +4
The advancement of large language models (LLMs) has made it difficult to differentiate human-written text from AI-generated text. Several AI-text detectors have been developed in r…