9 papers
Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
Zhiqi Wang, Yichi Zhang, Dongwon Lee +1
When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued instructions, Session Constraints (SCs),…
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?
Jiaxi Yang, Chaewan Chun, Jason Lucas +2
Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio tasks. As they are increasingly deployed in real-world applications, ensuring…
Context-Aware Multimodal Claim Verification in Spoken Dialogues
Chaewan Chun, Delvin Ce Zhang, Dongwon Lee
Every day, millions absorb claims from podcasts and streams that no fact-checker ever sees. Spoken misinformation is built through conversation, where credibility comes not from fa…
Moltbook Moderation: Uncovering Hidden Intent Through Multi-Turn Dialogue
Ali Al-Lawati, Nafis Tripto, Abolfazl Ansari +3
The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that a…
M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
Abolfazl Ansari, Delvin Ce Zhang, Zhuoyang Zou +2
Evaluating scientific arguments requires assessing the strict consistency between a claim and its underlying multimodal evidence. However, existing benchmarks lack the scale, domai…
DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning
Zhuoyang Zou, Abolfazl Ansari, Delvin Ce Zhang +2
Paper weakness identification using single-agent or multi-agent LLMs has attracted increasing attention, yet existing approaches exhibit key limitations. Many multi-agent systems s…