3 citations · 6 across the 19 of their papers we have counts for
17 papers · 1 filter
The Hallucination Signal Is a Mean Shift: Why Simple Probes Suffice
Jungseob Lee, Jaehyung Seo, Heuiseok Lim
Hidden-state probes effectively detect LLM hallucinations, but the geometry of the signal remains poorly characterized, driving increasingly complex probe architectures. Across thr…
CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia
Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan +30
Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking pract…
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
Jimin Jung, MyoungJin Kim, Jaehyung Seo +1
The Plain Writing Act in the United States requires government documents to be accessible in clear and simple language that the general public can easily understand, yet existing s…
The Impact of Negated Text on Hallucination with Large Language Models
Jaehyung Seo, Hyeonseok Moon, Heuiseok Lim
Recent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing. However, the impact of negated text on hallucination…
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
Hyeonseok Moon, Seongtae Hong, Jaehyung Seo +1
Recent frontier-level LLMs have saturated many previously difficult benchmarks, leaving little room for further differentiation. This progress highlights the need for challenging b…
MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
Weihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty +32
Large language models (LLMs) are now used worldwide, yet their multimodal understanding and reasoning often degrade outside Western, high-resource settings. We propose MMA-ASIA, a…