8 papers
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
Kyubyung Chae, Jewon Yeom, Jeongjae Park +5
Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory domains, relevant evidence is di…
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
Joosung Lee, Hwiyeol Jo, Donghyeon Ko +3
While large language models (LLMs) demonstrate strong capabilities across diverse user queries, they still suffer from hallucinations, often arising from knowledge misalignment bet…
Robust Domain Generalization under Divergent Marginal and Conditional Distributions
Jewon Yeom, Kyubyung Chae, Hyunggyu Lim +3
Domain generalization (DG) aims to learn predictive models that can generalize to unseen domains. Most existing DG approaches focus on learning domain-invariant representations und…
What If TSF: A Benchmark for Reframing Forecasting as Scenario-Guided Multimodal Forecasting
Jinkwan Jang, Hyunbin Jin, Hyungjin Park +2
Time series forecasting is critical to real-world decision making, yet most existing approaches remain unimodal and rely on extrapolating historical patterns. While recent progress…
KoSimpleQA: A Korean Factuality Benchmark with an Analysis of Reasoning LLMs
Donghyeon Ko, Yeguk Jin, Kyubyung Chae +6
We present , a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cultural knowledge. KoSimpleQA is d…
Assessing Socio-Cultural Alignment and Technical Safety of Sovereign LLMs
Kyubyung Chae, Gihoon Kim, Gyuseong Lee +3
Recent trends in LLMs development clearly show growing interest in the use and application of sovereign LLMs. The global debate over sovereign LLMs highlights the need for governme…