5 papers
Korean Canonical Legal Benchmark: Toward Knowledge-Independent Evaluation of LLMs' Legal Reasoning Capabilities
Hongseok Oh, Wonseok Hwang, Kyoung-Woon On
We introduce the Korean Canonical Legal Benchmark (KCL), a benchmark designed to assess language models' legal reasoning capabilities independently of domain-specific knowledge. KC…
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
Hongseok Oh, Wonseok Hwang
Recently, Large Vision-Language Models (LVLMs) show remarkable performance across various domains. However, these models suffer from object hallucination. In this work, we study ob…
LRAGE: Legal Retrieval Augmented Generation Evaluation Tool
Minhu Park, Hongseok Oh, Eunkyung Choi +1
Recently, building retrieval-augmented generation (RAG) systems to enhance the capability of large language models (LLMs) has become a common practice. Especially in the legal doma…
Does Alignment Tuning Really Break LLMs' Internal Confidence?
Hongseok Oh, Wonseok Hwang
Large Language Models (LLMs) have shown remarkable progress, but their real-world application necessitates reliable calibration. This study conducts a comprehensive analysis of cal…
On the Consideration of AI Openness: Can Good Intent Be Abused?
Yeeun Kim, Hyunseo Shin, Eunkyung Choi +3
Open source is a driving force behind scientific advancement.However, this openness is also a double-edged sword, with the inherent risk that innovative technologies can be misused…