7 papers
Evidence-Consistent Generative Detection under Scenario-Level Distribution Shift
San Kim, JinYeong Bak
Conventional in-distribution evaluation can overestimate robustness when training and test data share recurring task-specific patterns or surface cues. This risk is especially rele…
MindTailor: Personalized Emotional Support via Post History-Grounded Case Formulation and Collaborative Refinement
Suhyun Han, Kyunghyun Cho, JinYeong Bak
As mental health concerns continue to rise globally, social media has emerged as a vital space where individuals seek emotional support. While prior work on personalized emotional…
Temporal Preference Optimization for Unsupervised Retrieval
HyunJin Kim, Jaejun Shim, Young Jin Kim +1
Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled documents via contrastive learning, but they struggle to capture the temporal relevan…
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
Jaehyeok Lee, Xiaoyuan Yi, Jing Yao +4
As LLMs are globally deployed, aligning their cultural value orientations is critical for safety and user engagement. However, existing benchmarks face the Construct-Composition-Co…
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
Tarek Naous, Anagha Savit, Carlos Rafael Catalan +17
As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensitivity to culturally diverse entities becomes increasingly important. Prior work by Naous et…
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
Jaehyeok Lee, Keisuke Sakaguchi, JinYeong Bak
Self-training approach for large language models (LLMs) improves reasoning abilities by training the models on their self-generated rationales. Previous approaches have labeled rat…