activity
20242026
collaborators

8 papers

cs.CL2026

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim +1

Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a…

cs.AI2026

MERIT Feedback Elicits Better Bargaining in LLM Negotiators

Jihwan Oh, Murad Aghazada, Yooju Shin +2

Bargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic…

cs.LG2025

AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners

Woosung Koh, Wonbeen Oh, Jaein Jang +7

Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (…

cs.LG2025

LLM Agents for Bargaining with Utility-based Feedback

Jihwan Oh

Bargaining, a critical aspect of real-world interactions, presents challenges for large language models (LLMs) due to limitations in strategic depth and adaptation to complex human…

cs.CL2025

Guiding Reasoning in Small Language Models with LLM Assistance

Yujin Kim, Euiin Yi, Minu Kim +2

The limited reasoning capabilities of small language models (SLMs) cast doubt on their suitability for tasks demanding deep, multi-step logical deduction. This paper introduces a f…

cs.LG2025

: Scalable Auto-Feedback for LLM-based Chart Generation

Woosung Koh, Jang Han Yoon, MinHyung Lee +7

Generating high-quality charts with Large Language Models (LLMs) presents significant challenges due to limited data and the high cost of scaling through human curation. $\langle \…