8 papers
Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models
Taeyeong Kim, Ahhyun Kim, TaeHyeon Kim +1
Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable count from billions down to a…
MERIT Feedback Elicits Better Bargaining in LLM Negotiators
Jihwan Oh, Murad Aghazada, Yooju Shin +2
Bargaining is often regarded as a logical arena rather than an art or a matter of intuition, yet Large Language Models (LLMs) still struggle to navigate it due to limited strategic…
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Woosung Koh, Wonbeen Oh, Jaein Jang +7
Self-Taught Reasoners (STaR), synonymously known as Rejection sampling Fine-Tuning (RFT), is an integral part of the training pipeline of self-improving reasoning Language Models (…
LLM Agents for Bargaining with Utility-based Feedback
Jihwan Oh
Bargaining, a critical aspect of real-world interactions, presents challenges for large language models (LLMs) due to limitations in strategic depth and adaptation to complex human…
Guiding Reasoning in Small Language Models with LLM Assistance
Yujin Kim, Euiin Yi, Minu Kim +2
The limited reasoning capabilities of small language models (SLMs) cast doubt on their suitability for tasks demanding deep, multi-step logical deduction. This paper introduces a f…
: Scalable Auto-Feedback for LLM-based Chart Generation
Woosung Koh, Jang Han Yoon, MinHyung Lee +7
Generating high-quality charts with Large Language Models (LLMs) presents significant challenges due to limited data and the high cost of scaling through human curation. $\langle \…