12 citations · 16 across the 14 of their papers we have counts for
14 papers
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
Hyungjoo Chae, Taeyoon Kwon, Seungjun Moon +7
This paper presents Coffee-Gym, a comprehensive RL environment for training models that provide feedback on code editing. Coffee-Gym includes two major components: (1) Coffee, a da…
Evaluating Robustness of Reward Models for Mathematical Reasoning
Sunghwan Kim, Dongjin Kang, Taeyoon Kwon +4
Reward models are key in reinforcement learning from human feedback (RLHF) systems, aligning the model behavior with human preferences. Particularly in the math domain, there have…
YA-TA: Towards Personalized Question-Answering Teaching Assistants using Instructor-Student Dual Retrieval-augmented Knowledge Fusion
Dongil Yang, Suyeon Lee, Minjin Kim +4
Engagement between instructors and students plays a crucial role in enhancing students'academic performance. However, instructors often struggle to provide timely and personalized…
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
Heejae Chon, Seonghyeon Lee, Jinyoung Yeo +1
Language models (LMs) have exhibited impressive abilities in generating codes from natural language requirements. In this work, we highlight the diversity of code generated by LMs…
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
Kai Tzu-iunn Ong, Taeyoon Kwon, Jinyoung Yeo
Guiding large language models with a selected set of human-authored demonstrations is a common practice for improving LLM applications. However, human effort can be costly, especia…
Ever-Evolving Memory by Blending and Refining the Past
Seo Hyun Kim, Keummin Ka, Yohan Jo +3
For a human-like chatbot, constructing a long-term memory is crucial. However, current large language models often lack this capability, leading to instances of missing important u…