most citedDEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

1 citations · 1 across the 1 of their papers we have counts for

collaborators

11 papers

cs.CL20261 cited

DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation

Janghoon Han, Heegyu Kim, Changho Lee +6

Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, e…

cs.AI2026

A Regret Minimization Framework on Preference Learning in Large Language Models

Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim +4

Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness sig…

cs.CL2026

Early Decisions Matter: Proximity Bias and Initial Trajectory Shaping in Non-Autoregressive Diffusion Language Models

Jiyeon Kim, Sungik Choi, Yongrae Jo +2

Diffusion-based language models (dLLMs) have emerged as a promising alternative to autoregressive language models, offering the potential for parallel token generation and bidirect…

cs.SE2026

DuET: Dual Execution for Test Output Prediction with Generated Code and Pseudocode

Hojae Han, Jaejin Kim, Seung-won Hwang +2

This work addresses test output prediction, a key challenge in test case generation. To improve the reliability of predicted outputs by LLMs, prior approaches generate code first t…

cs.CL2026

SPRIG: Improving Large Language Model Performance by System Prompt Optimization

Lechen Zhang, Tolga Ergen, Lajanugen Logeswaran +2

Large Language Models (LLMs) have shown impressive capabilities in many scenarios, but their performance depends, in part, on the choice of prompt. Past research has focused on opt…

cs.LG2026

SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety

Geon-Hyeong Kim, Yu Jin Kim, Byoungjip Kim +4

As Large Language Models (LLMs) are increasingly deployed in real-world applications, balancing helpfulness and safety has become a central challenge. A natural approach is to inco…