10 papers
AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
Minbyul Jeong, Chanwoong Yoon
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmar…
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
Junha Jung, Minbyul Jeong, Suhyeon Lim +5
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain predominantly outcome-centric, relying on…
Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents
Minbyul Jeong
Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- while breadth, exhaustively enumerating a closed set and filling ea…
OpenBioRQ: Unsolved Biomedical Research Questions for Agents
Minbyul Jeong
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citat…
Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
Yein Park, Minbyul Jeong, Jaewoo Kang
The remarkable capabilities of modern large reasoning models are largely unlocked through post-training techniques such as supervised fine-tuning (SFT) and reinforcement learning (…
Trustworthy Agents for Electronic Health Records through Confidence Estimation
Yongwoo Song, Minbyul Jeong, Mujeen Sung
Large language models (LLMs) show promise for extracting information from Electronic Health Records (EHR) and supporting clinical decisions. However, deployment in clinical setting…