17 papers
ScholaWrite: A Dataset of End-to-End Scholarly Writing Process
Khanh Chi Le, Linghe Wang, Minhwa Lee +3
Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build…
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Skylar Zhai, Jingcheng Liang, Dongyeop Kang
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missin…
The Amazing Agent Race: Strong Tool Users, Weak Navigators
Zae Myung Kim, Dongseok Lee, Jaehyung Kim +2
Existing tool-use benchmarks for LLM agents are overwhelmingly linear: our analysis of six benchmarks shows 55 to 100% of instances are simple chains of 2 to 5 steps. We introduce…
Align to Structure: Aligning Large Language Models with Structural Information
Zae Myung Kim, Anand Ramachandran, Farideh Tavazoee +3
Generating long, coherent text remains a challenge for large language models (LLMs), as they lack hierarchical planning and structured organization in discourse generation. We intr…
Scaling Unverifiable Rewards: A Case Study on Visual Insights
Shuyu Gan, James Mooney, Pan Hao +4
Large Language Model (LLM) agents can increasingly automate complex reasoning through Test-Time Scaling (TTS), iterative refinement guided by reward signals. However, many real-wor…
A2P-Vis: an Analyzer-to-Presenter Agentic Pipeline for Visual Insights Generation and Reporting
Shuyu Gan, Renxiang Wang, James Mooney +1
Automating end-to-end data science pipeline with AI agents still stalls on two gaps: generating insightful, diverse visual evidence and assembling it into a coherent, professional…