4 papers
CacheRL:Multi-Turn Tool-Calling Agents via Cached Rollouts and Hybrid Reward
Md Amirul Islam, Sumiran Thakur, Huancheng Chen +3
We present CacheRL, a system for training small agent foundation models that achieves 92 percent process accuracy on multi-step tool-calling tasks, approaching GPT-5's 94 percent w…
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
Myeongseob Ko, Jihyun Jeong, Sumiran Singh Thakur +2
Anonymization is often assumed to protect privacy once explicit identifiers are removed, because re-identification has historically required specialized expertise, tailored algorit…
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
Gyuhak Kim, Sumiran Singh Thakur, Su Min Park +2
Supervised fine-tuning (SFT) has become an essential step in tailoring large language models (LLMs) to align with human expectations and specific downstream tasks. However, existin…
Harnessing Business and Media Insights with Large Language Models
Yujia Bao, Ankit Parag Shah, Neeru Narang +30
This paper introduces Fortune Analytics Language Model (FALM). FALM empowers users with direct access to comprehensive business analysis, including market trends, company performan…