activity
20242026
collaborators

6 papers

cs.LG2026

ISO: An RLVR-Native Optimization Stack

Hanqing Zhu, Wenyan Cong, Zhizhou Sha +8

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback i…

cs.LG2026

Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs

Sagnik Mukherjee, Lifan Yuan, Pavan Jayasinha +2

Reinforcement learning (RL), particularly RL from verifiable reward (RLVR), has become a crucial phase of training large language models (LLMs) and a key focus of current scaling e…

cs.NE2026

Graph Neural Network Assisted Genetic Algorithm for Structural Dynamic Response and Parameter Optimization

Sagnik Mukherjee, Indrajit Barua

The optimization of structural parameters, such as mass(m), stiffness(k), and damping coefficient(c), is critical for designing efficient, resilient, and stable structures. Convent…

cs.LG2025

Reinforcement Learning Finetunes Small Subnetworks in Large Language Models

Sagnik Mukherjee, Lifan Yuan, Dilek Hakkani-Tur +1

Reinforcement learning (RL) yields substantial improvements in large language models (LLMs) downstream task performance and alignment with human values. Surprisingly, such large ga…

cs.CL2025

ReasoningFlow: Semantic Structure of Complex Reasoning Traces

Jinu Lee, Sagnik Mukherjee, Dilek Hakkani-Tur +1

Large reasoning models (LRMs) generate complex reasoning traces with planning, reflection, verification, and backtracking. In this work, we introduce ReasoningFlow, a unified schem…

cs.AI2024

Infogent: An Agent-Based Framework for Web Information Aggregation

Revanth Gangi Reddy, Sagnik Mukherjee, Jeonghwan Kim +3

Despite seemingly performant web agents on the task-completion benchmarks, most existing methods evaluate the agents based on a presupposition: the web navigation task consists of…