activity
20232026
most citedMVMTnet: A Multi-variate Multi-modal Transformer for Multi-class Classification of Cardiac Irregularities Using ECG Waveforms and Clinical Notes

2 citations · 4 across the 13 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Priyank Agrawal, Ankur Samanta, Shervin Ghasemlou +5

Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on difficult problems: when a model cannot gene…

cs.LG2026

Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias

Ofek I. Cohen, Lior Shani, Aviv Rosenberg +3

Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. Howeve…

cs.LG2026

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale

Amit Roth, Ankur Samanta, Matan Halevy +2

Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful und…

cs.LG2025

Iterative GRPO: Batch-Online Multi-Turn RL via Single-Turn RLHF

Daniel R. Jiang, Ankur Samanta, Yukai Yang +3

Practical LLM agents often operate over multi-turn conversations where success is determined only after the full interaction ends. Most multi-turn RL methods train via on-policy ro…

cs.LG2025

Imbalanced Gradients in RL Post-Training of Multi-Task LLMs

Runzhe Wu, Ankur Samanta, Ayush Jain +7

Multi-task post-training of large language models (LLMs) is typically performed by mixing datasets from different tasks and optimizing them jointly. This approach implicitly assume…

cs.LG2025

FragmentNet: Adaptive Graph Fragmentation for Graph-to-Sequence Molecular Representation Learning

Ankur Samanta, Rohan Gupta, Aditi Misra +2

Molecular representation learning methods typically tokenize molecules as individual atoms or use rigid, rule-based fragment decompositions, limiting their ability to capture meani…