collaborators

7 papers

cs.LG2026

Entropy-Gated Latent Recursion

Soham Bhattacharjee, Dushyant Singh Chauhan, Salem Lahlou +2

Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-le…

cs.CL2026

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1

Data curation is a critical part of post-training pipelines for large language models, yet existing tools often treat ingestion, deduplication, synthetic generation, and quality fi…

cs.CL2026

Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation

Soham Bhattacharjee, Karun Sharma, Vinay Kumar Sankarapu +1

Synthetic post-training pipelines commonly filter generated samples with reward models or holistic LLM judges, yet two practices remain rarely examined together: whether the filter…

cs.CL2026

AlignTune: Modular Toolkit for Post-Training Alignment of Large Language Models

R E Zera Marveen Lyngkhoi, Chirag Chawla, Pratinav Seth +5

Post-training alignment is central to deploying large language models (LLMs), yet practical workflows remain split across backend-specific tools and ad-hoc glue code, making experi…

cs.CL2025

CAuSE: Decoding Multimodal Classifiers using Faithful Natural Language Explanation

Dibyanayan Bandyopadhyay, Soham Bhattacharjee, Mohammed Hasanuzzaman +1

Multimodal classifiers function as opaque black box models. While several techniques exist to interpret their predictions, very few of them are as intuitive and accessible as natur…

cs.CL2025

CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems

Soham Bhattacharjee, Mukund K Roy, Yathish Poojary +19

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognize…