Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data
Srinivasan Manoharan, Dilipkumar Nallusamy, Sachin Kumar +1
Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks incurs prohibitive latency, cost, and data-privacy overhead. We present a hybrid fra…
cs.LG2026
Leveraging Pretrained Language Models as Energy Functions for Glauber Dynamics Text Diffusion
Tarun Kathuria, Sachin Kumar
We present a discrete diffusion-based language model using Glauber dynamics from statistical physics. Our main insight is that instead of trying to train a discrete state space dif…
cs.LG2024
RewardBench: Evaluating Reward Models for Language Modeling
Nathan Lambert, Valentina Pyatkin, Jacob Morrison +9
Reward models (RMs) are at the crux of successfully using RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluatio…