Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning
Mukul Singh, Ananya Singha, Arjun Radhakrishna +1
We analyze reasoning in language models during task-specific fine-tuning and draws parallel between reasoning tokens--intermediate steps generated while solving problem and the hum…
cs.CL2025
An Empirical Study of Validating Synthetic Data for Formula Generation
Usneek Singh, José Cambronero, Sumit Gulwani +5
Large language models (LLMs) can be leveraged to help with writing formulas in spreadsheets, but resources on these formulas are scarce, impacting both the base performance of pre-…
cs.CL2025
Ordered Semantically Diverse Sampling for Textual Data
Ashish Tiwari, Mukul Singh, Ananya Singha +1
The goal of diversity sampling is to select a representative subset of data in a way that maximizes information contained in the subset while keeping its cardinality small. We intr…