activity
20232025
most citedFrom Words to Code: Harnessing Data for Program Synthesis from Natural Language

6 citations · 11 across the 16 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2025

Scaling Competence, Shrinking Reasoning: Cognitive Signatures in Language Model Learning

Mukul Singh, Ananya Singha, Arjun Radhakrishna +1

We analyze reasoning in language models during task-specific fine-tuning and draws parallel between reasoning tokens--intermediate steps generated while solving problem and the hum…

cs.CL2025

Ordered Semantically Diverse Sampling for Textual Data

Ashish Tiwari, Mukul Singh, Ananya Singha +1

The goal of diversity sampling is to select a representative subset of data in a way that maximizes information contained in the subset while keeping its cardinality small. We intr…

cs.CL2024

An Empirical Study of Validating Synthetic Data for Formula Generation

Usneek Singh, José Cambronero, Sumit Gulwani +5

Large language models (LLMs) can be leveraged to help with writing formulas in spreadsheets, but resources on these formulas are scarce, impacting both the base performance of pre-…

cs.CL20233 cited

Assessing GPT4-V on Structured Reasoning Tasks

Mukul Singh, José Cambronero, Sumit Gulwani +2

Multi-modality promises to unlock further uses for large language models. Recently, the state-of-the-art language model GPT-4 was enhanced with vision capabilities. We carry out a…

cs.CL2023

InstructExcel: A Benchmark for Natural Language Instruction in Excel

Justin Payan, Swaroop Mishra, Mukul Singh +7

With the evolution of Large Language Models (LLMs) we can solve increasingly more complex NLP tasks across various domains, including spreadsheets. This work investigates whether L…