Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
Aman Sharma, Sushrut Thorat, Paras Chopra
LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories. These benchmarks remain important, but…
cs.AI2026
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
Simardeep Singh, Paras Chopra
While large language models (LLMs) are trained purely on textual data, prior work has shown that their internal representations can exhibit rich geometric structure in embedding sp…
cs.AI2026
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
Aman Sharma, Paras Chopra
Large language models achieve near-ceiling performance on code generation benchmarks, yet most of the programming languages used by popular benchmarks such as SWE-bench and HumanEv…