9 papers
Denoising Models Develop Human-Like Perceptual Illusion Representations Across Architectures
Gautam Ranka, Paras Chopra
Deep neural networks trained on natural images are shown to produce outputs consistent with human observers for brightness illusions. While this phenomenon has been documented acro…
Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution
Aadi Narayana Varma Dantuluri, Sushrut Thorat, Paras Chopra
Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI systems from crediting work. As a boundary tes…
Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages
Aman Sharma, Sushrut Thorat, Paras Chopra
LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories. These benchmarks remain important, but…
Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark
Pratham Singla, Shivank Garg, Vihan Singh +1
Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors, the question's phrasing toge…
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
Simardeep Singh, Paras Chopra
While large language models (LLMs) are trained purely on textual data, prior work has shown that their internal representations can exhibit rich geometric structure in embedding sp…
EsoLang-Bench: Evaluating Genuine Reasoning in Large Language Models via Esoteric Programming Languages
Aman Sharma, Paras Chopra
Large language models achieve near-ceiling performance on code generation benchmarks, yet most of the programming languages used by popular benchmarks such as SWE-bench and HumanEv…