5 papers
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
Bassam Adnan, Aviral Gupta, Sreemaee Akshathala +1
Benchmarks for large language models (LLMs) have progressed from snippet-level function generation to repository-level issue resolution, yet they overwhelmingly target implementati…
TabPFN Through The Looking Glass: An interpretability study of TabPFN and its internal representations
Aviral Gupta, Armaan Sethi, Dhruv Kumar
Tabular foundational models are pre-trained models designed for a wide range of tabular data tasks. They have shown strong performance across domains, yet their internal representa…
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
Abhipsa Basu, Aviral Gupta, Abhijnya Bhat +1
Image classification systems often inherit biases from uneven group representation in training data. For example, in face datasets for hair color classification, blond hair may be…
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
Aviral Gupta, Armaan Sethi, Ameesh Sethi
Neural network classifiers trained on datasets with uneven group representation often inherit class biases and learn spurious correlations. These models may perform well on average…
Emergent Stack Representations in Modeling Counter Languages Using Transformers
Utkarsh Tiwari, Aviral Gupta, Michael Hahn
Transformer architectures are the backbone of most modern language models, but understanding the inner workings of these models still largely remains an open problem. One way that…