4 papers
Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection
Chinmay Pushkar, Sanchit Kabra, Dhruv Kumar +1
Large Language Models (LLMs) have demonstrated significant potential in automated software security, particularly in vulnerability detection. However, existing benchmarks primarily…
RUST-BENCH: Benchmarking LLM Reasoning on Unstructured Text within Structured Tables
Nikhil Abhyankar, Purvi Chaurasia, Sanchit Kabra +3
Existing tabular reasoning benchmarks mostly test models on small, uniform tables, underrepresenting the complexity of real-world data and giving an incomplete view of Large Langua…
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
Sanchit Kabra, Akshita Jha, Chandan K. Reddy
Recent advances in large-scale generative language models have shown that reasoning capabilities can significantly improve model performance across a variety of tasks. However, the…
Biased or Flawed? Mitigating Stereotypes in Generative Language Models by Addressing Task-Specific Flaws
Akshita Jha, Sanchit Kabra, Chandan K. Reddy
Recent studies have shown that generative language models often reflect and amplify societal biases in their outputs. However, these studies frequently conflate observed biases wit…