2 papers
cs.AI2026
LinAlg-Bench: A Forensic Benchmark Revealing Structural Failure Modes in LLM Mathematical Reasoning
Shradha Agarwal, Deepak Rajbhar, Tariq J
We introduce LinAlg-Bench, a diagnostic benchmark evaluating 10 frontier large language models on structured linear algebra computation across a strict dimensional gradient of 3x3,…
cs.IR2025
A Comparative Study of PDF Parsing Tools Across Diverse Document Categories
Narayan S. Adhikari, Shradha Agarwal
PDF is one of the most prominent data formats, making PDF parsing crucial for information extraction and retrieval, particularly with the rise of RAG systems. While various PDF par…