9 papers · 1 filter
FinBalance: A Multi-Document Accounting Reconciliation Benchmark
Sasank Tumpati, Devansh Agarwal, Ayush Kedia +6
Existing financial-NLP benchmarks mostly evaluate prepared artifacts such as filings, tables, or extracted values. Real accounting begins earlier: source documents must be reconcil…
Navigating the Reality Gap: On-Device Continual Adaptation of ASR for Clinical Telephony
Darshil Chauhan, Adityasinh Solanki, Vansh Patel +5
Automatic Speech Recognition (ASR) can significantly reduce documentation burden in clinical workflows, but standard models degrade sharply in real-world telephony settings where n…
Cultural Fidelity in English-to-Hindi Translation: A Preservation-Fluency Frontier for Gender Recoverability
Samyak Savi, Chavi Gupta, Shreyas Gantayet +2
Generative translation systems are cultural technologies because they decide how socially meaningful cues are rendered within culturally specific grammatical systems. We study one…
Beyond Accuracy: Diagnosing Algebraic Reasoning Failures in LLMs Across Nine Complexity Dimensions
Parth Patil, Dhruv Kumar, Yash Sinha +1
Algebraic reasoning remains one of the most informative stress tests for large language models, yet current benchmarks provide no mechanism for attributing failure to a specific ca…
Measuring Representation Robustness in Large Language Models for Geometry
Vedant Jawandhia, Yash Sinha, Murari Mandal +2
Large language models (LLMs) are increasingly evaluated on mathematical reasoning, yet their robustness to equivalent problem representations remains poorly understood. In geometry…
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
Devanshu Sahoo, Manish Prasad, Vasudev Majhi +5
The rapid integration of Large Language Models (LLMs) into educational assessment rests on the unverified assumption that instruction following capability translates directly to ob…