Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FinRule-Bench: A Benchmark for Joint Reasoning over Financial Tables and Principles
Arun Vignesh Malarkkan, Manan Roy Choudhury, Guangwei Zhang +4
Large language models (LLMs) are increasingly applied to financial analysis, yet their ability to audit structured financial statements under explicit accounting principles remains…
cs.AI2026
Better Call CLAUSE: A Discrepancy Benchmark for Auditing LLMs Legal Reasoning Capabilities
Manan Roy Choudhury, Adithya Chandramouli, Mannan Anand +1
The rapid integration of large language models (LLMs) into high-stakes legal work has exposed a critical gap: no benchmark exists to systematically stress-test their reliability ag…