From the 1 of 17 linked papers with an AI index.
17 papers
Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
Arnav Hiray, Agam Shah, Caleb Lu +3
The paper presents CreditCardQA, a benchmark of 1,800 real‑world credit‑card agreement questions for testing numerical reasoning in language models, and shows that Program‑of‑Thoug…
IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents
Michael Galarnyk, Siddharth Lohani, Vidhyakshaya Kannan +8
An Initial Public Offering (IPO) filing is a document released when a private firm goes public, allowing individual (retail) investors to purchase its shares. These filings describ…
Evolutionary Task Discovery: Advancing Reasoning Frontiers via Skill Composition and Complexity Scaling
Liqin Ye, Yanbin Yin, Michael Galarnyk +3
The reasoning frontier of Large Language Models (LLMs) has advanced significantly through modern post-training paradigms (e.g., Reinforcement Learning from Verifiable Rewards (RLVR…
Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editing
Rongzhi Zhang, Liqin Ye, Yuzhao Heng +5
Precise attribute intensity control--generating Large Language Model (LLM) outputs with specific, user-defined attribute intensities--is crucial for AI systems adaptable to diverse…
FinForge: Semi-Synthetic Financial Benchmark Generation
Glenn Matlin, Akhil Theerthala, Anant Gupta +4
Evaluating Language Models (LMs) in specialized, high-stakes domains such as finance remains a significant challenge due to the scarcity of open, high-quality, and domain-specific…
KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation
Nikita Tatarinov, Vidhyakshaya Kannan, Haricharana Srinivasa +7
We introduce KG-MuLQA (Knowledge-Graph-based Multi-Level Question-Answer Extraction): a framework that (1) extracts QA pairs at multiple complexity levels (2) along three key dimen…