From the 1 of 6 linked papers with an AI index.
6 papers
Evaluating Rational Contracting in Natural Language
Bhavyesh Sajja, Max Kleiman-Weiner, Roger Zimmermann +1
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agent…
FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction
Chaoqun Yang, Fengbin Zhu, Xinyu Lin +5
Financial indicators are essential tools for transforming raw financial data into interpretable measures for various downstream tasks, such as valuation, risk assessment, and econo…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
Make Your LVLM KV Cache More Lightweight
Xihao Chen, Yangyang Guo, Roger Zimmermann
Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in Large Language Models (L…
MM-PhyRLHF: Reinforcement Learning Framework for Multimodal Physics Question-Answering
Janak Kapuriya, Chhavi Kirtani, Apoorv Singh +7
Recent advancements in LLMs have shown their significant potential in tasks like text summarization and generation. Yet, they often encounter difficulty while solving complex physi…
Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring
Avinash Anand, Raj Jaiswal, Abhishek Dharmadhikari +6
This paper presents GPSM4K, a comprehensive geometry multimodal dataset tailored to augment the problem-solving capabilities of Large Vision Language Models (LVLMs). GPSM4K encompa…