8 papers
What MLLMs Learn about When they Learn about Multimodal Reasoning
Jiwan Chung, Neel Joshi, Pratyusha Sharma +2
Evaluation of multimodal reasoning models is typically reduced to a single accuracy score, implicitly treating reasoning as a unitary capability. We introduce MathLens, a benchmark…
Phi-4-reasoning-vision-15B Technical Report
Jyoti Aneja, Michael Harrison, Neel Joshi +3
We present Phi-4-reasoning-vision-15B, a compact open-weight multimodal reasoning model, and share the motivations, design choices, experiments, and learnings that informed its dev…
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
Natasha Butt, Varun Chandrasekaran, Neel Joshi +2
Evaluation insights are limited by the availability of high-quality benchmarks. As models evolve, there is a need to create benchmarks that can measure progress on new and complex…
Phi-4-reasoning Technical Report
Marah Abdin, Sahaj Agarwal, Ahmed Awadallah +20
We introduce Phi-4-reasoning, a 14-billion parameter reasoning model that achieves strong performance on complex reasoning tasks. Trained via supervised fine-tuning of Phi-4 on car…
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
Vidhisha Balachandran, Jingya Chen, Lingjiao Chen +8
Inference-time scaling can enhance the reasoning capabilities of large language models (LLMs) on complex problems that benefit from step-by-step problem solving. Although lengtheni…
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
Siddharth Joshi, Besmira Nushi, Vidhisha Balachandran +4
Vision-language models (VLMs) are highly effective but often underperform on specialized tasks; for example, Llava-1.5 struggles with chart and diagram understanding due to scarce…