activity
20242026
collaborators

6 papers

cs.AI2026

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Tianyou Wang, Chongyang Gao, Kezhen Chen +7

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the…

cs.CL2026

Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO

Mengyu Xu, Qiaoxin Yang, Zhihan Liu +4

Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can receive answers of considerably different q…

cs.AI2026

Classroom Final Exam: An Instructor-Tested Reasoning Benchmark

Chongyang Gao, Diji Yang, Shuyan Zhou +4

We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench…

cs.CV2025

SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning

Andrew Li, Rahul Thapa, Rahul Chalamala +3

Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-sou…

cs.CL2024

RedPajama: an Open Dataset for Training Large Language Models

Maurice Weber, Daniel Fu, Quentin Anthony +16

Large language models are increasingly becoming a cornerstone technology in artificial intelligence, the sciences, and society as a whole, yet the optimal strategies for dataset co…

cs.CV2024

Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Rahul Thapa, Kezhen Chen, Ian Covert +4

Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native re…