6 papers
AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems
Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov +4
Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating complicated descriptions of these problems…
Benchmarking Deflection and Hallucination in Large Vision-Language Models
Nicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro +2
Large Vision-Language Models (LVLMs) increasingly rely on retrieval to answer knowledge-intensive multimodal questions. Existing benchmarks overlook conflicts between visual and te…
GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
Ionut-Teodor Sorodoc, Leonardo F. R. Ribeiro, Rexhina Blloshmi +2
We present GaRAGe, a large RAG benchmark with human-curated long-form answers and annotations of each grounding passage, allowing a fine-grained evaluation of whether LLMs can iden…
Prompting open-source and commercial language models for grammatical error correction of English learner text
Christopher Davis, Andrew Caines, Ãistein Andersen +6
Thanks to recent advances in generative AI, we are able to prompt large language models (LLMs) to produce texts which are fluent and grammatical. In addition, it has been shown tha…
The Amazon Nova Family of Models: Technical Report and Model Card
Amazon AGI, Aaron Langford, Aayush Shah +783
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout
Chiyu Max Jiang, Yijing Bai, Andre Cornman +12
Realistic and interactive scene simulation is a key prerequisite for autonomous vehicle (AV) development. In this work, we present SceneDiffuser, a scene-level diffusion prior desi…