collaborators

7 papers

cs.CR2026

Evaluating Nova 2.0 Lite model under Amazon's Frontier Model Safety Framework

Satyapriya Krishna, Matteo Memelli, Tong Wang +5

Amazon published its Frontier Model Safety Framework (FMSF) as part of the Paris AI summit, following which we presented a report on Amazon's Premier model. In this report, we pres…

cs.AI2026

Embedded AI Companion System on Edge Devices

Rahul Gupta, Stephen D. H. Hsu

Computational resource constraints on edge devices make it difficult to develop a fully embedded AI companion system with a satisfactory user experience. AI companion and memory sy…

cs.CV2025

VMDT: Decoding the Trustworthiness of Video Foundation Models

Yujin Potter, Zhun Wang, Nicholas Crispino +11

As foundation models become more sophisticated, ensuring their trustworthiness becomes increasingly critical; yet, unlike text and image, the video modality still lacks comprehensi…

cs.CL2025

D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models

Satyapriya Krishna, Andy Zou, Rahul Gupta +6

The safety and alignment of Large Language Models (LLMs) are critical for their responsible deployment. Current evaluation methods predominantly focus on identifying and preventing…

cs.CL2025

Diagnosing Memorization in Chain-of-Thought Reasoning, One Token at a Time

Huihan Li, You Chen, Siyuan Wang +4

Large Language Models (LLMs) perform well on reasoning benchmarks but often fail when inputs alter slightly, raising concerns about the extent to which their success relies on memo…

cs.CR2025

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework

Satyapriya Krishna, Ninareh Mehrabi, Abhinav Mohanty +4

Nova Premier is Amazon's most capable multimodal foundation model and teacher for model distillation. It processes text, images, and video with a one-million-token context window,…