most citedMagentic-UI: Towards Human-in-the-loop Agentic Systems

3 citations · 4 across the 7 of their papers we have counts for

collaborators

7 papers

cs.AI2026

SentinelBench: A Benchmark for Long-Running Monitoring Agents

Matheus Kunzler Maldaner, Adam Fourney, Amanda Swearngin +5

AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: issuing tool calls, refreshin…

cs.CL2026

Plato's Cave: A Human-Centered Research Verification System

Matheus Kunzler Maldaner, Raul Valle, Junsung Kim +9

The growing publication rate of research papers has created an urgent need for better ways to fact-check information, assess writing quality, and identify unverifiable claims. We p…

cs.CV2026

MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment

Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin +8

Vision-Language Models (VLMs) continue to struggle to make morally salient judgments in multimodal and socially ambiguous contexts. Prior works typically rely on binary or pairwise…

cs.HC2025

Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies

Matheus Kunzler Maldaner, Wesley Hanwen Deng, Jason I. Hong +2

While generative AI systems have gained popularity in diverse applications, their potential to produce harmful outputs limits their trustworthiness and utility. A small but growing…

cs.AI2025★ 3 cited

Magentic-UI: Towards Human-in-the-loop Agentic Systems

Hussein Mozannar, Gagan Bansal, Cheng Tan +17

AI agents powered by large language models are increasingly capable of autonomously completing complex, multi-step tasks using external tools. Yet, they still fall short of human-l…

cs.HC2025★ 1 cited

MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI

Matheus Kunzler Maldaner, Wesley Hanwen Deng, Jason Hong +2

While generative AI systems have gained popularity in diverse applications, their potential to produce harmful outputs limits their trustworthiness and usability in different appli…