works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.SE2026

CoGate: Confidence-Gated Co-Decoding for Secure Code Generation

Minghao Hu, Lannan Luo, Allen Roush +1

The paper introduces CoGate, a method that uses the confidence of a security expert model to gate its influence during co-decoding for generating more secure code with large langua…

cs.AI2026

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

Phillip Howard, Xin Su, Allen Roush +2

High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques have helped mitigate drawbacks…

cs.CV2026

Mirage Probes: How Vision Models Fake Visual Understanding

Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao +5

Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirage behavior inflates benchmark scores with…

cs.AI2026

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Alexandra Yost, Shreyans Jain, Shivam Raval +6

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We prop…

cs.AI2026

Do Multi-Agents Dream of Electric Screens? Achieving Perfect Accuracy on AndroidWorld Through Task Decomposition

Pierre-Louis Favreau, Jean-Pierre Lo, Clement Guiguet +5

We present Minitap, a multi-agent system that achieves 100% success on the AndroidWorld benchmark, the first to fully solve all 116 tasks and surpassing human performance (80%). We…

cs.CL2025

A superpersuasive autonomous policy debating system

Allen Roush, Devin Gonier, John Hines +4

The capacity for highly complex, evidence-based, and strategically adaptive persuasion remains a formidable great challenge for artificial intelligence. Previous work, like IBM Pro…