3 papers
cs.LG2026
What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus
Edward Lue Chee Lip, Boden Moraski, Tim Knappe +4
Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capa…
cs.CL2026
Detecting Hidden Chain-of-Thought in Large Language Models with Linguistic, Behavioral, and Mechanistic Indicators
Armaan Singh, Ryan Trinh Le, Jasmine Kaur +5
Large language models often answer complex reasoning questions without revealing intermediate steps, raising whether they reason latently or complete patterns. We propose the Hidde…
cs.CR2025
Factor(U,T): Controlling Untrusted AI by Monitoring their Plans
Edward Lue Chee Lip, Anthony Channg, Diana Kim +2
As AI capabilities advance, we increasingly rely on powerful models to decompose complex tasks $\unicode{x2013}$ but what if the decomposer itself is malicious? Factored cognition…