4 papers
Factor(U,T): Controlling Untrusted AI by Monitoring their Plans
Edward Lue Chee Lip, Anthony Channg, Diana Kim +2
As AI capabilities advance, we increasingly rely on powerful models to decompose complex tasks $\unicode{x2013}$ but what if the decomposer itself is malicious? Factored cognition…
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
Glenn Zhang, Treasure Mayowa, Jason Fan +4
Producing trustworthy and reliable Large Language Models (LLMs) has become increasingly important as their usage becomes more widespread. Calibration seeks to achieve this by impro…
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
Aaron Sandoval, Cody Rushing
The field of AI Control seeks to develop robust control protocols, deployment safeguards for untrusted AI which may be intentionally subversive. However, existing protocols that re…
Adaptive Linguistic Prompting (ALP) Enhances Phishing Webpage Detection in Multimodal Large Language Models
Atharva Bhargude, Ishan Gonehal, Dave Yoon +4
Phishing attacks represent a significant cybersecurity threat, necessitating adaptive detection techniques. This study explores few-shot Adaptive Linguistic Prompting (ALP) in dete…