3 papers
cs.AI2026
Agents of Chaos
Natalie Shapira, Chris Wendler, Avery Yen +35
We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord acc…
cs.CL2026
Emergent Search and Backtracking in Latent Reasoning Models
Jasmine Cui, Charles Ye
What happens when a language model thinks without words? Standard reasoning LLMs verbalize intermediate steps as chain-of-thought; latent reasoning transformers (LRTs) instead perf…
cs.LG2026
Efficient Representations are Controllable Representations
Charles Ye, Jasmine Cui
What is the most brute-force way to install interpretable, controllable features into a model's activations? Controlling how LLMs internally represent concepts typically requires s…