5 papers
An Emergent Mirage: Is Emergent Misalignment and Realignment Indeed a Robust Phenomenon?
Abhinav Rao, Liancheng Gong, Bin Hu +1
Recent work has reported Emergent Misalignment (EM), where language models fine-tuned on narrow, domain-specific misaligned datasets abruptly acquire broadly misaligned behavior, a…
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
Weidi Luo, Tianyu Lu, Qiming Zhang +8
Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models enable impressive reasoning cap…
Strong Memory, Weak Control: An Empirical Study of Executive Functioning in LLMs
Karin de Langis, Jong Inn Park, Bin Hu +5
Working memory, or the ability to hold and manipulate information in the mind, is a critical component of human intelligence and executive functioning. It is correlated with perfor…
Your Harness is Not Secure: Benchmarking Real-world Threat of Command Line Interface Agent
Weidi Luo, Qiming Zhang, Tianyu Lu +9
Command-line interface (CLI) agents powered by large language models (LLMs) can interpret natural-language requests, plan multi-step tasks, execute shell commands, and modify files…
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
Karin de Langis, Jong Inn Park, Andreas Schramm +5
Large language models (LLMs) exhibit increasingly sophisticated linguistic capabilities, yet the extent to which these behaviors reflect human-like cognition versus advanced patter…