4 papers
Reviewer Scores Are Not Comparable Across Research Areas in ML Peer Review
Binyan Xu, Xilin Dai, Fan Yang +1
Peer review at ML conferences increasingly relies on reviewer scores as the primary decision instrument. As submissions have scaled from thousands to tens of thousands per year, no…
LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via State Proprioception
Binyan Xu, Haitao Li, Kehuan Zhang
Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management agent- or system-controlled, bu…
Beyond Nodes vs. Edges: A Multi-View Fusion Framework for Provenance-Based Intrusion Detection
Fan Yang, Binyan Xu, Di Tang +1
Provenance-based intrusion detection has emerged as a promising approach for analyzing complex attack behaviors through system-level provenance graphs. However, existing defense me…
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
Weikai Lu, Ziqian Zeng, Kehua Zhang +5
Multimodal Large Language Models (MLLMs) are increasingly vulnerable to multimodal Indirect Prompt Injection (IPI) attacks, which embed malicious instructions in images, videos, or…