4 papers
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
Ãmer Faruk Akgül, Rajgopal Kannan, Willie Neiswanger +1
Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not teach new strategies; it redist…
Estimating LLM Consistency: A User Baseline vs Surrogate Metrics
Xiaoyuan Wu, Weiran Lin, Omer Akgul +1
Large language models (LLMs) are prone to hallucinations and sensitive to prompt perturbations, often resulting in inconsistent or unreliable generated text. Different methods have…
When AIOps Become "AI Oops": Subverting LLM-driven IT Operations via Telemetry Manipulation
Dario Pasquini, Evgenios M. Kornaropoulos, Giuseppe Ateniese +3
AI for IT Operations (AIOps) is transforming how organizations manage complex software systems by automating anomaly detection, incident diagnosis, and remediation. Modern AIOps so…
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
Weiran Lin, Anna Gerchanovsky, Omer Akgul +3
Writing effective prompts for large language models (LLM) can be unintuitive and burdensome. In response, services that optimize or suggest prompts have emerged. While such service…