6 papers
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
Miles Q. Li, Benjamin C. M. Fung, Martin Weiss +3
As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Cur…
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
Miles Q. Li, Benjamin C. M. Fung, Boyang Li +2
Existing white-box jailbreak attacks against aligned LLMs typically append discrete adversarial suffixes to the user prompt, which visibly alters the prompt and operates in a combi…
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
Miles Q. Li, Benjamin C. M. Fung, Boyang Li +2
The rapid deployment of LLM-based autonomous agents has introduced safety risks that extend far beyond traditional LLM concerns, prompting a proliferation of safety benchmarks sinc…
Security Concerns for Large Language Models: A Survey
Miles Q. Li, Benjamin C. M. Fung
Large Language Models (LLMs) such as ChatGPT and its competitors have caused a revolution in natural language processing, but their capabilities also introduce new security vulnera…
Training Dynamics of a 1.7B LLaMa Model: A Data-Efficient Approach
Miles Q. Li, Benjamin C. M. Fung, Shih-Chia Huang
Pretraining large language models is a complex endeavor influenced by multiple factors, including model architecture, data quality, training continuity, and hardware constraints. I…
On the Effectiveness of Incremental Training of Large Language Models
Miles Q. Li, Benjamin C. M. Fung, Shih-Chia Huang
Training large language models is a computationally intensive process that often requires substantial resources to achieve state-of-the-art results. Incremental layer-wise training…