6 papers · 1 filter
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
Divij Handa, Zehua Zhang, Amir Saeidi +4
Recent advancements in Large Language Model (LLM) safety have primarily focused on mitigating attacks crafted in natural language or common ciphers (e.g. Base64), which are likely…
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on -bench
Venkatesh Mishra, Amir Saeidi, Satyam Raj +5
Recent advances in reasoning and planning capabilities of large language models (LLMs) have enabled their potential as autonomous agents capable of tool use in dynamic environments…
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
Md Nayem Uddin, Amir Saeidi, Divij Handa +5
This paper introduces UnSeenTimeQA, a novel data contamination-free time-sensitive question-answering (TSQA) benchmark. It differs from existing TSQA benchmarks by avoiding web-sea…
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
Amir Saeidi, Shivanshu Verma, Aswin RRV +2
Reinforcement Learning with Human Feedback (RLHF) enhances the alignment of Large Language Models (LLMs). However, its limitations have led to the development of Direct Preference…
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
Amir Saeidi, Shivanshu Verma, Md Nayem Uddin +1
This study evaluates Direct Preference Optimization (DPO) and its variants for aligning Large Language Models (LLMs) with human preferences, testing three configurations: (1) with…
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation
Neeraj Varshney, Satyam Raj, Venkatesh Mishra +4
Large Language Models (LLMs) have achieved remarkable performance across a wide variety of natural language tasks. However, they have been shown to suffer from a critical limitatio…