activity
20182026
most citedAegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

2 citations · 6 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL20252 cited

Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails

Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar +4

As Large Language Models (LLMs) and generative AI become increasingly widespread, concerns about content safety have grown in parallel. Currently, there is a clear lack of high-qua…

cs.CL20232 cited

Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language

Di Jin, Shikib Mehri, Devamanyu Hazarika +4

Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF)…

cs.CL2022

Dialog Acts for Task-Driven Embodied Agents

Spandana Gella, Aishwarya Padmakumar, Patrick Lange +1

Embodied agents need to be able to interact in natural language understanding task descriptions and asking appropriate follow up questions to obtain necessary information to be eff…

cs.CL2022

On the Limits of Evaluating Embodied Agent Model Generalization Using Validation Sets

Hyounghun Kim, Aishwarya Padmakumar, Di Jin +2

Natural language guided embodied task completion is a challenging problem since it requires understanding natural language instructions, aligning them with egocentric visual observ…

cs.CL20211 cited

Generative Conversational Networks

Alexandros Papangelis, Karthik Gopalakrishnan, Aishwarya Padmakumar +3

Inspired by recent work in meta-learning and generative teaching networks, we propose a framework called Generative Conversational Networks, in which conversational agents learn to…

cs.CL20201 cited

Dialog as a Vehicle for Lifelong Learning

Aishwarya Padmakumar, Raymond J. Mooney

Dialog systems research has primarily been focused around two main types of applications - task-oriented dialog systems that learn to use clarification to aid in understanding a go…