3 papers
cs.CL2025
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems
Jisoo Lee, Raeyoung Chang, Dongwook Kwon +2
Multi-agent systems built on language models have shown strong performance on collaborative reasoning tasks. However, existing evaluations focus only on the correctness of the fina…
cs.AI2025
OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic Reflections
Manasa Bharadwaj, Nikhil Verma, Kevin Ferreira
Efforts to improve Large Language Model (LLM) agent performance on complex tasks have largely focused on fine-tuning and iterative self-correction. However, these approaches often…
cs.CL2025
The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
Nikhil Verma, Manasa Bharadwaj
Alignment tuning has enabled large language models to excel in reasoning, instruction-following, and minimizing harmful generations. However, despite their widespread deployment, t…