activity
20242026
collaborators

6 papers

cs.HC2026

PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data

Zeyu He, Xuan Qi, Subramanian Chidambaram +4

Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a n…

cs.CL2026

An Empirical Study of Automating Agent Evaluation

Kang Zhou, Sangmin Woo, Haibo Ding +14

Agent evaluation requires assessing complex multi-step behaviors involving tool use and intermediate reasoning, making it costly and expertise-intensive. A natural question arises:…

cs.CL2025

PILOT: Steering Synthetic Data Generation with Psychological & Linguistic Output Targeting

Caitlin Cisar, Emily Sheffield, Joshua Drake +5

Generative AI applications commonly leverage user personas as a steering mechanism for synthetic data generation, but reliance on natural language representations forces models to…

cs.CL2025

MetaSynth: Meta-Prompting-Driven Agentic Scaffolds for Diverse Synthetic Data Generation

Haris Riaz, Sourav Bhabesh, Vinayak Arannil +2

Recent smaller language models such Phi-3.5 and Phi-4 rely on synthetic data generated using larger Language models. Questions remain about leveraging synthetic data for other use…

cs.CL2024

ADEQA: A Question Answer based approach for joint ADE-Suspect Extraction using Sequence-To-Sequence Transformers

Vinayak Arannil, Tomal Deb, Atanu Roy

Early identification of Adverse Drug Events (ADE) is critical for taking prompt actions while introducing new drugs into the market. These ADEs information are available through va…

cs.CL2024

DoPAMine: Domain-specific Pre-training Adaptation from seed-guided data Mining

Vinayak Arannil, Neha Narwal, Sourav Sanjukta Bhabesh +5

Large Language Models (LLMs) have shown remarkable ability to generalize effectively across numerous industry domains while executing a range of tasks. Many of these competencies a…