4 papers · 1 filter
TabReX : Tabular Referenceless eXplainable Evaluation
Tejas Anvekar, Junha Park, Aparna Garimella +1
Evaluating the quality of tables generated by large language models (LLMs) remains an open challenge: existing metrics either flatten tables into text, ignoring structure, or rely…
TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation
Vihang Pancholi, Jainit Bafna, Tejas Anvekar +2
Evaluating tables qualitatively and quantitatively poses a significant challenge, as standard metrics often overlook subtle structural and content-level discrepancies. To address t…
DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity
Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi +3
Multimodal Large Language Models (MLLMs) can directly consume exam documents, threatening conventional assessments and academic integrity. We present DoPE (Decoy-Oriented Perturbat…
Integrity Shield A System for Ethical AI Use & Authorship Transparency in Assessments
Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi +3
Large Language Models (LLMs) can now solve entire exams directly from uploaded PDF assessments, raising urgent concerns about academic integrity and the reliability of grades and c…