15 papers
A Good Talk Does not Look Like a Summary, It Teaches You! Measuring Takeaways from Paper-to-Video Talks
Ishani Mondal, Aparna Garimella, Ananya Sai +2
Automatically generated videos from scientific papers are increasingly used for education and research dissemination. However, existing evaluation metrics mainly measure visual qua…
An Answer is just the Start: Related Insight Generation for Open-Ended Document-Grounded QA
Saransh Sharma, Pritika Ramu, Aparna Garimella +1
Answering open-ended questions remains challenging for AI systems because it requires synthesis, judgment, and exploration beyond factual retrieval, and users often refine answers…
TabReX : Tabular Referenceless eXplainable Evaluation
Tejas Anvekar, Junha Park, Aparna Garimella +1
Evaluating the quality of tables generated by large language models (LLMs) remains an open challenge: existing metrics either flatten tables into text, ignoring structure, or rely…
Decisive: Guiding User Decisions with Optimal Preference Elicitation from Unstructured Documents
Akriti Jain, Anish Mulay, Divyansh Verma +3
Decision-making is a cognitively intensive task that requires synthesizing relevant information from multiple unstructured sources, weighing competing factors, and incorporating su…
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
Jeonghyun Park, Ingeol Baek, Seunghyun Yoon +5
Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resolution. Since ambiguity can occur…
Moneyball with LLMs: Analyzing Tabular Summarization in Sports Narratives
Ritam Upadhyay, Naman Ahuja, Rishabh Baral +2
Large language model (LLM) approaches to tabular summarization rely on extensive prompt engineering, decomposition pipelines, or entity-level intermediate representations to achiev…