2 papers
cs.AI2026
LH-Bench: Skill-Grounded Evaluation of Long-Horizon Agents on Subjective Enterprise Tasks
Abhishek Chandwani, Ishan Gupta
Large language models excel on objectively verifiable tasks such as math and programming, where evaluation reduces to unit tests or a single correct answer. In contrast, real-world…
cs.CV2025
Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation
Arnav Gupta, Gurekas Singh Sahney, Hardik Rathi +4
Evaluating short-form video content requires moving beyond surface-level quality metrics toward human-aligned, multimodal reasoning. While existing frameworks like VideoScore-2 ass…