5 papers
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li, Yimin Liu, Wenbo Chen +75
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapid adoption, there is no standard way to m…
DeployBench: Benchmarking LLM Agents for Research Artifact Deployment
Yuanli Wang, Yaoyao Qian, Yue Zhang +8
LLM agents have made rapid progress on software engineering and ML research tasks, but these advances often assume access to a working runnable environment. For research artifacts…
Speech Recognition on TV Series with Video-guided Post-ASR Correction
Haoyuan Yang, Yue Zhang, Liqiang Jing +1
Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assi…
Can Large Vision-Language Models Understand Multimodal Sarcasm?
Xinyu Wang, Yue Zhang, Liqiang Jing
Sarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive…
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
Yue Zhang, Liqiang Jing, Vibhav Gogate
We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hy…