5 papers
Speech Recognition on TV Series with Video-guided Post-ASR Correction
Haoyuan Yang, Yue Zhang, Liqiang Jing +1
Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assi…
Fine-grained and Explainable Factuality Evaluation for Multimodal Summarization
Yue Zhang, Jingxuan Zuo, Ke Su +1
Multimodal summarization aims to generate a concise summary based on the input text and image. However, the existing methods potentially suffer from unfactual output. To evaluate t…
Can Large Vision-Language Models Understand Multimodal Sarcasm?
Xinyu Wang, Yue Zhang, Liqiang Jing
Sarcasm is a complex linguistic phenomenon that involves a disparity between literal and intended meanings, making it challenging for sentiment analysis and other emotion-sensitive…
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
Shuo Yan, Ruochen Li, Ziming Luo +11
Large language model (LLM) agents have demonstrated remarkable potential in advancing scientific discovery. However, their capability in the fundamental yet crucial task of reprodu…
Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
Yue Zhang, Liqiang Jing, Vibhav Gogate
We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hy…