activity
20242026
collaborators

6 papers

cs.CV2026

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Yuancheng Wei, Haojie Zhang, Linli Yao +7

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change…

cs.CV2025

ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models

Xusen Hei, Jiali Chen, Jinyu Yang +2

As multimodal large language models (MLLMs) frequently exhibit errors in complex video reasoning scenarios, correcting these errors is critical for uncovering their weaknesses and…

cs.CV2025

ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments

Jiali Chen, Yujie Jia, Zihan Wu +6

Experiment commentary is crucial in describing the experimental procedures, delving into underlying scientific principles, and incorporating content-related safety guidelines. In p…

cs.CV2025

CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction

Jiali Chen, Xusen Hei, HongFei Liu +5

Computer-aided design (CAD) is crucial in prototyping 3D objects through geometric instructions (i.e., CAD programs). In practical design workflows, designers often engage in time-…

cs.CL2025

Classic4Children: Adapting Chinese Literary Classics for Children with Large Language Model

Jiali Chen, Xusen Hei, Yuqi Xue +3

Chinese literary classics hold significant cultural and educational value, offering deep insights into morality, history, and human nature. These works often include classical Chin…

cs.CV2024

Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor

Jiali Chen, Xusen Hei, Yuqi Xue +4

Large multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual com…