collaborators

5 papers

cs.CE2026

MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark

Cui Yakun, Yanting Zhang, Zhu Lei +5

The advent of multi-modal language models (MLLMs) has spurred research into their application across various table understanding tasks. However, their performance in credit table u…

cs.AI2025

MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data

Zhenghao Zhu, Chuxue Cao, Sirui Han +4

In medical data analysis, extracting deep insights from complex, multi-modal datasets is essential for improving patient care, increasing diagnostic accuracy, and optimizing health…

cs.CV2025

Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection

Cui Yakun, Peng Qi, Fushuo Huo +6

The advent of multi-modal large language models (MLLMs) has greatly advanced research on video fake news detection (VFND) tasks. Existing benchmarks typically focus on the detectio…

cs.CL2025

SafeLawBench: Towards Safe Alignment of Large Language Models

Chuxue Cao, Han Zhu, Jiaming Ji +7

With the growing prevalence of large language models (LLMs), the safety of LLMs has raised significant concerns. However, there is still a lack of definitive standards for evaluati…

cs.CL2025

Measuring Hong Kong Massive Multi-Task Language Understanding

Chuxue Cao, Zhenghao Zhu, Junqi Zhu +6

Multilingual understanding is crucial for the cross-cultural applicability of Large Language Models (LLMs). However, evaluation benchmarks designed for Hong Kong's unique linguisti…