activity
20242026
collaborators

5 papers

cs.CV2026

Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA

Tong Wu, Thanet Markchom

Visual Question Answering (VQA) for stylised cartoon imagery presents challenges, such as interpreting exaggerated visual abstraction and narrative-driven context, which are not ad…

cs.SD2025

Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment

Jiaying Hong, Ting Zhu, Thanet Markchom +1

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches o…

cs.CL2025

NCL-UoR at SemEval-2025 Task 3: Detecting Multilingual Hallucination and Related Observable Overgeneration Text Spans with Modified RefChecker and Modified SeflCheckGPT

Jiaying Hong, Thanet Markchom, Jianfei Xu +2

SemEval-2025 Task 3 (Mu-SHROOM) focuses on detecting hallucinations in content generated by various large language models (LLMs) across multiple languages. This task involves not o…

cs.CL2025

UoR-NCL at SemEval-2025 Task 1: Using Generative LLMs and CLIP Models for Multilingual Multimodal Idiomaticity Representation

Thanet Markchom, Tong Wu, Liting Huang +1

SemEval-2025 Task 1 focuses on ranking images based on their alignment with a given nominal compound that may carry idiomatic meaning in both English and Brazilian Portuguese. To a…

cs.CV2024

CRRG-CLIP: Automatic Generation of Chest Radiology Reports and Classification of Chest Radiographs

Jianfei Xu, Thanet Markchom, Huizhi Liang

The complexity of stacked imaging and the massive number of radiographs make writing radiology reports complex and inefficient. Even highly experienced radiologists struggle to mai…