works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cs.CL2026

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Tuo Liang, Zhe Hu, Disheng Liu +2

Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms, shared cultural knowledge, and communic…

cs.RO2026

ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response

Xiaomeng Zhu, Fengming Zhu, Weijie Zhou +8

The paper introduces ProAct-75, a benchmark of 75 proactive tasks with step‑level annotations and task graphs, and presents ProAct-Helper, a multimodal LLM that uses these graphs f…

cs.CL2026

VIABLE: A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

Yi Zhao, Siqi Wang, Zhe Hu +2

AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer a promising alternative, al…

cs.CV2026

When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Tuo Liang, Zhe Hu, Jing Li +8

Understanding humor-particularly when it involves complex, contradictory narratives that require comparative reasoning-remains a significant challenge for large vision-language mod…

cs.CL2026

Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

Zhe Hu, Tuo Liang, Jing Li +5

Recent advancements in large multimodal language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models still struggle with understanding t…

cs.CL2025

Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning

Zhe Hu, Jing Li, Zhongzhu Pu +2

Vision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This pape…