works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.MM2026

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +12

The paper presents RetroHolmes, a benchmark for testing vision‑language models on retrospective physical process reasoning—inferring hidden causes from image outcomes—and shows tha…

cs.CV2026

TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection

Jian-Yu Jiang-Lin, Kang-Yang Huang, Ling Zou +11

Advances in generative modeling have made it increasingly easy to fabricate realistic portrayals of individuals, creating serious risks for security, communication, and public trus…

cs.AI2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou +7

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision…

cs.CV2025

CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation

Ruoxuan Zhang, Bin Wen, Hongxia Xie +5

Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusi…

cs.CV2025

RecipeGen: A Benchmark for Real-World Recipe Image Generation

Ruoxuan Zhang, Hongxia Xie, Yi Yao +6

Recipe image generation is an important challenge in food computing, with applications from culinary education to interactive recipe platforms. However, there is currently no real-…

cs.MM2024

ReCorD: Reasoning and Correcting Diffusion for HOI Generation

Jian-Yu Jiang-Lin, Kang-Yang Huang, Ling Lo +5

Diffusion models revolutionize image generation by leveraging natural language to guide the creation of multimedia content. Despite significant advancements in such generative mode…