works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.MM2026

RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +12

The paper presents RetroHolmes, a benchmark for testing vision‑language models on retrospective physical process reasoning—inferring hidden causes from image outcomes—and shows tha…

cs.AI2026

Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM Agents

Bin Wen, Ruoxuan Zhang, Yang Chen +2

Large language models (LLMs) have demonstrated strong potential in long-horizon decision-making tasks, such as embodied manipulation and web interaction. However, agents frequently…

cs.AI2026

MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents

Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou +7

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision…

cs.CV2025

CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation

Ruoxuan Zhang, Bin Wen, Hongxia Xie +5

Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusi…

cs.CV2025

RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

Ruoxuan Zhang, Jidong Gao, Bin Wen +4

Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained…

cs.CV2025

EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation

Cheng Zhang, Hongxia xie, Bin Wen +3

With the rapid advancement of diffusion models, text-to-image generation has achieved significant progress in image resolution, detail fidelity, and semantic alignment, particularl…