From the 1 of 7 linked papers with an AI index.
7 papers
RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning
Ruoxuan Zhang, Qiyun Zheng, Siyu Wu +12
The paper presents RetroHolmes, a benchmark for testing vision‑language models on retrospective physical process reasoning—inferring hidden causes from image outcomes—and shows tha…
Aligning Progress and Feasibility: A Neuro-Symbolic Dual Memory Framework for Long-Horizon LLM Agents
Bin Wen, Ruoxuan Zhang, Yang Chen +2
Large language models (LLMs) have demonstrated strong potential in long-horizon decision-making tasks, such as embodied manipulation and web interaction. However, agents frequently…
MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
Ruoxuan Zhang, Qiyun Zheng, Zhiyu Zhou +7
Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision…
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
Ruoxuan Zhang, Bin Wen, Hongxia Xie +5
Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusi…
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
Ruoxuan Zhang, Jidong Gao, Bin Wen +4
Creating recipe images is a key challenge in food computing, with applications in culinary education and multimodal recipe assistants. However, existing datasets lack fine-grained…
EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation
Cheng Zhang, Hongxia xie, Bin Wen +3
With the rapid advancement of diffusion models, text-to-image generation has achieved significant progress in image resolution, detail fidelity, and semantic alignment, particularl…