5 papers
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
Shuai Wang, Daoan Zhang, Tianyi Bai +3
Humans can perceive and understand 3D space and long videos from sequential visual observations. But do vision-language models (VLMs) can? Recent work demonstrates that even state-…
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
Shengming Yuan, Xinyu Lyu, Shuailong Wang +3
Multimodal large language models (MLLMs) face an inherent trade-off between faithfulness and creativity, as different tasks require varying degrees of associative reasoning. Howeve…
AgenticMath: Enhancing LLM Reasoning via Agentic-based Math Data Generation
Xianyang Liu, Yilin Liu, Shuai Wang +5
The creation of high-quality datasets to improve Large Language Model (LLM) reasoning remains a significant challenge, as current methods often suffer from generating low-quality/i…
Seeing The Words: Evaluating AI-generated Biblical Art
Hidde Makimei, Shuai Wang, Willem van Peursen
The past years witnessed a significant amount of Artificial Intelligence (AI) tools that can generate images from texts. This triggers the discussion of whether AI can generate acc…
Surrealistic-like Image Generation with Vision-Language Models
Elif Ayten, Shuai Wang, Hjalmar Snoep
Recent advances in generative AI make it convenient to create different types of content, including text, images, and code. In this paper, we explore the generation of images in th…