7 papers
From Image Generation to Infrastructure Design: a Multi-agent Pipeline for Street Design Generation
Chenguang Wang, Xiang Yan, Yilong Dai +2
Realistic visual renderings of street-design scenarios are essential for public engagement in active transportation planning. Traditional approaches are labor-intensive, hindering…
MusiXQA: Advancing Visual Music Understanding in Multimodal Large Language Models
Jian Chen, Wenye Ma, Penghang Liu +7
Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to…
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
Ming Li, Chenguang Wang, Yijun Liang +6
Recent agentic Multi-Modal Large Language Models (MLLMs) such as GPT-o3 have achieved near-ceiling scores on various existing benchmarks, motivating a demand for more challenging t…
Mosaic-IT: Cost-Free Compositional Data Synthesis for Instruction Tuning
Ming Li, Pei Chen, Chenguang Wang +5
Finetuning large language models with a variety of instruction-response pairs has enhanced their capability to understand and follow instructions. Current instruction tuning primar…
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
Ruxiao Chen, Chenguang Wang, Yuran Sun +2
Evacuation decision prediction is critical for efficient and effective wildfire response by helping emergency management anticipate traffic congestion and bottlenecks, allocate res…
Where You Go is Who You Are: Behavioral Theory-Guided LLMs for Inverse Reinforcement Learning
Yuran Sun, Susu Xu, Chenguang Wang +1
Big trajectory data hold great promise for human mobility analysis, but their utility is often constrained by the absence of critical traveler attributes, particularly sociodemogra…