4 papers
MessyKitchens: Contact-rich object-level 3D scene reconstruction
Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati +1
Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large-scale data, recent methods achieve high performance i…
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
Yue Zhou, Jue Chen, Zilun Zhang +10
Remote sensing (RS) large vision-language models (LVLMs) have shown strong promise across visual grounding (VG) tasks. However, existing RS VG datasets predominantly rely on explic…
3D-CovDiffusion: 3D-Aware Diffusion Policy for Coverage Path Planning
Chenyuan Chen, Haoran Ding, Ran Ding +6
Diffusion models have shown strong potential for robot skill learning, yet their role in coverage path planning remains underexplored. In industrial surface processing (painting, p…
Narrative-Driven Travel Planning: Geoculturally-Grounded Script Generation with Evolutionary Itinerary Optimization
Ziyu Zhang, Ran Ding, Ying Zhu +2
To enhance tourists' experiences and immersion, this paper proposes a narrative-driven travel planning framework called NarrativeGuide, which generates a geoculturally-grounded nar…