From the 1 of 6 linked papers with an AI index.
6 papers
CinemaTraj: Composing Atomic Camera Trajectories for 3D Scenes with LLM Agents
Qianru Li, Xuyang Chen, Erkin Türköz +5
CinemaTraj generates cinematic camera movements in 3D environments from natural language prompts by using an LLM agent that reasons over a structured 3D scene graph and produces co…
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
Hari Krishna Gadi, Daniel Matos, Hongyi Luo +4
Visual geolocalization, the task of predicting where an image was taken, remains challenging due to global scale, visual ambiguity, and the inherently hierarchical structure of geo…
Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving
Xuyang Chen, Conglang Zhang, Chuanheng Fu +10
Driven by the emergence of Controllable Video Diffusion, existing Sim2Real methods for autonomous driving video generation typically rely on explicit intermediate representations t…
MeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusion
Xuyang Chen, Zhijun Zhai, Kaixuan Zhou +9
Mesh models have become increasingly accessible for numerous cities; however, the lack of realistic textures restricts their application in virtual urban navigation and autonomous…
TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
Hongyi Luo, Qing Cheng, Daniel Matos +7
Humans can interpret geospatial information through natural language, while the geospatial cognition capabilities of Large Language Models (LLMs) remain underexplored. Prior resear…
Box2Poly: Memory-Efficient Polygon Prediction of Arbitrarily Shaped and Rotated Text
Xuyang Chen, Dong Wang, Konrad Schindler +4
Recently, Transformer-based text detection techniques have sought to predict polygons by encoding the coordinates of individual boundary vertices using distinct query features. How…