7 papers
APD-Agents: A Large Language Model-Driven Multi-Agents Collaborative Framework for Automated Page Design
Xinpeng Chen, Xiaofeng Han, Kaihao Zhang +6
Layout design is a crucial step in developing mobile app pages. However, crafting satisfactory designs is time-intensive for designers: they need to consider which controls and con…
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
Junhao Wu, Xiuer Gu, Zhiying Li +6
Evaluating human actions with clear and detailed feedback is important in areas such as sports, healthcare, and robotics, where decisions rely not only on final outcomes but also o…
SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
Xianghui Ze, Beiyi Zhu, Zhenbo Song +2
Generating multiview-consistent ground-level scenes from satellite imagery is a challenging task with broad applications in simulation, autonomous navigation, and digit…
Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control
Xianghui Ze, Zhenbo Song, Qiwei Wang +2
Generating street-view images from satellite imagery is a challenging task, particularly in maintaining accurate pose alignment and incorporating diverse environmental conditions.…
GIDP: Learning a Good Initialization and Inducing Descriptor Post-enhancing for Large-scale Place Recognition
Zhaoxin Fan, Zhenbo Song, Hongyan Liu +1
Large-scale place recognition is a fundamental but challenging task, which plays an increasingly important role in autonomous driving and robotics. Existing methods have achieved a…
Incorporating Orientations into End-to-end Driving Model for Steering Control
Peng Wan, Zhenbo Song, Jianfeng Lu
In this paper, we present a novel end-to-end deep neural network model for autonomous driving that takes monocular image sequence as input, and directly generates the steering cont…