From the 1 of 10 linked papers with an AI index.
10 papers
Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning
Jiahao Shao, Yuanbo Yang, Yiyi Liao +3
Tool-augmented vision-language models increasingly "think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind te…
AutoPath: Learning Transferable Goal-Conditioned Stochastic Path Prior for Safe Navigation Without Human Demonstrations
Ziyang Zhang, Boyang Zhou, Zesong Yang +8
The paper proposes a goal‑conditioned stochastic path prior that learns a transferable distribution over local navigation paths from limited observations, enabling safe, multimodal…
GARDEN: Gravity-Aligned Reconstruction of Disentangled ENvironments from RGB images
Jiahao Sun, Dingkun Wei, Zehong Shen +3
Converting multi-view RGB observations into simulation-ready 3D environments remains challenging because current reconstruction pipelines produce monolithic scene representations w…
AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation
Haobo Li, Yanhong Zeng, Yunhong Lu +6
We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt adversarial distillation bu…
FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views
Shangzhan Zhang, Jianyuan Wang, Yinghao Xu +5
We present FLARE, a feed-forward model designed to infer high-quality camera poses and 3D geometry from uncalibrated sparse-view images (i.e., as few as 2-8 inputs), which is a cha…
ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generation
Hanlei Guo, Jiahao Shao, Xinya Chen +4
Recent advancements in 3D object generation using diffusion models have achieved remarkable success, but generating realistic 3D urban scenes remains challenging. Existing methods…