From the 1 of 9 linked papers with an AI index.
9 papers
Infinity-Parser2 Technical Report
Zuming Huang, Jun Huang, Kexuan Ren +12
Infinity-Parser2 is a large multimodal model that uses a controllable synthetic data pipeline and multi‑task reinforcement learning to parse documents, offering two variants (Flash…
HetScene: Heterogeneity-Aware Diffusion for Dense Indoor Scene Generation
Zini Chen, Junming Huang, Rong Zhang +4
Generating controllable and physically plausible indoor scenes is a pivotal prerequisite for constructing high-fidelity simulation environments for embodied AI. However, existing d…
BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
Haozhe Zhang, Kaichen Liu, Miaomiao Chen +4
Industrial Computer-Aided Design (CAD) code generation requires models to produce executable parametric programs from visual or textual inputs. Beyond recognizing the outer shape o…
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
Yutao Tang, Cheng Zhao, Gaurav Mittal +4
Recent advances in 3D vision-language models (VLMs) highlight a strong potential for 3D scene understanding and reasoning. However, effectively tokenizing 3D scenes into holistic s…
FusionRF: High-Fidelity Satellite Neural Radiance Fields from Multispectral and Panchromatic Acquisitions
Michael Sprintson, Rama Chellappa, Cheng Peng
We introduce FusionRF, a novel framework for digital surface reconstruction from satellite multispectral and panchromatic images. Current work has demonstrated the increased accura…
World-in-World: World Models in a Closed-Loop World
Jiahan Zhang, Muqing Jiang, Nanru Dai +14
Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive pe…