From the 1 of 12 linked papers with an AI index.
12 papers
FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry
Muxin Liu, Xiaoyang Lyu, Tianhe Ren +7
FoundationGeo is a two‑stage framework that first learns an affine‑invariant geometry model from a large multi‑domain dataset, then refines metric depth using lightweight pixel‑wis…
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Yicheng Xiao, Wenhu Zhang, Lin Song +10
Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained…
DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization
Zijian Ye, Wei Huang, Yifei Yu +3
Large language models (LLMs) demonstrate remarkable performance but face substantial computational and memory challenges that limit their practical deployment. Quantization has eme…
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
Aiden Yiliu Li, Bizhi Yu, Daoan Lei +2
GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual G…
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
Jinyuan Qu, Hongyang Li, Xingyu Chen +5
In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training i…
Detect Anything via Next Point Prediction
Qing Jiang, Junan Huo, Xingyu Chen +6
Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…