works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry

Muxin Liu, Xiaoyang Lyu, Tianhe Ren +7

FoundationGeo is a two‑stage framework that first learns an affine‑invariant geometry model from a large multi‑domain dataset, then refines metric depth using lightweight pixel‑wis…

cs.CV2026

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

Yicheng Xiao, Wenhu Zhang, Lin Song +10

Image spatial editing performs geometry-driven transformations, allowing precise control over object layout and camera viewpoints. Current models are insufficient for fine-grained…

cs.LG2026

DBellQuant: Breaking the Bell with Double-Bell Transformation for LLMs Post Training Binarization

Zijian Ye, Wei Huang, Yifei Yu +3

Large language models (LLMs) demonstrate remarkable performance but face substantial computational and memory challenges that limit their practical deployment. Quantization has eme…

cs.AI2025

Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback

Aiden Yiliu Li, Bizhi Yu, Daoan Lei +2

GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual G…

cs.CV2025

SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features

Jinyuan Qu, Hongyang Li, Xingyu Chen +5

In this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training i…

cs.CV2025

Detect Anything via Next Point Prediction

Qing Jiang, Junan Huo, Xingyu Chen +6

Object detection has long been dominated by traditional coordinate regression-based models, such as YOLO, DETR, and Grounding DINO. Although recent efforts have attempted to levera…