2 papers
cs.CV2025
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
Jiajie Guo, Qingpeng Zhu, Jin Zeng +3
Multimodal large language models (MLLMs) have achieved significant progress in image and language tasks due to the strong reasoning capability of large language models (LLMs). Neve…
cs.CV2025
Towards Robust Time-of-Flight Depth Denoising with Confidence-Aware Diffusion Model
Changyong He, Jin Zeng, Jiawei Zhang +1
Time-of-Flight (ToF) sensors efficiently capture scene depth, but the nonlinear depth construction procedure often results in extremely large noise variance or even invalid areas.…