paper

Rapid-Deployment Crack Measurement Based on SAM3 Semantic-Edge Response Decoding

arXiv:2607.12292

Abstract

Reliable crack measurement is essential for infrastructure condition assessment, yet existing image-based approaches typically depend on pixel-wise annotations, task-specific segmentation training, and mask-based geometric measurement, making cross-scene deployment costly and sensitive to segmentation errors. We identify an output-interface mismatch in SAM3: its prompt-conditioned semantic response preserves crack evidence that is often suppressed or spatially distorted in the final candidate masks. Across six public crack datasets, the internal response achieves 82.66% average crack-pixel recall, compared with 74.66% for the retained SAM3 proposals, with an average mismatch ratio of 8.73%. Based on this observation, we propose Semantic-Edge Response Decoding (SERD) to calibrate the semantic response using a fixed Sobel structural field, and further develop SERD-DQ, a training-free framework that directly estimates crack centerline and transverse geometry from the continuous decoded response without generating an intermediate predicted mask. Experiments verify both segmentation fidelity and direct geometric measurement against manually established pixel-level references. Compared with native SAM3 mask-based quantification, SERD-DQ reduces width MAE from 6.072 to 5.547 pixels, length relative error from 22.245% to 17.355%, and area relative error from 33.921% to 26.228%, while achieving a latent geometry recovery rate of 0.394. The results indicate that continuous semantic-edge responses provide a more reliable interface for training-free crack quantification than conventional mask-mediated measurement.

Submitted to Elsevier