2 papers
cs.CV2026
Tokenizing Semantic Segmentation with Run Length Encoding
Abhineet Singh, Justin Rozeboom, Nilanjan Ray
This paper presents a new unified approach to semantic segmentation in both images and videos by using language modeling to output the masks as sequences of discrete tokens. We use…
cs.CV2025
Improving Token-based Object Detection with Video
Abhineet Singh, Nilanjan Ray
This paper improves upon the Pix2Seq object detector by extending it for videos. In the process, it introduces a new way to perform end-to-end video object detection that improves…