4 citations · 4 across the 7 of their papers we have counts for
9 papers · 1 filter
MobileSAM2: Lightweight Segment Anything for Spatial Intelligence
Kai Jiang, Jiaxing Huang, Jingyi Zhang +5
The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such…
PRPO: Perception-Reinforced Policy Optimization via Token-Level Dynamic Advantage Reshaping
Qiming Li, Tianlun Li, Xiaolong Cheng +5
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective paradigm for improving the reasoning capability of Large Vision-Language Models (LVLMs). However, exis…
ABot-OCR Technical Report
Kaitao Jiang, Ruiyan Gong, Xiaolong Cheng +3
We introduce ABot-OCR, an end-to-end vision-language model that transcribes a page image directly into clean Markdown in a single forward pass. By doing so, our approach completely…
A Survey on Vision Autoregressive Model
Kai Jiang, Jiaxing Huang
Autoregressive models have demonstrated great performance in natural language processing (NLP) with impressive scalability, adaptability and generalizability. Inspired by their not…
Open-Vocabulary Object Detection via Language Hierarchy
Jiaxing Huang, Jingyi Zhang, Kai Jiang +1
Recent studies on generalizable object detection have attracted increasing attention with additional weak supervision from large-scale datasets with image-level labels. However, we…
Learning to Prompt Segment Anything Models
Jiaxing Huang, Kai Jiang, Jingyi Zhang +4
Segment Anything Models (SAMs) like SEEM and SAM have demonstrated great potential in learning to segment anything. The core design of SAMs lies with Promptable Segmentation, which…