8 papers
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
Ruike Hu, Shulei Wu
The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Struct, a lightweight framework using Gradie…
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
Lianghui Zhu, Bin Ouyang, Yuxuan Zhang +8
Text-prompted image segmentation enables fine-grained visual understanding and is critical for applications such as human-computer interaction and robotics. However, existing super…
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
Rui Hu, Lianghui Zhu, Yuxuan Zhang +7
Pixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between v…
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
Rui Hu, Xiaolong Lin, Jiawang Liu +2
In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our a…
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
Han Xiao, Guozhi Wang, Yuxiang Chai +12
In this paper, we introduce UI-Genie, a self-improving framework addressing two key challenges in GUI agents: verification of trajectory outcome is challenging and high-quality tra…
Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
Han Xiao, Yina Xie, Guanxin Tan +12
Visual Document Understanding has become essential with the increase of text-rich visual content. This field poses significant challenges due to the need for effective integration…