2 papers
cs.CL2026
Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models
Zongji Yu, Wenshui Luo, Yiliu Sun +4
Post-training has significantly enhanced the reasoning capability of Large Reasoning Models (LRMs), especially with Reinforcement Learning (RL) like Group Relative Policy Optimizat…
cs.CV2025
UIS-Mamba: Exploring Mamba for Underwater Instance Segmentation via Dynamic Tree Scan and Hidden State Weaken
Runmin Cong, Zongji Yu, Hao Fang +2
Underwater Instance Segmentation (UIS) tasks are crucial for underwater complex scene detection. Mamba, as an emerging state space model with inherently linear complexity and globa…