30 papers
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
Kai Yang, Jingwei Xu, Wanyu Wang +4
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradat…
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
Zichuan Fu, Shirong Wang, Wenlin Zhang +10
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated inter…
BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking
Bowen Yu, Sheng Zhang, Binhao Wang +8
Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabilities into smaller models remains challengin…
T-GINEE: A Tensor-Based Multilayer Graph Representation Learning
Maolin Wang, Ziting Mai, Xuhui Chen +9
Traditional network analysis focuses on single-layer networks, real-world systems often form multilayer networks with multiple relationship types. However, existing methods typical…
BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations
Mengyang Ma, Xiaopeng Li, Wanyu Wang +9
Transformer structures have been widely used in sequential recommender systems (SRS). However, as user interaction histories increase, computational time and memory requirements al…
SkyLink: A Large Vision-Language Model Driven Re-ranking Framework for Cross-View UAV geolocalization
Bowen Liu, Pengyue Jia, Wanyu Wang +9
Cross-view UAV geolocalization is fundamentally a challenging large-scale image retrieval task, aiming to determine the geographic coordinates of Unmanned Aerial Vehicle (UAV) quer…