12 papers
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
Zichuan Fu, Shirong Wang, Wenlin Zhang +10
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated inter…
BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking
Bowen Yu, Sheng Zhang, Binhao Wang +8
Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabilities into smaller models remains challengin…
LLM-as-Code: Agentic Programming for Agent Harness
Junjia Qi, Zichuan Fu, Jingtong Gao +4
Every major LLM agent framework gives the LLM the role of orchestrator; the model decides what to do next, when to call tools, and when to stop. We argue that token explosion, cont…
Reinforced Preference Optimization for Reasoning-Augmented Recommendations
Jingtong Gao, Zeyu Song, Chi Lu +7
Recommender systems are critical for delivering personalized content across digital platforms, and recent advances in Large Language Models (LLMs) offer new opportunities to enhanc…
Strategic Over-Parameterization for Generalizable Low-Rank Adaptation
Jing Gao, Zhong-Yi Lu, Pan Zhang +1
Adapting large language models (LLMs) to downstream tasks via full fine-tuning is increasingly impractical due to its computational and memory demands. Parameter-efficient fine-tun…
LegalDrill: Diagnosis-Driven Synthesis for Legal Reasoning in Small Language Models
Tianchun Li, Haochen Liu, Vishwa Pardeshi +5
Small language models (SLMs) are promising for real-world deployment due to their efficiency and low operational cost. However, their limited capacity struggles with high-stakes le…