activity
20242026
collaborators

6 papers

cs.CV2026

When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning

Zhengxian Wu, Kai Shi, Chuanrui Zhang +10

Recent progress in multimodal large language models has led to strong performance on reasoning tasks, but these improvements largely rely on high-quality annotated data or teacher-…

cs.RO2025

MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation

Ning Li, Xiangmou Qu, Jiamu Zhou +6

Recent advances in Multimodal Large Language Models (MLLMs) have enabled the development of mobile agents that can understand visual inputs and follow user instructions, unlocking…

cs.CL2025

DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL

Haoyuan Ma, Yongliang Shen, Hengwei Liu +5

Recent text-to-SQL systems powered by large language models (LLMs) have demonstrated remarkable performance in translating natural language queries into SQL. However, these systems…

cs.CL2025

STaR-SQL: Self-Taught Reasoner for Text-to-SQL

Mingqian He, Yongliang Shen, Wenqi Zhang +3

Generating step-by-step "chain-of-thought" rationales has proven effective for improving the performance of large language models on complex reasoning tasks. However, applying such…

cs.CL2025

HammerBench: Fine-Grained Function-Calling Evaluation in Real Mobile Device Scenarios

Jun Wang, Jiamu Zhou, Muning Wen +7

Evaluating the performance of LLMs in multi-turn human-agent interactions presents significant challenges, particularly due to the complexity and variability of user behavior. In t…

cs.CL2024

P3: A Policy-Driven, Pace-Adaptive, and Diversity-Promoted Framework for data pruning in LLM Training

Yingxuan Yang, Huayi Wang, Muning Wen +4

In the rapidly advancing field of Large Language Models (LLMs), effectively leveraging existing datasets during fine-tuning to maximize the model's potential is of paramount import…