activity
20242026
most citedDAMe: Personalized Federated Social Event Detection with Dual Aggregation Mechanism

1 citations · 2 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

8 papers · 1 filter

cs.CL2026

Towards Compositional Generalization of LLMs via Skill Taxonomy Guided Data Synthesis

Yifan Wei, Li Du, Xiaoyan Yu +2

Large Language Models (LLMs) and agent-based systems often struggle with compositional generalization due to a data bottleneck in which complex skill combinations follow a long-tai…

cs.CL20251 cited

AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning

Yifan Wei, Xiaoyan Yu, Yixuan Weng +3

Large Language Models (LLMs), when enhanced through reasoning-oriented post-training, evolve into powerful Large Reasoning Models (LRMs). Tool-Integrated Reasoning (TIR) further ex…

cs.CL2025

Structural Entropy Guided Agent for Detecting and Repairing Knowledge Deficiencies in LLMs

Yifan Wei, Xiaoyan Yu, Tengfei Pan +2

Large language models (LLMs) have achieved unprecedented performance by leveraging vast pretraining corpora, yet their performance remains suboptimal in knowledge-intensive domains…

cs.CL2025

SetKE: Knowledge Editing for Knowledge Elements Overlap

Yifan Wei, Xiaoyan Yu, Ran Song +2

Large Language Models (LLMs) excel in tasks such as retrieval and question answering but require updates to incorporate new knowledge and reduce inaccuracies and hallucinations. Tr…

cs.CL2024

Towards Effective, Efficient and Unsupervised Social Event Detection in the Hyperbolic Space

Xiaoyan Yu, Yifan Wei, Shuaishuai Zhou +5

The vast, complex, and dynamic nature of social message data has posed challenges to social event detection (SED). Despite considerable effort, these challenges persist, often resu…

cs.CL2024

DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models

Yiming Huang, Jianwen Luo, Yan Yu +8

We introduce DA-Code, a code generation benchmark specifically designed to assess LLMs on agent-based data science tasks. This benchmark features three core elements: First, the ta…