activity
20242026
most cited"Ghost of the past": identifying and resolving privacy leakage from LLM's memory through proactive user interaction

4 citations · 9 across the 15 of their papers we have counts for

collaborators
Showing 2025 · cs.AIShow all

7 papers · 2 filters

cs.AI2025

DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas

Zhen Wang, Yufan Zhou, Zhongyan Luo +5

Simulating human profiles by instilling personas into large language models (LLMs) is rapidly transforming research in agentic behavioral simulation, LLM personalization, and human…

cs.AI2025

Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training

Dayuan Fu, Yunze Wu, Xiaojie Cai +13

Large Language Model (LLM) agents have recently shown strong potential in domains such as automated coding, deep research, and graphical user interface manipulation. However, train…

cs.AI2025

InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research

Yunze Wu, Dayuan Fu, Weiye Si +13

AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills…

cs.AI2025★ 3 cited

Context Engineering 2.0: The Context of Context Engineering

Qishuo Hua, Lyumanshan Ye, Dayuan Fu +6

Karl Marx once wrote that ``the human essence is the ensemble of social relations'', suggesting that individuals are not isolated entities but are fundamentally shaped by their int…

cs.AI2025★ 1 cited

AlphaGo Moment for Model Architecture Discovery

Yixiu Liu, Yang Nan, Weixian Xu +4

While AI systems demonstrate exponentially improving capabilities, the pace of AI research itself remains linearly bounded by human cognitive capacity, creating an increasingly sev…

cs.AI2025

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

Tianze Xu, Pengrui Lu, Lyumanshan Ye +2

The emergence of deep research systems presents significant capabilities in problem-solving, extending from basic queries to sophisticated research tasks. However, existing benchma…