37 citations · 69 across the 30 of their papers we have counts for
31 papers · 1 filter
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement
Siwei Wu, Jincheng Ren, Yizhi Li +11
Recent work extends recursive self-improvement (RSI) to agent harnesses for long-horizon coding and terminal tasks, enabling agents to improve execution mechanisms from experience.…
MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL
Beiyu Xu, Zhenyu Wu, Jiaoyan Chen +1
Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, heterogeneous database collecti…
The BD-LSC Dataset: Facilitating the Benchmarking of Models for Lexical Semantic Change Detection in Slang and Standard Usage
Afnan Aloraini, Viktor Schlegel, Goran Nenadic +1
Automatic semantic change detection aims to identify how word meanings shift over time, offering insights into both linguistic and societal change. Despite recent progress in compu…
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
Jincheng Ren, Siwei Wu, Yizhi Li +8
As terminal agents scale to long-horizon, multi-turn workflows, a key bottleneck is not merely limited context length, but the accumulation of noisy terminal observations in the in…
Large-Scale Terminal Agentic Trajectory Generation from Dockerized Environments
Siwei Wu, Yizhi Li, Yuyang Song +8
Training agentic models for terminal-based tasks critically depends on high-quality terminal trajectories that capture realistic long-horizon interactions across diverse domains. H…
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
Yulong Wu, Viktor Schlegel, Riza Batista-Navarro
How does the natural evolution of context paragraphs affect question answering in generative Large Language Models (LLMs)? To investigate this, we propose a framework for curating…