collaborators

6 papers

cs.LG2026

DeepLoop: Depth Scaling for Looped Transformers

Shuzhen Li, Yifan Zhang, Jiacheng Guo +2

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters.…

cs.AI2026

Interactive Benchmarks

Baoqing Yue, Zihan Zhu, Yutong Han +6

Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evalu…

cs.LG2026

Deep Delta Learning

Yifan Zhang, Yifeng Liu, Mengdi Wang +1

Transformer residual streams evolve through additive updates. Although a sufficiently expressive residual block can represent content replacement, standard architectures do not par…

cs.AI2025

Web World Models

Jichen Feng, Yifan Zhang, Chenggong Zhang +3

Language agents increasingly require persistent worlds in which they can act, remember, and learn. Existing approaches sit at two extremes: conventional web frameworks provide reli…

cs.AI2025

Monadic Context Engineering

Yifan Zhang, Yang Yuan, Mengdi Wang +1

The proliferation of Large Language Models (LLMs) has catalyzed a shift towards autonomous agents capable of complex reasoning and tool use. However, current agent architectures ar…

cs.CL2025

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

Jiacheng Guo, Suozhi Huang, Zixin Yao +16

This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in t…