6 citations · 7 across the 18 of their papers we have counts for
22 papers
Dr. Claw: An AI Scientist Workspace for Vibe Research
Dingjie Song, Hanrong Zhang, Dawei Liu +10
Command-line coding agents (e.g., Claude Code, Gemini CLI) can already read and write files and sustain long sessions, yet end-to-end research still fragments across chat tools, ID…
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Jiajun Shi, Siyuan Tao, Yuhao Wu +18
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them…
WebWorld: The Browser as a World Model for Self-Improving Web Code
Jiajun Wu, Jian Yang, Yaxin Du +7
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor pr…
VGI-Bench: Probing Visual Intelligence in Video Generation Models
Xuan He, Cong Wei, Yuhao Cheng +20
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: b…
MedClaw: Heuristic Agent Harness for Long-Horizon Surgical Video Reasoning
Yingying Fan, Penghui Du, Leyan Zhu +10
Understanding tens-of-minutes surgical videos requires long-horizon temporal reasoning, answering what happens before, after, or across stages of a procedure by grounding the quest…
Distributional Matching for Vector Quantization: A Unified Theoretical and Empirical Framework
Xianghong Fang, Litao Guo, Hengchao Chen +8
The effectiveness of modern visual representation learning and autoregressive models critically depends on vector quantization (VQ), which discretizes continuous feature representa…