activity
20242026
collaborators

7 papers

cs.AI2026

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Tingyang Chen, Shuo Lu, Kang Zhao +11

AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet to…

cs.CV2026

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

Haozhan Shen, Tiancheng Zhao, Kangjia Zhao +1

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schem…

cs.SE2026

EpiDroid: Dependency-Guided Recomposition for Deep State Discovery in Mobile GUI Testing

Jiahui Song, Jiaxin Zhi, Kangjia Zhao +6

The increasing scale and complexity of mobile applications make automated GUI exploration essential for software quality assurance. However, existing methods often neglect state de…

cs.CV2025

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao +4

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding. Recently, with the integration of test-time scaling techniques,…

cs.LG2025

HORAE: A Domain-Agnostic Language for Automated Service Regulation

Yutao Sun, Mingshuai Chen, Tiancheng Zhao +8

Artificial intelligence is rapidly encroaching on the field of service regulation. However, existing AI-based regulation techniques are often tailored to specific application domai…

cs.CV2025

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Haozhan Shen, Peng Liu, Jingcheng Li +9

Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective…