collaborators

6 papers

cs.CV2026

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Yifan Shen, Jian Xu, Boyi Li +6

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stem…

cs.AI2026

Brick-Composer: Using MLLMs for Assembly with Diverse Bricks

Jiateng Liu, Bingxuan Li, Zhenhailong Wang +8

We dream of AI agents that can read arbitrary designs and construct real-world objects from reusable building blocks. As a first step toward this vision, we study whether multimoda…

cs.HC2026

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents

Jiateng Liu, Rushi Wang, Bingxuan Li +6

Recent advances have enabled general computer-use agents that interpret screens and execute grounded actions from human instructions, yet they still struggle to generalize to unsee…

cs.AI2026

OSExpert: Computer-Use Agents Learning Professional Skills via Exploration

Jiateng Liu, Zhenhailong Wang, Rushi Wang +6

General-purpose computer-use agents have shown impressive performance across diverse digital environments. However, our new benchmark, OSExpert-Eval, indicates they remain far less…

cs.CL2025

Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts

Rushi Wang, Jiateng Liu, Cheng Qian +6

Incorporating external context can significantly enhance the response quality of Large Language Models (LLMs). However, real-world contexts often mix relevant information with disp…

cs.IR2025

Automating Financial Statement Audits with Large Language Models

Rushi Wang, Jiateng Liu, Weijie Zhao +2

Financial statement auditing is essential for stakeholders to understand a company's financial health, yet current manual processes are inefficient and error-prone. Even with exten…