collaborators

6 papers

cs.CV2025

Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models

Yolo Y. Tang, Jing Bi, Pinxin Liu +24

Video understanding represents the most challenging frontier in computer vision, requiring models to reason about complex spatiotemporal relationships, long-term dependencies, and…

cs.HC2025

UI-Evol: Automatic Knowledge Evolving for Computer Use Agents

Ziyun Zhang, Xinyi Liu, Xiaoyi Zhang +3

External knowledge has played a crucial role in the recent development of computer use agents. We identify a critical knowledge-execution gap: retrieved knowledge often fails to tr…

cs.HC2025

UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis

Xinyi Liu, Xiaoyi Zhang, Ziyun Zhang +1

Recent advancements in Large Vision-Language Models are accelerating the development of Graphical User Interface (GUI) agents that utilize human-like vision perception capabilities…

cs.CL2025

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

Shuhao Guan, Moule Lin, Cheng Xu +5

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual con…

cs.CY2025

Crafting Synthetic Realities: Examining Visual Realism and Misinformation Potential of Photorealistic AI-Generated Images

Qiyao Peng, Yingdan Lu, Yilang Peng +3

Advances in generative models have created Artificial Intelligence-Generated Images (AIGIs) nearly indistinguishable from real photographs. Leveraging a large corpus of 30,824 AIGI…

cs.RO2024

SplaTraj: Camera Trajectory Generation with Semantic Gaussian Splatting

Xinyi Liu, Tianyi Zhang, Matthew Johnson-Roberson +1

Many recent developments for robots to represent environments have focused on photorealistic reconstructions. This paper particularly focuses on generating sequences of images from…