collaborators

14 papers

cs.CR2026

CallScreenBench: Benchmarking On-Device Models as Phone Secretaries

Simiao Ren

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting for their user, making on-device task automation newly plausible. One…

cs.CV2026

GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment

Kidus Zewde, Simiao Ren, Xingyu Shen +6

The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content has never been more difficult…

cs.CV2026

When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents

Jiaqi Wu, Yuchen Zhou, Dennis Tsang Ng +5

OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be replaced in under a second for…

cs.CV2026

A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator

Kidus Zewde, Yuchen Zhou, Dennis Ng +6

Large vision-language models have achieved remarkable capabilities by training on massive internet-scale data, yet a fundamental asymmetry persists: while LLMs can leverage self-su…

cs.CL2026

Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate

Simiao Ren, Xingyu Shen, Yuchen Zhou +3

A claim has been circulating on social media and practitioner forums that Chinese prompts are more token-efficient than English for LLM coding tasks, potentially reducing costs by…

cs.AI2026

GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics

Yan Zhang, Simiao Ren, Ankit Raj +6

Can humans detect AI-generated financial documents better than machines? We present GPT4o-Receipt, a benchmark of 1,235 receipt images pairing GPT-4o-generated receipts with authen…