4 papers
Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence
Sankalp Nagaonkar, Rohit Garg, Ankit Raj +2
Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems p…
When the Forger Is the Judge: GPT-Image-2 Cannot Recognize Its Own Faked Documents
Jiaqi Wu, Yuchen Zhou, Dennis Tsang Ng +5
OpenAI's GPT-Image-2 has effectively erased the visual boundary between authentic and AI-edited document images: a single number on a receipt can be replaced in under a second for…
A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator
Kidus Zewde, Yuchen Zhou, Dennis Ng +6
Large vision-language models have achieved remarkable capabilities by training on massive internet-scale data, yet a fundamental asymmetry persists: while LLMs can leverage self-su…
GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document Forensics
Yan Zhang, Simiao Ren, Ankit Raj +6
Can humans detect AI-generated financial documents better than machines? We present GPT4o-Receipt, a benchmark of 1,235 receipt images pairing GPT-4o-generated receipts with authen…