4 papers
Open-Jev Judgments on CallScreenBench: Calibrated One-Pass Scam Screening with a Small Language Model
Simiao Ren, Kidus Zewde, Xingyu Shen +6
Screening a phone call for fraud needs a trustworthy probability after every caller turn, in milliseconds. Jev-style typed decisions promise exactly that: declared options go in, o…
Agents That Edit Documents: Measuring Agentic PDF Forgery Against a Non-Agentic Control
Simiao Ren, Ankit Raj, Tommy Duong +6
AI agents that carry a multi-step computer task through on their own became ordinary tools in the past year, and the same autonomy is available to anyone whose task is harmful. We…
ChatGPT Images 2.5 in the Wild: A Launch-Period Dataset and Detector Evaluation
Dennis Ng, Xingyu Shen, Ankit Raj +6
An image tool can change its underlying generator while retaining its public name, making version attribution from online posts ambiguous. We study this problem after the ChatGPT I…
A Synthetic Eye Movement Dataset for Script Reading Detection: Real Trajectory Replay on a 3D Simulator
Kidus Zewde, Yuchen Zhou, Dennis Ng +6
Large vision-language models have achieved remarkable capabilities by training on massive internet-scale data, yet a fundamental asymmetry persists: while LLMs can leverage self-su…