9 papers
CoCo-IR: Contextual Composed Image Retrieval
Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding +6
Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual search…
A Two-Validator Web Interface for Structured Geometry Figure Annotation
Sabin-Codrut Badea, Adrian-Marius Dumitran
Annotating geometric figures from scanned documents has long been addressed by adapting generic annotation tools, tools not originally designed for such tasks, to use cases where t…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
OSGym: Scalable OS Infra for Computer Use Agents
Zengyi Qin, Jinyuan Chen, Yunze Man +25
Training computer use agents requires full-featured OS sandboxes with GUI environments, which consume substantial hardware resources as the number of sandboxes scales. Stochastic e…
DCNNAnaCal: Physics-Informed Machine Learning for Accurate and Precise Weak Lensing Shear Estimation
Shurui Lin, Xiangchong Li, Ji Li +3
Traditional weak gravitational lensing shear estimators are carefully calibrated but struggle to fully capture realistic galaxy morphologies, point-spread-function (PSF) effects, b…
Refer to Any Segmentation Mask Group With Vision-Language Prompts
Shengcao Cao, Zijun Wei, Jason Kuen +6
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for c…