3 papers
cs.AI2026
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
Junjie Zhao, Jingyi Liang, Zhenyang Cai +22
While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios remain critically underexplore…
cs.DC2026
OSGym: Scalable OS Infra for Computer Use Agents
Zengyi Qin, Jinyuan Chen, Yunze Man +25
Training computer use agents requires full-featured OS sandboxes with GUI environments, which consume substantial hardware resources as the number of sandboxes scales. Stochastic e…
cs.DL2025
Web Archives Metadata Generation with GPT-4o: Challenges and Insights
Ashwin Nair, Zhen Rong Goh, Tianrui Liu +1
Current metadata creation for web archives is time consuming and costly due to reliance on human effort. This paper explores the use of gpt-4o for metadata generation within the We…