3 papers
cs.AI2026
UI-Venus-2 Technical Report
Venus Team, Zhuohan Cai, Haoxing Chen +28
Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remai…
cs.SE2025
SIADAFIX: issue description response for adaptive program repair
Xin Cao, Nan Yu
We propose utilizing fast and slow thinking to enhance the capabilities of large language model-based agents on complex tasks such as program repair. In particular, we design an ad…
cs.CL2025
OpenHuEval: Evaluating Large Language Model on Hungarian Specifics
Haote Yang, Xingjian Wei, Jiang Wu +18
We introduce OpenHuEval, the first benchmark for LLMs focusing on the Hungarian language and specifics. OpenHuEval is constructed from a vast collection of Hungarian-specific mater…