2 papers
cs.LG2026
WARC-Bench: Web Archive Based Benchmark for GUI Subtask Executions
Sanjari Srivastava, Gang Li, Cheng Chang +8
Training web agents to navigate complex, real-world websites requires them to master - short-horizon interactions on multiple UI components (e.g., choosing the…
cs.CL2026
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Mengzhao Jia, Zhihan Zhang, Ignacio Cases +3
Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning with verifiable rewards (RLVR) ofte…