2 papers
cs.CL2025
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
Yeganeh Kordi, Nihal V. Nayak, Max Zuo +2
We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation. Existing research is…
cs.AI2025
Tur[k]ingBench: A Challenge Benchmark for Web Agents
Kevin Xu, Yeganeh Kordi, Tanay Nayak +7
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-ta…