3 papers
cs.CL2026
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
Adam Byerly, Daniel Khashabi
Large language models (LLMs) exhibit pronounced position bias in long-context needle-in-haystack problems, systematically prioritizing the location of information over its relevanc…
cs.CL2025
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
Adam Byerly, Daniel Khashabi
Self-consistency (SC) improves the performance of large language models (LLMs) across various tasks and domains that involve short content. However, does this support its effective…
cs.AI2025
Tur[k]ingBench: A Challenge Benchmark for Web Agents
Kevin Xu, Yeganeh Kordi, Tanay Nayak +7
Can advanced multi-modal models effectively tackle complex web-based tasks? Such tasks are often found on crowdsourcing platforms, where crowdworkers engage in challenging micro-ta…