2 papers
cs.AI2026
BrowseComp-: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents
Huanyao Zhang, Jiepeng Zhou, Bo Li +22
Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimod…
cs.CL2026
Long-Context Long-Form Question Answering for Legal Domain
Anagha Kulkarni, Parin Rajesh Jhaveri, Prasha Shrestha +3
Legal documents have complex document layouts involving multiple nested sections, lengthy footnotes and further use specialized linguistic devices like intricate syntax and domain-…