Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Can You Trust the Confidence? ConfBench for Vision-Language Models on Document Extraction
Priyashree Roy, Sujitha Martin, Mohammad Rostami +6
Intelligent document processing (IDP) with vision-language models (VLMs) hinges on confidence scores trustworthy enough to route extractions between automation and human review. Ex…
cs.AI2026
Brick-DICL: Dynamic In-Context Learning for Automated Brick Schema Classification
Yiyue Qian, Shinan Zhang, Huan Song +3
Building Management Systems (BMS) are essential for optimizing energy efficiency and operational performance in modern buildings. However, the lack of standardization across BMS po…
cs.AI2026
CollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration
Yiyue Qian, Shinan Zhang, Yun Zhou +3
Large Language Models (LLMs) have revolutionized AI-generated content evaluation, with the LLM-as-a-Judge paradigm becoming increasingly popular. However, current single-LLM evalua…