5 papers · 1 filter
Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
TaoLive AIGC LLM Team, Yuhan Sun, Wenhao Lin +7
AI-powered digital avatar streamers must answer product questions, engage viewers, and execute marketing strategies in real time, demanding low latency, frequent strategy updates,…
Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy
Denghao Ma, Qing Liu, Zulong Chen +5
Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings wit…
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
Xiang Feng, Jiawei Zhou, Zhangfeng Huang +6
Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer ac…
CC-OCR V2: Fine-Grained Attribution of LMM Failures in Real-World Visual Document Understanding
Chunyi Peng, Zhipeng Xu, Yuqi Xiong +12
Recent Large Multimodal Models (LMMs) have achieved remarkable progress on OCR-centric document understanding and processing tasks. Existing benchmarks primarily evaluate LMMs acro…
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
Zhenghao Liu, Zhuoyang Wu, Xinze Li +6
Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, particularly in solving complex mathematical problems. Recent studies show that distilling long re…