2 citations · 3 across the 19 of their papers we have counts for
7 papers · 1 filter
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Hongchen Wei, Yuanzhe Wang, Bei Liu +8
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually selec…
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Hongchen Wei, Yuanzhe Wang, Bei Liu +9
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands o…
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Yijia Fan, Zonglin Di, Zimo Wen +8
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, te…
A Comprehensive Ecosystem for Open-Domain Customized Video Generation
Jingxu Zhang, Yuqian Hong, Daneul Kim +6
Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale,…
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
Ziwei Zhou, Zeyuan Lai, Rui Wang +6
Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and v…
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo +5
Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image to…