6 papers
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Hongchen Wei, Yuanzhe Wang, Bei Liu +8
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually selec…
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding
Hongchen Wei, Yuanzhe Wang, Bei Liu +9
Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands o…
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Yijia Fan, Zonglin Di, Zimo Wen +8
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, te…
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
Yasheng Sun, Zezi Zeng, Yifan Yang +4
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as the…
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
Ziwei Zhou, Zeyuan Lai, Rui Wang +6
Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and v…
DynaVid: Learning to Generate Highly Dynamic Videos using Synthetic Motion Data
Wonjoon Jin, Jiyun Won, Janghyeok Han +4
Despite recent progress, video diffusion models still struggle to synthesize realistic videos involving highly dynamic motions or requiring fine-grained motion controllability. A c…