activity
20242026
most citedLLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation

2 citations · 3 across the 19 of their papers we have counts for

collaborators
Showing 2026Show all

7 papers · 1 filter

cs.CL2026

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Hongchen Wei, Yuanzhe Wang, Bei Liu +8

Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually selec…

cs.CL2026

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding

Hongchen Wei, Yuanzhe Wang, Bei Liu +9

Real-world document tasks often ask professionals to answer questions from annual reports, regulations, clinical guidelines, and technical manuals that span hundreds or thousands o…

cs.SE2026

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

Yijia Fan, Zonglin Di, Zimo Wen +8

Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, te…

cs.CV2026

A Comprehensive Ecosystem for Open-Domain Customized Video Generation

Jingxu Zhang, Yuqian Hong, Daneul Kim +6

Recent progress in video generation has shown impressive visual synthesis capabilities. However, open-domain customized video generation remains limited by the lack of large-scale,…

cs.CV2026

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Ziwei Zhou, Zeyuan Lai, Rui Wang +6

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and v…

cs.CV2026

High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding

Ji Woo Hong, Hee Suk Yoon, Gwanhyeong Koo +5

Recent large-scale vision-language models (VLMs) have shown remarkable text-to-image generation capabilities, yet their visual fidelity remains constrained by the discrete image to…