10 papers · 1 filter
Combating Label Noise With A General Surrogate Model For Sample Selection
Chao Liang, Linchao Zhu, Humphrey Shi +1
Modern deep learning systems are data-hungry. Learning with web data is one of the feasible solutions, but will introduce label noise inevitably, which can hinder the performance o…
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
Yu Lu, Yuanzhi Liang, Linchao Zhu +1
Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computa…
VillagerAgent: A Graph-Based Multi-Agent Framework for Coordinating Complex Task Dependencies in Minecraft
Yubo Dong, Xukun Zhu, Zhengzhe Pan +2
In this paper, we aim to evaluate multi-agent systems against complex dependencies, including spatial, causal, and temporal constraints. First, we construct a new benchmark, named…
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models
Xihang Yue, Linchao Zhu, Yi Yang
To process contexts with unlimited length using Large Language Models (LLMs), recent studies explore hierarchically managing the long text. Only several text fragments are taken fr…
CapHuman: Capture Your Moments in Parallel Universes
Chao Liang, Fan Ma, Linchao Zhu +2
We concentrate on a novel human-centric image synthesis task, that is, given only one reference facial photograph, it is expected to generate specific individual images with divers…
AudioScenic: Audio-Driven Video Scene Editing
Kaixin Shen, Ruijie Quan, Linchao Zhu +2
Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current…