1 citations · 1 across the 3 of their papers we have counts for
5 papers
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
Susan Liang, Chao Huang, Filippos Bellos +7
State-of-the-art text-to-video generation models such as Sora 2 and Veo 3 can now produce high-fidelity videos with synchronized audio directly from a textual prompt, marking a new…
Towards Consistent Long-Term Pose Generation
Yayuan Li, Filippos Bellos, Jason Corso
Current approaches to pose generation rely heavily on intermediate representations, either through two-stage pipelines with quantization or autoregressive models that accumulate er…
Towards Effective Human-in-the-Loop Assistive AI Agents
Filippos Bellos, Yayuan Li, Cary Shu +3
Effective human-AI collaboration for physical task completion has significant potential in both everyday activities and professional domains. AI agents equipped with informative gu…
Transparent and Coherent Procedural Mistake Detection
Shane Storks, Itamar Bar-Yossef, Yayuan Li +3
Procedural mistake detection (PMD) is a challenging problem of classifying whether a human user (observed through egocentric video) has successfully executed a task (specified by a…
HANDI: Hand-Centric Text-and-Image Conditioned Video Generation
Yayuan Li, Zhi Cao, Jason J. Corso
Despite the recent strides in video generation, state-of-the-art methods still struggle with elements of visual detail. One particularly challenging case is the class of videos in…