From the 2 of 10 linked papers with an AI index.
10 papers
ENCORE: Event-Assisted Complementary Motion Refinement for Learned Video Compression
Shuhan Ye, Hongbin Yu, Chenqi Kong +4
The paper introduces ENCORE, a framework that uses asynchronous event‑camera data to refine motion estimation in learned video compression, improving quality especially under chall…
Contrastive-Augmented Flow Matching for Style-Content Disentanglement
Yusong Li, Pingchuan Ma, Ming Gui +2
The paper proposes Contrastive Augmented Flow Matching (CAtFM), a method that adds contrastive regularization to invertible flow matching to learn disentangled content and style re…
Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
Johannes Schusterbauer, Ming Gui, Yusong Li +3
Diffusion- and flow-based models usually allocate compute uniformly across space, updating all patches with the same timestep and number of function evaluations. While convenient,…
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
Thai-Binh Nguyen, Katerina Zmolikova, Pingchuan Ma +3
We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a s…
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
Pingchuan Ma, Xiaopei Yang, Yusong Li +4
Explicitly disentangling style and content in vision models remains challenging due to their semantic overlap and the subjectivity of human perception. Existing methods propose sep…
Does VLM Classification Benefit from LLM Description Semantics?
Pingchuan Ma, Lennart Rietdorf, Dmytro Kotovenko +2
Accurately describing images with text is a foundation of explainable AI. Vision-Language Models (VLMs) like CLIP have recently addressed this by aligning images and texts in a sha…