works on

From the 1 of 6 linked papers with an AI index.

activity
20242026
collaborators

6 papers

cs.LG2026

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Minh-Quan Le, Armand Comas, Alexandros Lattas +7

The paper proposes a self‑correcting framework called SC‑CMJP that couples image understanding and generation via cross‑modal attention in masked diffusion models, and introduces a…

cs.CV2026

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

Minh-Quan Le, Gaurav Mittal, Cheng Zhao +3

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training…

cs.CV2025

What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards

Minh-Quan Le, Yuanzhi Zhu, Vicky Kalogeiton +1

Recent video diffusion models can synthesize visually compelling clips, yet often violate basic physical laws-objects float, accelerations drift, and collisions behave inconsistent…

cs.RO2025

DUViN: Diffusion-Based Underwater Visual Navigation via Knowledge-Transferred Depth Features

Jinghe Yang, Minh-Quan Le, Mingming Gong +1

Autonomous underwater navigation remains a challenging problem due to limited sensing capabilities and the difficulty of constructing accurate maps in underwater environments. In t…

cs.CV2025

Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment

Minh-Quan Le, Gaurav Mittal, Tianjian Meng +5

While diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Que…

cs.CV2024

-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions

Minh-Quan Le, Alexandros Graikos, Srikar Yellapragada +3

Synthesizing high-resolution images from intricate, domain-specific information remains a significant challenge in generative modeling, particularly for applications in large-image…