activity
20172026
most citedFine-grained Video-Text Retrieval with Hierarchical Graph Reasoning

26 citations · 131 across the 37 of their papers we have counts for

collaborators
Showing cs.CVShow all

35 papers · 1 filter

cs.CV2026

Flow Matching in Feature Space for Stochastic World Modeling

Francois Porcher, Nicolas Carion, Karteek Alahari +1

World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world models often struggle to satisfy both goa…

cs.CV2026

MAGICIAN: Efficient Long-Term Planning with Imagined Gaussians for Active Mapping

Shiyao Li, Antoine Guédon, Shizhe Chen +1

Active mapping aims to determine how an agent should move to efficiently reconstruct unknown environments. Most existing approaches rely on greedy next-best-view prediction, result…

cs.CV2026

HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +3

Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plau…

cs.CV2025

ComposeAnything: Composite Object Priors for Text-to-Image Generation

Zeeshan Khan, Shizhe Chen, Cordelia Schmid

Generating images from text involving complex and novel object arrangements remains a significant challenge for current text-to-image (T2I) models. Although prior layout-based meth…

cs.CV2025

HORT: Monocular Hand-held Objects Reconstruction with Transformers

Zerui Chen, Rolandos Alexandros Potamias, Shizhe Chen +1

Reconstructing hand-held objects in 3D from monocular images remains a significant challenge in computer vision. Most existing approaches rely on implicit 3D representations, which…

cs.CV2025

NextBestPath: Efficient 3D Mapping of Unseen Environments

Shiyao Li, Antoine Guédon, Clémentin Boittiaux +2

This work addresses the problem of active 3D mapping, where an agent must find an efficient trajectory to exhaustively reconstruct a new scene. Previous approaches mainly predict t…