4 papers
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
Pengcheng Fang, Yuxia Chen, Xiaohao Cai
Video temporal grounding (VTG) aims to localize the start and end timestamps of the event described by a given query within an untrimmed video. Despite the strong open-world video…
CHASM: Cross-frequency Harmonized Axis-Separable Mixing for Spectral Token Operators
Pengcheng Fang, Hongli Chen, Yuxia Chen +3
Spectral token mixers based on Fourier transforms provide an efficient way to model global interactions in visual feature maps. Existing designs often either apply filter-wise spec…
HiFi-Mamba: Dual-Stream W-Laplacian Enhanced Mamba for High-Fidelity MRI Reconstruction
Hongli Chen, Pengcheng Fang, Yuxia Chen +6
Reconstructing high-fidelity MR images from undersampled k-space data remains a challenging problem in MRI. While Mamba variants for vision tasks offer promising long-range modelin…
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
Pengcheng Fang, Yuxia Chen, Rui Guo
Understanding videos requires more than answering open ended questions, it demands the ability to pinpoint when events occur and how entities interact across time. While recent Vid…