From the 2 of 5 linked papers with an AI index.
5 papers
VersaQ-3D: Architecture Support for Visual Geometry Grounded Transformers via Versatile Quantization
Yipu Zhang, Jintao Cheng, Xingyu Liu +8
The paper introduces VersaQ-3D, a co-designed quantization algorithm and reconfigurable accelerator that enables low‑bit (4‑bit) inference of Visual Geometry Grounded Transformers…
The SonicAGI System for the REAL-TSE Challenge
Kai Li, Wendi Sang, Jintao Cheng +1
The paper presents SonicAGI, a system for real-world target speaker extraction that combines simulated and real meeting data, using a low‑latency SwiftNet-Lookahead model for onlin…
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
Kai Li, Jintao Cheng, Chang Zeng +5
Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods c…
Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer
Yipu Zhang, Jintao Cheng, Weilun Feng +5
Feed-forward 3D reconstruction models, represented by Visual Geometry Grounded Transformer (VGGT), jointly predict multiple visual geometry tasks such as depth estimation, camera p…
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
Jintao Cheng, Weibin Li, Zhijian He +3
Visual Place Recognition (VPR) demands representations robust to drastic environmental and viewpoint shifts. Existing aggregation paradigms either depend on extensive supervised tr…