Publications (5)
InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
Qi Zhang, Xinquan Yu, Kaiyi Zhang +1
Camera-only 4D occupancy forecasting enables autonomous vehicles to predict future 3D semantic scenes solely from historical multi-view images, which is critical for driving safety…
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
Xinquan Yu, Wei Lu, Xiangyang Luo
The task of Detecting and Grounding Multi-Modal Media Manipulation (DGM) is a branch of misinformation detection. Unlike traditional binary classification, it includes complex…
Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
Qi Zhang, Jixuan Chen, Kaiyi Zhang +3
Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking archit…
CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization
Xinquan Yu, Wei Lu, Xiangyang Luo +1
To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-…
RaCMC: Residual-Aware Compensation Network with Multi-Granularity Constraints for Fake News Detection
Xinquan Yu, Ziqi Sheng, Wei Lu +2
Multimodal fake news detection aims to automatically identify real or fake news, thereby mitigating the adverse effects caused by such misinformation. Although prevailing approache…