collaborators

6 papers

cs.CV2026

Towards Accurate and Robust Surveillance Roadside IVD via Trackletized Audio-Visual Reasoning

Xiwen Li, Xiaoya Tang, Bodong Zhang +1

Idling Vehicle Detection (IVD) seeks to determine, at the final frame of a video clip, whether any vehicle is idling, meaning the vehicle is stationary with its engine running, usi…

cs.CV2026

Weakly Supervised Contrastive Learning for Histopathology Patch Embeddings

Bodong Zhang, Xiwen Li, Hamid Manoochehri +4

Digital histopathology whole slide images (WSIs) provide gigapixel-scale high-resolution images that are highly useful for disease diagnosis. However, digital histopathology image…

cs.CV2026

How to Build Robust, Scalable Models for GSV-Based Indicators in Neighborhood Research

Xiaoya Tang, Xiaohe Yue, Heran Mane +3

A substantial body of health research demonstrates a strong link between neighborhood environments and health outcomes. Recently, there has been increasing interest in leveraging a…

cs.CV2025

HAVT-IVD: Heterogeneity-Aware Cross-Modal Network for Audio-Visual Surveillance: Idling Vehicles Detection With Multichannel Audio and Multiscale Visual Cues

Xiwen Li, Xiaoya Tang, Tolga Tasdizen

Idling vehicle detection (IVD) uses surveillance video and multichannel audio to localize and classify vehicles in the last frame as moving, idling, or engine-off in pick-up zones.…

cs.CV2025

DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer

Xiaoya Tang, Bodong Zhang, Man Minh Ho +2

Despite the widespread adoption of transformers in medical applications, the exploration of multi-scale learning through transformers remains limited, while hierarchical representa…

cs.LG2025

Hierarchical Transformer for Electrocardiogram Diagnosis

Xiaoya Tang, Jake Berquist, Benjamin A. Steinberg +1

We propose a hierarchical Transformer for ECG analysis that combines depth-wise convolutions, multi-scale feature aggregation via a CLS token, and an attention-gated module to lear…