6 papers
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
Donghuo Zeng, Hao Niu, Masato Taya
Learning aligned multimodal embeddings from weakly paired, label-free corpora is challenging: pipelines often provide only pre-extracted features, clips contain multiple events, an…
Variance & Greediness: A comparative study of metric-learning losses
Donghuo Zeng, Hao Niu, Zhi Li +1
Metric learning is central to retrieval, yet its effects on embedding geometry and optimization dynamics are not well understood. We introduce a diagnostic framework, VARIANCE (int…
Learning Audio-Visual Embeddings with Inferred Latent Interaction Graphs
Donghuo Zeng, Hao Niu, Yanan Wang +1
Learning robust audio-visual embeddings requires bringing genuinely related audio and visual signals together while filtering out incidental co-occurrences - background noise, unre…
Comparing Contrastive and Triplet Loss: Variance Analysis and Optimization Behavior
Donghuo Zeng
Contrastive loss and triplet loss are widely used objectives in deep metric learning, yet their effects on representation quality remain insufficiently understood. We present a the…
Top-down Activity Representation Learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrat…
Multi-object event graph representation learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Video question answering (VideoQA) is a task to predict the correct answer to questions posed about a given video. The system must comprehend spatial and temporal relationships amo…