5 papers
Codebook Capacity Governs Perceptual Quality Across Resolutions in Hierarchical Discrete Video Compression
Manikanta Kotthapalli, Banafsheh Rekabdar
Learned video codecs based on continuous latent representations typically require resolution-specific retraining or rate-distortion (RD) recalibration when scaling to new spatial r…
Entropy-Coded MS-VQ-VAE with Learned Priors for Ultra-Low Bitrate Video Compression
Manikanta Kotthapalli, Banafsheh Rekabdar
Learned video codecs based on continuous latent representations struggle to operate reliably below 0.1 bits per pixel~(bpp): without a differentiable rate signal, Lagrangian optimi…
Hierarchical Vector-Quantized Latents for Perceptual Low-Resolution Video Compression
Manikanta Kotthapalli, Banafsheh Rekabdar
The exponential growth of video traffic has placed increasing demands on bandwidth and storage infrastructure, particularly for content delivery networks (CDNs) and edge devices. W…
YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
Manikanta Kotthapalli, Deepika Ravipati, Reshma Bhatia
Over the past decade, object detection has advanced significantly, with the YOLO (You Only Look Once) family of models transforming the landscape of real-time vision applications t…
Self-Supervised YOLO: Leveraging Contrastive Learning for Label-Efficient Object Detection
Manikanta Kotthapalli, Reshma Bhatia, Nainsi Jain
One-stage object detectors such as the YOLO family achieve state-of-the-art performance in real-time vision applications but remain heavily reliant on large-scale labeled datasets…