6 papers
DiscussLLM: Teaching Large Language Models When to Speak
Deep Anil Patel, Iain Melvin, Christopher Malon +1
Large Language Models (LLMs) have demonstrated remarkable capabilities in understanding and generating human-like text, yet they largely operate as reactive agents, responding only…
CalibFree: Self-Supervised View Feature Separation for Calibration-Free Multi-Camera Multi-Object Tracking
Ruiqi Xian, Deep Patel, Iain Melvin +3
Multi-camera multi-object tracking (MCMOT) faces significant challenges in maintaining consistent object identities across varying camera perspectives, particularly when precise ca…
Uncertainty-Aware Knowledge Distillation for Multimodal Large Language Models
Jingchen Sun, Shaobo Han, Deep Patel +3
Knowledge distillation establishes a learning paradigm that leverages both data supervision and teacher guidance. However, determining the optimal balance between learning from dat…
MCTR: Multi Camera Tracking Transformer
Alexandru Niculescu-Mizil, Deep Patel, Iain Melvin
Multi-camera tracking plays a pivotal role in various real-world applications. While end-to-end methods have gained significant interest in single-camera tracking, multi-camera tra…
Object-Aware 4D Human Motion Generation
Shurui Gui, Deep Anil Patel, Xiner Li +1
Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations,…
Group Relative Augmentation for Data Efficient Action Detection
Deep Anil Patel, Iain Melvin, Zachary Izzo +1
Adapting large Video-Language Models (VLMs) for action detection using only a few examples poses challenges like overfitting and the granularity mismatch between scene-level pre-tr…