2 citations · 3 across the 5 of their papers we have counts for
4 papers · 1 filter
VideoPCDNet: Video Parsing and Prediction with Phase Correlation Networks
Noel José Rodrigues Vicente, Enrique Lehner, Angel Villar-Corrales +2
Understanding and predicting video content is essential for planning and reasoning in dynamic environments. Despite advancements, unsupervised learning of object representations an…
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
Bastian Pätzold, Jan Nogga, Sven Behnke
Vision-language models (VLMs) excel in visual understanding but often lack reliable grounding capabilities and actionable inference rates. Integrating them with open-vocabulary obj…
Semantic Prediction: Which One Should Come First, Recognition or Prediction?
Hafez Farazi, Jan Nogga, and Sven Behnke
The ultimate goal of video prediction is not forecasting future pixel-values given some previous frames. Rather, the end goal of video prediction is to discover valuable internal r…
Local Frequency Domain Transformer Networks for Video Prediction
Hafez Farazi, Jan Nogga, Sven Behnke
Video prediction is commonly referred to as forecasting future frames of a video sequence provided several past frames thereof. It remains a challenging domain as visual scenes evo…