4 papers
Question-guided Visual Compression with Memory Feedback for Long-Term Video Understanding
Sosuke Yamao, Natsuki Miyahara, Yuankai Qi +1
In the context of long-term video understanding with large multimodal models, many frameworks have been proposed. Although transformer-based visual compressors and memory-augmented…
Enhancing Multi-Camera Gymnast Tracking Through Domain Knowledge Integration
Fan Yang, Shigeyuki Odashima, Shoichi Masui +3
We present a robust multi-camera gymnast tracking, which has been applied at international gymnastics championships for gymnastics judging. Despite considerable progress in multi-c…
YOWO: You Only Walk Once to Jointly Map An Indoor Scene and Register Ceiling-mounted Cameras
Fan Yang, Sosuke Yamao, Ikuo Kusajima +3
Using ceiling-mounted cameras (CMCs) for indoor visual capturing opens up a wide range of applications. However, registering CMCs to the target scene layout presents a challenging…
IQViC: In-context, Question Adaptive Vision Compressor for Long-term Video Understanding LMMs
Sosuke Yamao, Natsuki Miyahara, Yuki Harazono +1
With the increasing complexity of video data and the need for more efficient long-term temporal understanding, existing long-term video understanding methods often fail to accurate…