2 citations · 2 across the 9 of their papers we have counts for
14 papers
Large VLM-based Stylized Sports Captioning
Sauptik Dhar, Nicholas Buoncristiani, Joe Anakata +2
The advent of large (visual) language models (LLM / LVLM) have led to a deluge of automated human-like systems in several domains including social media content generation, search…
Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
Qiaohui Chu, Haoyu Zhang, Meng Liu +3
Long-term action anticipation from egocentric video is critical for applications such as human-computer interaction and assistive technologies, where anticipating user intent enabl…
SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
Yunfei Xie, Yuxuan Cheng, Juncheng Wu +3
Recent advancements in adapting vision-language pre-training models like CLIP for person re-identification (ReID) tasks often rely on complex adapter design or modality-specific tu…
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
Qiaohui Chu, Haoyu Zhang, Yisen Feng +4
In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our met…
OSGNet @ Ego4D Episodic Memory Challenge 2025
Yisen Feng, Haoyu Zhang, Qiaohui Chu +4
In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise…
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
Haoyu Zhang, Meng Liu, Zaijing Li +4
Visual-spatial understanding, the ability to infer object relationships and layouts from visual input, is fundamental to downstream tasks such as robotic navigation and embodied in…