activity
20242026
most citedBridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

4 citations · 4 across the 2 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning

Junha Song, Yongsik Jo, So Yeon Min +4

Systems such as video chatbots and navigation robots often depend on streaming image captioning to interpret visual inputs. Existing approaches typically employ large multimodal la…

cs.CV2025

Unsupervised Discovery of Long-Term Spatiotemporal Periodic Workflows in Human Activities

Fan Yang, Quanting Xie, Atsunori Moteki +5

Periodic human activities with implicit workflows are common in manufacturing, sports, and daily life. While short-term periodic activities -- characterized by simple structures an…

cs.CV2025

T*: Re-thinking Temporal Search for Long-Form Video Understanding

Jinhui Ye, Zihan Wang, Haosen Sun +9

Efficiently understanding long-form videos remains a significant challenge in computer vision. In this work, we revisit temporal search paradigms for long-form video understanding…

cs.CV2024

DiffusionPID: Interpreting Diffusion via Partial Information Decomposition

Rushikesh Zawar, Shaurya Dewan, Prakanshul Saxena +3

Text-to-image diffusion models have made significant progress in generating naturalistic images from textual inputs, and demonstrate the capacity to learn and represent complex vis…

cs.CV2024

DegustaBot: Zero-Shot Visual Preference Estimation for Personalized Multi-Object Rearrangement

Benjamin A. Newman, Pranay Gupta, Kris Kitani +3

De gustibus non est disputandum ("there is no accounting for others' tastes") is a common Latin maxim describing how many solutions in life are determined by people's personal pref…