activity
20232025
most citedAnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

7 citations · 11 across the 8 of their papers we have counts for

collaborators
Showing 2024 · cs.CVShow all

5 papers · 2 filters

cs.CV2024

Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang +4

Goal-oriented planning, or anticipating a series of actions that transition an agent from its current state to a predefined objective, is crucial for developing intelligent assista…

cs.CV2024★ 1 cited

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

Zi-Yi Dou, Xitong Yang, Tushar Nagarajan +5

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-s…

cs.CV2024★ 2 cited

ExpertAF: Expert Actionable Feedback from Video

Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos +2

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demon…

cs.CV2024

Step Differences in Instructional Video

Tushar Nagarajan, Lorenzo Torresani

Comparing a user video to a reference how-to video is a key requirement for AR/VR technology delivering personalized assistance tailored to the user's progress. However, current ap…

cs.CV2024★ 1 cited

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang +3

Most video captioning models are designed to process short video clips of few seconds and output text describing low-level visual concepts (e.g., objects, scenes, atomic actions).…