6 papers
EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms
Brian VanVoorst, Nicholas Walczak, Christopher Gilleo +9
This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled…
STORM: End-to-End Referring Multi-Object Tracking in Videos
Zijia Lu, Jingru Yi, Jue Wang +4
Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMO…
Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
Bin Hu, Zijian Lu, Haicheng Liao +7
Motion planning for autonomous driving must handle multiple plausible futures while remaining computationally efficient. Recent end-to-end systems and world-model-based planners pr…
PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving
Chengran Yuan, Zijian Lu, Zhanqi Zhang +8
End-to-end motion planning is promising for simplifying complex autonomous driving pipelines. However, challenges such as scene understanding and effective prediction for decision-…
OptiPrune: Boosting Prompt-Image Consistency with Attention-Guided Noise and Dynamic Token Selection
Ziji Lu
Text-to-image diffusion models often struggle to achieve accurate semantic alignment between generated images and text prompts while maintaining efficiency for deployment on resour…
DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos
Zijia Lu, A S M Iftekhar, Gaurav Mittal +6
Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach ta…