collaborators

6 papers

cs.CV2026

EgoMAGIC- An Egocentric Video Field Medicine Dataset for Training Perception Algorithms

Brian VanVoorst, Nicholas Walczak, Christopher Gilleo +9

This paper introduces EgoMAGIC (Medical Assistance, Guidance, Instruction, and Correction), an egocentric medical activity dataset collected as part of DARPA's Perceptually-enabled…

cs.CV2026

STORM: End-to-End Referring Multi-Object Tracking in Videos

Zijia Lu, Jingru Yi, Jue Wang +4

Referring multi-object tracking (RMOT) is a task of associating all the objects in a video that semantically match with given textual queries or referring expressions. Existing RMO…

cs.CV2025

Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving

Bin Hu, Zijian Lu, Haicheng Liao +7

Motion planning for autonomous driving must handle multiple plausible futures while remaining computationally efficient. Recent end-to-end systems and world-model-based planners pr…

cs.RO2025

PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving

Chengran Yuan, Zijian Lu, Zhanqi Zhang +8

End-to-end motion planning is promising for simplifying complex autonomous driving pipelines. However, challenges such as scene understanding and effective prediction for decision-…

cs.CV2025

OptiPrune: Boosting Prompt-Image Consistency with Attention-Guided Noise and Dynamic Token Selection

Ziji Lu

Text-to-image diffusion models often struggle to achieve accurate semantic alignment between generated images and text prompts while maintaining efficiency for deployment on resour…

cs.CV2025

DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long Videos

Zijia Lu, A S M Iftekhar, Gaurav Mittal +6

Long Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach ta…