5 papers
ROSER: Few-Shot Robotic Sequence Retrieval for Scalable Robot Learning
Zillur Rahman, Eddison Pham, Alejandro Daniel Noel +1
A critical bottleneck in robot learning is the scarcity of task-labeled, segmented training data, despite the abundance of large-scale robotic datasets recorded as long, continuous…
Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
Zillur Rahman, Alex Sheng, Cristian Meo
While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prom…
DynaStride: Dynamic Stride Windowing with MMCoT for Instructional Multi-Scene Captioning
Eddison Pham, Prisha Priyadarshini, Adrian Maliackel +3
Scene-level captioning in instructional videos can enhance learning by requiring an understanding of both visual cues and temporal structure. By aligning visual cues with textual g…
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
Nicholas Babey, Tiffany Gu, Yiheng Li +2
For embodied agents to effectively understand and interact within the world around them, they require a nuanced comprehension of human actions grounded in physical space. Current a…
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
Haochen Su, Cristian Meo, Francesco Stella +3
Robotic systems are increasingly expected to operate in human-centered, unstructured environments where safety, adaptability, and generalization are essential. Vision-Language-Acti…