3 papers
cs.CV2026
ViLL-E: Video LLM Embeddings for Retrieval
Rohit Gupta, Jayakrishnan Unnikrishnan, Fan Fei +3
Video Large Language Models (VideoLLMs) excel at video understanding tasks where outputs are textual, such as Video Question Answering and Video Captioning. However, they underperf…
eess.SY2026
Modular Safety Guardrails Are Necessary for Foundation-Model-Enabled Robots in the Real World
Joonkyung Kim, Wenxi Chen, Davood Soleymanzadeh +9
The integration of foundation models (FMs) into robotics has accelerated real-world deployment, while introducing new safety challenges arising from open-ended semantic reasoning a…
cs.LG2025
GFocal: A Global-Focal Neural Operator for Solving PDEs on Arbitrary Geometries
Fangzhi Fei, Jiaxin Hu, Qiaofeng Li +1
Transformer-based neural operators have emerged as promising surrogate solvers for partial differential equations, by leveraging the effectiveness of Transformers for capturing lon…