3 papers
cs.RO2026
EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations
Yangcen Liu, Shuo Cheng, Xinchen Yin +6
Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but…
cs.CV2026
Do All Vision Transformers Need Registers? A Cross-Architectural Reassessment
Spiros Baxevanakis, Platon Karageorgis, Ioannis Dravilas +1
Training Vision Transformers (ViTs) presents significant challenges, one of which is the emergence of artifacts in attention maps, hindering their interpretability. Darcet et al. (…
cs.CV2026
BLM-Guard: Explainable Multimodal Ad Moderation with Chain-of-Thought and Policy-Aligned Rewards
Yiran Yang, Zhaowei Liu, Yuan Yuan +10
Short-video platforms now host vast multimodal ads whose deceptive visuals, speech and subtitles demand finer-grained, policy-driven moderation than community safety filters. We pr…