14 papers
Zero-shot 2D Grounding with Novel Affordance Types
Haomeng Zhang, Raymond A. Yeh
2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types seen during training and do…
Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
Amber Yijia Zheng, Lu Liu, Raymond A. Yeh +1
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often und…
DREAM-Chunk: Reactive Action Chunking with Latent World Model
Wenxi Chen, Kaidi Zhang, Chi Lin +6
Action chunking has become a common interface for vision-language-action (VLA) models, enabling low-frequency policy inference to drive high-frequency robot execution. However, onc…
4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
Chiao-An Yang, Ryo Hachiuma, Sifei Liu +4
Despite advances in Multimodal LLMs (MLLMs), their ability to reason over 3D structures and temporal dynamics remains limited, constrained by weak 4D perception and temporal unders…
Designing to Forget: Deep Semi-parametric Models for Unlearning
Amber Yijia Zheng, Yu-Shan Tai, Raymond A. Yeh
Recent advances in machine unlearning have focused on developing algorithms to remove specific training samples from a trained model. In contrast, we observe that not all models ar…
WebAccessVL: Violation-Aware VLM for Web Accessibility
Amber Yijia Zheng, Jae Joong Lee, Bedrich Benes +1
We present a vision-language model (VLM) that automatically edits website HTML to address violations of the Web Content Accessibility Guidelines 2 (WCAG2) while preserving the orig…