2 papers
cs.CV2026
Memory Tree Guided Key Frame Querying for Efficient 3D Question Answering
Hsiang-Wei Huang, Fu-Chen Chen, Li-Wu Tsao +7
Answering questions accurately and efficiently in embodied scenarios presents significant challenges due to limited computational and memory resources for Vision Language Model (VL…
cs.CV2024
A Cat Is A Cat (Not A Dog!): Unraveling Information Mix-ups in Text-to-Image Encoders through Causal Analysis and Embedding Optimization
Chieh-Yun Chen, Chiang Tseng, Li-Wu Tsao +1
This paper analyzes the impact of causal manner in the text encoder of text-to-image (T2I) diffusion models, which can lead to information bias and loss. Previous works have focuse…