38 citations · 38 across the 4 of their papers we have counts for
5 papers
SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis
Jongbeom Lee, Hyunwoo Yu, Jincheol Yang +2
InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context, however, leave late-scale a…
VIPA: Visual Informative Part Attention for Referring Image Segmentation
Yubin Cho, Hyunwoo Yu, Kyeongbo Kong +3
Referring Image Segmentation (RIS) aims to segment a target object described by a natural language expression. Existing methods have evolved by leveraging the vision information in…
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
Beoungwoo Kang, Seunghun Moon, Yubin Cho +2
Beyond the Transformer, it is important to explore how to exploit the capacity of the MetaFormer, an architecture that is fundamental to the performance improvements of the Transfo…
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
Yubin Cho, Hyunwoo Yu, Suk-ju Kang
Referring segmentation aims to segment a target object related to a natural language expression. Key challenges of this task are understanding the meaning of complex and ambiguous…
Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang +3
We present an Encoder-Decoder Attention Transformer, EDAFormer, which consists of the Embedding-Free Transformer (EFT) encoder and the all-attention decoder leveraging our Embeddin…