4 papers · 1 filter
SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis
Jongbeom Lee, Hyunwoo Yu, Jincheol Yang +2
InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context, however, leave late-scale a…
Dual Anchors, Do It Better: Hierarchical Group Merging for Zero-Shot Anomaly Detection
Jimin Roh, DongKyu Kim, Suk-Ju Kang
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts…
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
Beoungwoo Kang, Seunghun Moon, Yubin Cho +2
Beyond the Transformer, it is important to explore how to exploit the capacity of the MetaFormer, an architecture that is fundamental to the performance improvements of the Transfo…
Embedding-Free Transformer with Inference Spatial Reduction for Efficient Semantic Segmentation
Hyunwoo Yu, Yubin Cho, Beoungwoo Kang +3
We present an Encoder-Decoder Attention Transformer, EDAFormer, which consists of the Embedding-Free Transformer (EFT) encoder and the all-attention decoder leveraging our Embeddin…