works on

From the 2 of 5 linked papers with an AI index.

activity
20242026
collaborators

5 papers

cs.CV2026

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Fanghui Xue +5

The paper introduces VideoSEMA, a split space‑time attention model for video classification that combines a scalable Mamba‑like spatial attention block with softmax temporal attent…

cs.CV2026

SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging

Nhat Thanh Tran, Fanghui Xue, Shuai Zhang +4

The paper introduces SEMA, a new attention mechanism for vision transformers that combines token localization with arithmetic averaging to avoid the dispersion problem of linear at…

cs.CV2026

ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations

Yuhao Zhou, Yunpeng Zhu, Yang Zhou +7

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring a…

cs.CV2025

EgoPrivacy: What Your First-Person Camera Says About You?

Yijiang Li, Genpei Zhang, Jiacheng Cheng +7

While the rapid proliferation of wearable cameras has raised significant concerns about egocentric video privacy, prior work has largely overlooked the unique privacy threats posed…

cs.CV2024

AFIDAF: Alternating Fourier and Image Domain Adaptive Filters as an Efficient Alternative to Attention in ViTs

Yunling Zheng, Zeyi Xu, Fanghui Xue +5

We propose and demonstrate an alternating Fourier and image domain filtering approach for feature extraction as an efficient alternative to build a vision backbone without using th…