Publications (21)
Prefix-Adaptive Block Diffusion for Efficient Document Recognition
Mingxu Chai, Ziyu Shen, Chenyu Liu +9
Block Diffusion Models (BDMs) support parallel generation, flexible-length output, and KV caching, making them promising for efficient document parsing. However, existing BDMs bind…
A Filter of Minhash for Image Similarity Measures
Jun Long, Qunfeng Liu, Xinpan Yuan +2
Image similarity measures play an important role in nearest neighbor search and duplicate detection for large-scale image datasets. Recently, Minwise Hashing (or Minhash) and its r…
TriPAH: Imbalance-Aware Tri-Prompt Affinity Hashing for Cross-Modal Medical Retrieval
Jiaming Bian, Songming Li, Yurui Song +3
In the era of big medical data, efficient cross-modal retrieval is pivotal for evidence-based diagnosis and large-scale case management. Cross-modal medical hashing retrieval aims…
HOC-Tree: A Novel Index for efficient Spatio-temporal Range Search
Jun Long, Lei Zhu, Chengyuan Zhang +3
With the rapid development of mobile computing and Web services, a huge amount of data with spatial and temporal information have been collected everyday by smart mobile terminals,…
Hierarchical One Permutation Hashing: Efficient Multimedia Near Duplicate Detection
Chengyuan Zhang, Yunwu Lin, Lei Zhu +3
With advances in multimedia technologies and the proliferation of smart phone, digital cameras, storage devices, there are a rapidly growing massive amount of multimedia data colle…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
Omni-Seg: A Scale-aware Dynamic Network for Renal Pathological Image Segmentation
Ruining Deng, Quan Liu, Can Cui +9
Comprehensive semantic segmentation on renal pathological images is challenging due to the heterogeneous scales of the objects. For example, on a whole slide image (WSI), the cross…
Compound Figure Separation of Biomedical Images: Mining Large Datasets for Self-supervised Learning
Tianyuan Yao, Chang Qu, Jun Long +13
With the rapid development of self-supervised learning (e.g., contrastive learning), the importance of having large-scale images (even without annotations) for training a more gene…
Asymmetric Deep Semantic Quantization for Image Retrieval
Zhan Yang, Osolo Ian Raymond, WuQing Sun +1
Due to its fast retrieval and storage efficiency capabilities, hashing has been widely used in nearest neighbor retrieval tasks. By using deep learning based techniques, hashing ca…
Temporal Activity Path Based Character Correction in Social Networks
Jun Long, Lei Zhu, Zhan Yang +2
Vast amount of multimedia data contains massive and multifarious social information which is used to construct large-scale social networks. In a complex social network, a character…
Efficient Top K Temporal Spatial Keyword Search
Chengyuan Zhang, Lei Zhu, Weiren Yu +3
Massive amount of data that are geo-tagged and associated with text information are being generated at an unprecedented scale in many emerging applications such as location based s…
ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
Ke Ma, Jun Long, Hongxiao Fei +3
Pre-trained Vision-Language Models (VLMs) struggle with Zero-Shot Anomaly Detection (ZSAD) due to a critical adaptation gap: they lack the local inductive biases required for dense…
Glo-In-One: Holistic Glomerular Detection, Segmentation, and Lesion Characterization with Large-scale Web Image Mining
Tianyuan Yao, Yuzhe Lu, Jun Long +6
The quantitative detection, segmentation, and characterization of glomeruli from high-resolution whole slide imaging (WSI) play essential roles in the computer-assisted diagnosis a…
DFTerNet: Towards 2-bit Dynamic Fusion Networks for Accurate Human Activity Recognition
Zhan Yang, Osolo Ian Raymond, ChengYuan Zhang +2
Deep Convolutional Neural Networks (DCNNs) are currently popular in human activity recognition applications. However, in the face of modern artificial intelligence sensor-based gam…
HiFuse: Hierarchical Multi-Scale Feature Fusion Network for Medical Image Classification
Xiangzuo Huo, Gang Sun, Shengwei Tian +5
Medical image classification has developed rapidly under the impetus of the convolutional neural network (CNN). Due to the fixed size of the receptive field of the convolution kern…
An Efficient Approach for Geo-Multimedia Cross-Modal Retrieval
Lei Zhu, Jun Long, Chengyuan Zhang +3
Due to the rapid development of mobile Internet techniques, cloud computation and popularity of online social networking and location-based services, massive amount of multimedia d…
Deep Attention-guided Hashing
Zhan Yang, Osolo Ian Raymond, Wuqing Sun +1
With the rapid growth of multimedia data (e.g., image, audio and video etc.) on the web, learning-based hashing techniques such as Deep Supervised Hashing (DSH) have proven to be v…
Efficient Interactive Search for Geo-tagged Multimedia Data
Jun Long, Lei Zhu, Chengyuan Zhang +3
Due to the advances in mobile computing and multimedia techniques, there are vast amount of multimedia data with geographical information collected in multifarious applications. In…
An Enhanced Latent Semantic Analysis Approach for Arabic Document Summarization
Kamal Al-Sabahi, Zuping Zhang, Jun Long +1
The fast-growing amount of information on the Internet makes the research in automatic document summarization very urgent. It is an effective solution for information overload. Man…
Multi-object Tracking with a Hierarchical Single-branch Network
Fan Wang, Lei Luo, En Zhu +2
Recent Multiple Object Tracking (MOT) methods have gradually attempted to integrate object detection and instance re-identification (Re-ID) into a united network to form a one-stag…
Asymmetric Residual Neural Network for Accurate Human Activity Recognition
Jun Long, WuQing Sun, Zhan Yang +1
Human Activity Recognition (HAR) using deep neural network has become a hot topic in human-computer interaction. Machine can effectively identify human naturalistic activities by l…