6 papers
Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding
Xiaoli Yang, Huiyuan Tian, Yurui Li +3
Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited information bandwidth. This raises a central q…
AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
Bo Yang, Lanfei Feng, Yunkui Chen +5
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the lack of multilingual speech data, unified multimodal architectures,…
AgriGPT-VL: Agricultural Vision-Language Understanding Suite
Bo Yang, Yunkui Chen, Lanfei Feng +8
Despite rapid advances in multimodal large language models, agricultural applications remain constrained by the scarcity of domain-tailored models, curated vision-language corpora,…
AgriGPT: a Large Language Model Ecosystem for Agriculture
Bo Yang, Yu Zhang, Lanfei Feng +10
Despite the rapid progress of Large Language Models (LLMs), their application in agriculture remains limited due to the lack of domain-specific models, curated datasets, and robust…
Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation
Jianyu Zhang, Li Zhang, Shijian Li
The visual understanding are often approached from 3 granular levels: image, patch and pixel. Visual Tokenization, trained by self-supervised reconstructive learning, compresses vi…
A Framework For Image Synthesis Using Supervised Contrastive Learning
Yibin Liu, Jianyu Zhang, Li Zhang +2
Text-to-image (T2I) generation aims at producing realistic images corresponding to text descriptions. Generative Adversarial Network (GAN) has proven to be successful in this task.…