10 papers
E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning
Zihan Liao, Jun Wang, Hang Yu +3
Processing long contexts is increasingly important for Large Language Models (LLMs) in tasks like multi-turn dialogues, code generation, and document summarization. This paper addr…
MERIT: Memory-Enhanced Retrieval for Interpretable Knowledge Tracing
Runze Li, Kedi Chen, Guwei Feng +3
Knowledge Tracing (KT) models students' evolving knowledge states to predict future performance, serving as a foundation for personalized education. While traditional deep learning…
Markovian Pre-Trained Transformer for Next-Item Recommendation
Cong Xu, Guoliang Li, Jun Wang +1
We introduce the Markovian Pre-trained Transformer (MPT) for next-item recommendation, a transferable model fully pre-trained on synthetic Markov chains, yet capable of achieving s…
Large-Model AI for Near Field Beam Prediction: A CNN-GPT2 Framework for 6G XL-MIMO
Wang Liu, Cunhua Pan, Hong Ren +3
The emergence of extremely large-scale antenna arrays (ELAA) in millimeter-wave (mmWave) communications, particularly in high-mobility scenarios, highlights the importance of near-…
Attention Residual Fusion Network with Contrast for Source-free Domain Adaptation
Renrong Shao, Wei Zhang, Jun Wang
Source-free domain adaptation (SFDA) involves training a model on source domain and then applying it to a related target domain without access to the source data and labels during…
Conditional Pseudo-Supervised Contrast for Data-Free Knowledge Distillation
Renrong Shao, Wei Zhang, Jun wang
Data-free knowledge distillation~(DFKD) is an effective manner to solve model compression and transmission restrictions while retaining privacy protection, which has attracted exte…