6 papers
Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding
Changjiang Jiang, Qiannian Zhao, Lei Xin +3
Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However, this reliance introduces significant inf…
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
Lei Xin, Bin Gu, Peize Li +9
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still stores and routes over its full expert bank. W…
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
Zhenglun Kong, Yize Li, Fanhu Zeng +7
In Transformer architectures, tokens\textemdash discrete units derived from raw data\textemdash are formed by segmenting inputs into fixed-length chunks. Each token is then mapped…
HyTRec: A Hybrid Temporal-Aware Attention Architecture for Long Behavior Sequential Recommendation
Lei Xin, Yuhao Zheng, Ke Cheng +3
Modeling long sequences of user behaviors has emerged as a critical frontier in generative recommendation. However, existing solutions face a dilemma: linear attention mechanisms a…
Life-Code: Central Dogma Modeling with Multi-Omics Sequence Unification
Zicheng Liu, Siyuan Li, Zhiyuan Chen +6
The interactions between DNA, RNA, and proteins are fundamental to biological processes, as illustrated by the central dogma of molecular biology. Although modern biological pre-tr…
Artificial Intelligence for Central Dogma-Centric Multi-Omics: Challenges and Breakthroughs
Lei Xin, Caiyun Huang, Hao Li +8
With the rapid development of high-throughput sequencing platforms, an increasing number of omics technologies, such as genomics, metabolomics, and transcriptomics, are being appli…