From the 1 of 6 linked papers with an AI index.
6 papers
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results
Sheng Xu, Junhua Wang, Boyuan Huang +4
The paper proposes Amplitude Gating, an inference‑time method that adjusts FFN activation magnitudes to improve tool‑structured outputs of large language models without changing mo…
FlowTrain: Flow-Based Decoupled Training for Industrial-Grade Vision-Language Models
Zhida Jiang, Zhaolong Xing, Yang Pei +14
Industrial-grade distributed training of vision-language models (VLMs) remains far less efficient than that of unimodal LLMs. Existing solutions either follow a monolithic design t…
NestPipe: Large-Scale Recommendation Training on 1,500+ Accelerators via Nested Pipelining
Zhida Jiang, Zhaolong Xing, Huichao Chai +12
Modern recommendation models have increased to trillions of parameters. As cluster scales expand to O(1k), distributed training bottlenecks shift from computation and memory to dat…
XAttnRes: Cross-Stage Attention Residuals for Medical Image Segmentation
Xinyu Liu, Qing Xu, Zhen Chen
In the field of Large Language Models (LLMs), Attention Residuals have recently demonstrated that learned, selective aggregation over all preceding layer outputs can outperform fix…
Harnessing Lightweight Transformer with Contextual Synergic Enhancement for Efficient 3D Medical Image Segmentation
Xinyu Liu, Zhen Chen, Wuyang Li +2
Transformers have shown remarkable performance in 3D medical image segmentation, but their high computational requirements and need for large amounts of labeled data limit their ap…
Semantic Consistency Regularization with Large Language Models for Semi-supervised Sentiment Analysis
Kunrong Li, Xinyu Liu, Zhen Chen
Accurate sentiment analysis of texts is crucial for a variety of applications, such as understanding customer feedback, monitoring market trends, and detecting public sentiment. Ho…