Publications (20)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
Taiming Lu, Muhan Gao, Kuai Yu +2
Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts. Our study explores LLMs' long-context reasoning by…
Tunable bilayer dielectric metasurface via stacking magnetic mirrors
Hao Song, Binbin Hong, Yanbing Qiu +3
Functional tunability, environmental adaptability, and easy fabrication are highly desired properties in metasurfaces. Here we provide a tunable bilayer metasurface composed of two…
Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering
Crystal Su, Kuai Yu, Jingrui Zhang +2
This work presents an ontology-integrated large language model (LLM) framework for chemical engineering that unites structured domain knowledge with generative reasoning. The propo…
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
DeepSeek-AI, Aixin Liu, Aoxue Mei +260
We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 ar…
How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors
Kuai Yu, Naicheng Yu, Han Wang +2
Web agents have demonstrated strong performance on a wide range of web-based tasks. However, existing research on the effect of environmental variation has mostly focused on robust…
Strong Vibrational Coupling in Room Temperature Plasmonic Resonators
Junzhong Wang, Kuai Yu, Yang Yang +3
Strong vibrational coupling has been realized in a variety of mechanical systems from cavity optomechanics to electromechanics. It is an essential requirement for…
Points2Polygons: Context-Based Segmentation from Weak Labels Using Adversarial Networks
Kuai Yu, Hakeem Frank, Daniel Wilson
In applied image segmentation tasks, the ability to provide numerous and precise labels for training is paramount to the accuracy of the model at inference time. However, this over…
Personalize Your Large Vision-language Models With In-context Prompt Tuning
Yanshu Li, Jiaqian Li, Kuai Yu +4
Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This trend has driven growing inter…
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
DeepSeek-AI, Daya Guo, Dejian Yang +195
General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-tho…
DeepSeek-V3 Technical Report
DeepSeek-AI, Aixin Liu, Bei Feng +195
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effec…
LyTimeT: Towards Robust and Interpretable State-Variable Discovery
Kuai Yu, Crystal Su, Xiang Liu +3
Extracting the true dynamical variables of a system from high-dimensional video is challenging due to distracting visual factors such as background motion, occlusions, and texture…
mHC: Manifold-Constrained Hyper-Connections
Zhenda Xie, Yixuan Wei, Huanqi Cao +17
Recently, studies exemplified by Hyper-Connections (HC) have extended the ubiquitous residual connection paradigm established over the past decade by expanding the residual stream…
Hierarchical Bayesian Model for Gene Deconvolution and Functional Analysis in Human Endometrium Across the Menstrual Cycle
Crystal Su, Kuai Yu, Mingyuan Shao +1
Bulk tissue RNA sequencing of heterogeneous samples provides averaged gene expression profiles, obscuring cell type-specific dynamics. To address this, we present a probabilistic h…
Large Processor Chip Model
Kaiyan Chang, Mingzhi Chen, Yunji Chen +40
Computer System Architecture serves as a crucial bridge between software applications and the underlying hardware, encompassing components like compilers, CPUs, coprocessors, and R…
FWeb3: A Practical Incentive-Aware Federated Learning Framework
Peishen Yan, Shuang Liang, Yang Hua +9
Federated learning (FL) enables collaborative model training over distributed private data. However, sustaining open participation requires incentive mechanisms that compensate con…
POLAR: Policy-based Layerwise Reinforcement Learning Method for Stealthy Backdoor Attacks in Federated Learning
Kuai Yu, Xiaoyu Wu, Peishen Yan +6
Federated Learning (FL) enables decentralized model training across multiple clients without exposing local data, but its distributed feature makes it vulnerable to backdoor attack…
G-parking functions and tree inversions
David Perkinson, Qiaoyu Yang, Kuai Yu
A depth-first search version of Dhar's burning algorithm is used to give a bijection between the parking functions of a graph and labeled spanning trees, relating the degree of the…
Disorder-immune metasurfaces with constituents exhibiting the anapole mode
Hao Song, Neng Wang, Kuai Yu +2
Common optical metasurfaces are 2-dimensional functional devices composed of periodically arranged subwavelength constituents. Here, we achieved the positional-disorder-immune meta…
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
DeepSeek-AI, Anyi Xu, Bangcai Lin +315
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
Mingkai Chen, Tianhua Han, Cheng Liu +8
Approximate nearest neighbor search (ANNS) is essential for applications like recommendation systems and retrieval-augmented generation (RAG) but is highly I/O-intensive and memory…