papers

Publications (20)

cs.CL2024

Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell

Taiming Lu, Muhan Gao, Kuai Yu +2

Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts. Our study explores LLMs' long-context reasoning by…

physics.app-ph2022

Tunable bilayer dielectric metasurface via stacking magnetic mirrors

Hao Song, Binbin Hong, Yanbing Qiu +3

Functional tunability, environmental adaptability, and easy fabrication are highly desired properties in metasurfaces. Here we provide a tunable bilayer metasurface composed of two…

cs.LG2025

Integrating Ontologies with Large Language Models for Enhanced Control Systems in Chemical Engineering

Crystal Su, Kuai Yu, Jingrui Zhang +2

This work presents an ontology-integrated large language model (LLM) framework for chemical engineering that unites structured domain knowledge with generative reasoning. The propo…

cs.CL2025

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

DeepSeek-AI, Aixin Liu, Aoxue Mei +260

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 ar…

cs.AI2026

How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design Factors

Kuai Yu, Naicheng Yu, Han Wang +2

Web agents have demonstrated strong performance on a wide range of web-based tasks. However, existing research on the effect of environmental variation has mostly focused on robust…

physics.app-ph2019

Strong Vibrational Coupling in Room Temperature Plasmonic Resonators

Junzhong Wang, Kuai Yu, Yang Yang +3

Strong vibrational coupling has been realized in a variety of mechanical systems from cavity optomechanics to electromechanics. It is an essential requirement for…

cs.CV2021

Points2Polygons: Context-Based Segmentation from Weak Labels Using Adversarial Networks

Kuai Yu, Hakeem Frank, Daniel Wilson

In applied image segmentation tasks, the ability to provide numerous and precise labels for training is paramount to the accuracy of the model at inference time. However, this over…

cs.CV2026

Personalize Your Large Vision-language Models With In-context Prompt Tuning

Yanshu Li, Jiaqian Li, Kuai Yu +4

Large vision-language models (LVLMs) have demonstrated strong general multimodal capability and are increasingly deployed in downstream systems. This trend has driven growing inter…

cs.CL2026

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI, Daya Guo, Dejian Yang +195

General reasoning represents a long-standing and formidable challenge in artificial intelligence. Recent breakthroughs, exemplified by large language models (LLMs) and chain-of-tho…

cs.CL2025

DeepSeek-V3 Technical Report

DeepSeek-AI, Aixin Liu, Bei Feng +195

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effec…

cs.CV2025

LyTimeT: Towards Robust and Interpretable State-Variable Discovery

Kuai Yu, Crystal Su, Xiang Liu +3

Extracting the true dynamical variables of a system from high-dimensional video is challenging due to distracting visual factors such as background motion, occlusions, and texture…

cs.CL2026

mHC: Manifold-Constrained Hyper-Connections

Zhenda Xie, Yixuan Wei, Huanqi Cao +17

Recently, studies exemplified by Hyper-Connections (HC) have extended the ubiquitous residual connection paradigm established over the past decade by expanding the residual stream…

cs.LG2025

Hierarchical Bayesian Model for Gene Deconvolution and Functional Analysis in Human Endometrium Across the Menstrual Cycle

Crystal Su, Kuai Yu, Mingyuan Shao +1

Bulk tissue RNA sequencing of heterogeneous samples provides averaged gene expression profiles, obscuring cell type-specific dynamics. To address this, we present a probabilistic h…

cs.AR2025

Large Processor Chip Model

Kaiyan Chang, Mingzhi Chen, Yunji Chen +40

Computer System Architecture serves as a crucial bridge between software applications and the underlying hardware, encompassing components like compilers, CPUs, coprocessors, and R…

cs.DC2026

FWeb3: A Practical Incentive-Aware Federated Learning Framework

Peishen Yan, Shuang Liang, Yang Hua +9

Federated learning (FL) enables collaborative model training over distributed private data. However, sustaining open participation requires incentive mechanisms that compensate con…

cs.LG2025

POLAR: Policy-based Layerwise Reinforcement Learning Method for Stealthy Backdoor Attacks in Federated Learning

Kuai Yu, Xiaoyu Wu, Peishen Yan +6

Federated Learning (FL) enables decentralized model training across multiple clients without exposing local data, but its distributed feature makes it vulnerable to backdoor attack…

math.CO2014

G-parking functions and tree inversions

David Perkinson, Qiaoyu Yang, Kuai Yu

A depth-first search version of Dhar's burning algorithm is used to give a bijection between the parking functions of a graph and labeled spanning trees, relating the degree of the…

physics.optics2020

Disorder-immune metasurfaces with constituents exhibiting the anapole mode

Hao Song, Neng Wang, Kuai Yu +2

Common optical metasurfaces are 2-dimensional functional devices composed of periodically arranged subwavelength constituents. Here, we achieved the positional-disorder-immune meta…

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…

cs.PF2025

DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs

Mingkai Chen, Tianhua Han, Cheng Liu +8

Approximate nearest neighbor search (ANNS) is essential for applications like recommendation systems and retrieval-augmented generation (RAG) but is highly I/O-intensive and memory…