activity
20242026
most citedSIU: A Million-Scale Structural Small Molecule-Protein Interaction Dataset for Unbiased Bioactivity Prediction

2 citations · 4 across the 5 of their papers we have counts for

collaborators

10 papers

cs.CL2026

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

DeepSeek-AI, Anyi Xu, Bangcai Lin +315

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSe…

cs.GR2025

Learning-based density-equalizing map

Yanwen Huang, Lok Ming Lui, Gary P. T. Choi

Density-equalizing map (DEM) serves as a powerful technique for creating shape deformations with the area changes reflecting an underlying density function. In recent decades, DEM…

cs.CL2025

Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression

Yong Zhang, Heng Li, Yanwen Huang +6

Retrieval-augmented generation (RAG) often suffers from long and noisy retrieved contexts. Existing context compression methods typically rely on heuristic relevance estimation or…

q-bio.BM20252 cited

PharmAgents: Building a Virtual Pharma with Large Language Model Agents

Bowen Gao, Yanwen Huang, Yiqiao Liu +4

The discovery of novel small molecule drugs remains a critical scientific challenge with far-reaching implications for treating diseases and advancing human health. Traditional dru…

q-bio.BM2025

Pushing the boundaries of Structure-Based Drug Design through Collaboration with Large Language Models

Bowen Gao, Yanwen Huang, Yiqiao Liu +4

Structure-Based Drug Design (SBDD) has revolutionized drug discovery by enabling the rational design of molecules for specific protein targets. Despite significant advancements in…

cs.IR2025

Large Language Model as Universal Retriever in Industrial-Scale Recommender System

Junguang Jiang, Yanwen Huang, Bin Liu +6

In real-world recommender systems, different retrieval objectives are typically addressed using task-specific datasets with carefully designed model architectures. We demonstrate t…