◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Dongyang Ma

10 papers hereh-index 446 citations11 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author1
  • middle author9

Across the 10 of 10 papers where every author was matched, so the position is known.

fields
  • cs.AI3
  • cs.CL3
  • cs.LG2
  • cs.NE2

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.LGShow all

2 papers · 1 filter

cs.LG2026

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Yan Wang, Qifan Zhang, Jiachen Yu +12

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In this report, we propose \textbf{Lookahead…

cs.LG2025

Block-Attention for Efficient Prefilling

Dongyang Ma, Yan Wang, Lan Tian

We introduce Block-attention, an attention mechanism designed to address the increased inference latency and cost in Retrieval-Augmented Generation (RAG) scenarios. Traditional app…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.