2 papers
cs.AR2026
The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference
Michael J. Yuan, Ju Long
Every mainstream GPU is built compute-heavy and capacity-light: it pairs enormous arithmetic throughput with too little memory to hold a modern model. In contrast, large language m…
cs.AI2025
Trust, but verify
Michael J. Yuan, Carlos Lospoy, Sydney Lai +2
Decentralized AI agent networks, such as Gaia, allows individuals to run customized LLMs on their own computers and then provide services to the public. However, in order to mainta…