2 papers
cond-mat.other2026
A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
A. C. Opus, J. Q. Lu
The paper describes how a modern multimodal assistant model, MiniCPM-V-4.6, was fully implemented and run on an older 2011 NVIDIA Tesla C2075 GPU using all‑GPU CUDA inference, deta…
cond-mat.other2026
A 35B Hybrid-Attention Mixture-of-Experts Model on a 6GB 2011 GPU: Hand-Written 4-bit CUDA Inference for Fermi
A. C. Opus, J. Q. Lu
We report end-to-end inference of \textbf{Qwen3.6-35B-A3B} -- a 35-billion-parameter, 3B-active Mixture-of-Experts (MoE) model with a hybrid gated-delta-net / full-attention…