Here are
9 public repositories
matching this topic...
GLM-5.3-Flash (320B MoE) served with vLLM on 8x CMP 170HX (SM80, 64 GB, PCIe Gen2 x4): pipeline-parallel 8, NVFP4 weights, DFlash2 speculative decoding, 1M context. Patches, launch config and benchmarks.
Updated
Aug 31, 2026
Python
Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
Updated
Aug 1, 2026
Python
Reproducible vLLM deployment for GLM-5.3-Flash on 8x NVIDIA A800 SM80 GPUs
Updated
Aug 30, 2026
Python
Evidence-led vLLM tuning and qualification notes for NVIDIA CMP 170HX (sm80)
DeepSeek-V4-Flash-Vision-Exp on NVIDIA CMP 170HX/SM80 with vLLM, PP DSpark and 1M context
Updated
Sep 4, 2026
Python
Measured llama.cpp Q6_K CUDA optimizations for NVIDIA A100 (SM80)
Strict-FP32 CUDA SGEMM with shape-dispatched SM80 cp.async kernels and reproducible A100 evidence
Updated
Aug 21, 2026
Python
Native dual-GPU long-context inference runtime for Qwen3.8-Flash-Next
From-scratch SM80 (A100/A800) CUDA kernels for GDN (Gated DeltaNet) and QSA (query-key sparse attention). Apache-2.0.
Add this topic to your repo
To associate your repository with the
sm80
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.