v100
Here are 13 public repositories matching this topic...
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
-
Updated
Jul 10, 2026 - Python
ComfyUI custom node: run MiniMax H3 at near-fp16 speed with near-fp32 numerical stability on GPUs without bf16/fp8 hardware (V100 sm_70)
-
Updated
Aug 14, 2026 - Python
Found out that using A100 and V100 on Vicuna and Llama2 have a different result, while other model such as Falcon doesn't has such question.
-
Updated
Oct 19, 2023 - Jupyter Notebook
Performance of CUDA example benchmark code on NVIDIA A100.
-
Updated
Feb 1, 2021
SunFire V100 10 Inch Compact Case
-
Updated
Oct 6, 2025
Serve Qwen3.5-397B-A17B (AWQ) on 8x Tesla V100-SXM2-32GB (DGX-1, TP8) for agentic coding & ops — a downstream fork of 1Cat-vLLM.
-
Updated
Jul 3, 2026 - Python
Multi-GPU acceleration for MiniMax H3 video generation on NVIDIA V100 (sm_70). Ulysses sequence parallelism as a drop-in ComfyUI custom node — ~19 min to ~7 min on 8x V100.
-
Updated
Aug 14, 2026 - Python
Throughput-oriented Llama 3.1 405B inference on 2 Taiwania 2 nodes / 16 V100 GPUs: continuous batching, a V100-compatible FlashAttention backend, and cross-node tensor parallelism over NCCL InfiniBand/GDRDMA.
-
Updated
Aug 17, 2026 - Shell
VastLLM: a production-oriented FastLLM fork for native C++ inference, V100/SM70, long context, and Qwen3.5/3.6; upstream: ztxz16/fastllm
-
Updated
Aug 19, 2026 - C++
Improve this page
Add a description, image, and links to the v100 topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the v100 topic, visit your repo's landing page and select "manage topics."