#
int2
Here are 2 public repositories matching this topic...
SSD-resident INT2/INT4 KV cache with asynchronous staging and fused CUDA dequantizing decode attention for long-context LLM inference.
cuda pytorch ssd attention quantization nvme asynchronous-io gpu-kernels gqa kv-cache int4 long-context llm-inference int2 inference-systems
-
Updated
Jul 19, 2026 - Python
Add this topic to your repo
To associate your repository with the int2 topic, visit your repo's landing page and select "manage topics."