DGPP is a C++/CUDA inference engine for NVIDIA DGX Spark (GB10) systems. It serves GLM-5.3-Flash, full GLM-5.3, Qwen3.8-Flash-Next, GLM-4.7 and DeepSeek-V4.1-Flash through an OpenAI-compatible HTTP API, with tensor parallelism over RoCE for multi-node deployments. Supported configurations use one, two or four nodes, depending on the model and its memory requirements.
The repository contains the CUDA kernels, RDMA collectives, tokenizer, Jinja chat-template interpreter, scheduler, prefix cache and HTTP service. It uses the CUDA runtime, cuBLASLt and libibverbs. Rank 0 coordinates requests through an admission journal; peers check their operation streams against it throughout a run.
These serving configurations have deployment templates and recorded measurements. World size is the number of ranks, with one DGX Spark per rank; world 1 runs locally, while worlds 2 and 4 use tensor parallelism over RoCE. Each quant links to its specific Hugging Face model card.
| Model | Quant / Hugging Face model card | World sizes | Example configuration |
|---|---|---|---|
| GLM-5.3-Flash | unsloth/GLM-5.3-Flash-FP8 | 4 | Copy the base template to cluster_glm-5.3-flash_fp8_w4.json, set model to the linked FP8 repository and lower engine.kv_capacity to 393216 (the FP8 experts are 31 GiB larger per rank; the startup memory plan refuses the base template's context) |
| GLM-5.3-Flash (hybrid) | HawkBearPig/GLM-5.3-Flash-NVFP4-FP8 | 2, 4 | Two nodes, four nodes |
| Qwen3.8-Flash-Next | Qwen/Qwen3.8-Flash-Next-FP8 | 2, 4 | Two nodes, four nodes |
| Qwen3.8-Flash-Next | nvidia/Qwen3.8-Flash-Next-NVFP4 | 1, 2 | One node (256K shared KV pool, 4K busy/idle prefill), two nodes (mapped n-gram table, dense projections FP8 at load) |
| Qwen3.8-Flash-Next | Saren/Qwen3.8-Flash-Next-W4A16-AutoRound-hybrid-MTP_int4RTN | 1 | One node (AutoRound int4 g128 experts and an int8 g128 head served in their packed form, the n-gram table from the FP8 release's shards via engine.ngram_table_model; the template turns on the prefill_bf16_partials and prefill_fp8_gemm levers, whose measured quality cost is inside the default chain's band — see accuracy and correctness) |
| Qwen3.8-Flash-Next | RadixArk/Qwen3.8-Flash-Next-NVFP4 | 1, 2 | One node (tuned: 4K prefill chunks with matching prefill budgets, MTP depth 2), two nodes (same NVFP4 format and engine configuration as the NVIDIA release) |
| GLM-4.7 | nvidia/GLM-4.7-NVFP4 | 4 | Four nodes |
| GLM-5.3 | HawkBearPig/GLM-5.3-Int4-Int8Mix-RTN-g64 | 4 | Four nodes |
| DeepSeek-V4.1-Flash | deepseek-ai/DeepSeek-V4.1-Flash | 4 | Four nodes |
| MiMo-V2.6-Flash | XiaomiMiMo/MiMo-V2.6-Flash-RL | 2, 4 | Two nodes, four nodes |
| DeepSeek-V4-Flash | deepseek-ai/DeepSeek-V4-Flash-0731 | 2, 4 | Two nodes, four nodes |
| Qwen3.8-27B | Qwen/Qwen3.8-27B-FP8 | 1, 2, 4 | One node (the DFlash2 block drafter), two nodes and four nodes (MTP depth 3) |
The Qwen NVFP4 templates select streaming MMA for the FP8 vocabulary head
with engine.fp8_head: "mma", following matched one- and two-Spark
real-checkpoint numerical validation,
including the YaRN template. --fp8-head gemv restores the previous head
path. See operation and numerical constraints.
The Qwen NVFP4 templates use engine.ngram_table: "mmap" to read the
n-gram table from NVMe, engine.decode_graph: true for resident graph serving,
and engine.dense_weights: "fp8" to encode dense projections at load. On two
Sparks, mapping saves 23.84 GiB per rank while staying within 3.6% of resident
decode throughput and 2.9% of resident prefill time in the matched campaign.
Use --dense-weights checkpoint --fp8-head gemv to retain the checkpoint's
BF16 dense stack.
See the single-node guide and
two-node benchmark.
The GLM-5.3 hybrid takes the main-stack routed experts from
dabsLabs and the
remaining tensors, including MTP, from
Unsloth.
The GLM-5.3-Flash templates use
HawkBearPig/GLM-5.3-Flash-NVFP4-FP8. The setup command below downloads it
once on rank 0 and syncs the selected snapshot to peers. To use the FP8 release
instead, set
model to unsloth/GLM-5.3-Flash-FP8.
To reproduce the hybrid from its sources, use
the composition tool and
the NVFP4 notes.
MiMo-V2.6-Flash is served as shipped — MXFP4 routed experts, fp8 block-128
dense projections, BF16 o_proj / head / eh_proj (resident in their
lossless 12-bit form under engine.bf16_weights: "bf12") — with its hybrid
sliding-window (128, sink-biased) / global attention, one MTP draft layer
and the checkpoint's own chat template and <tool_call> format. The vision
and audio encoders in the checkpoint are not loaded; text prompts only. The
two-node template keeps its K/V cache in the fp8 row form
(engine.kv_dtype: "fp8", 58 KiB per token per rank against 109 in BF16)
for a 256K-token pool; --kv-dtype bf16 restores the BF16 cache. Decode
runs the attention as one launch per layer with the residual add fused
into each norm; prefill attention is query-tiled on the tensor cores (each
K/V tile staged — and under the fp8 cache dequantized — once per 64 query
vectors). See the model card and
plan. Native blocks 0/1/2 MTP3 and prefill
work reductions are opt-in; see MiMo operations
for settings, memory implications and measured workload tradeoffs.
DeepSeek-V4-Flash (the 0731 release) is served as shipped — MXFP4 routed
experts, FP8 block-128 attention projections and shared expert, BF16
compressors, router and head — with its 128-token sliding window, the
ratio-4 (indexed, 512 entries per query) and ratio-128 compressed caches,
token-table routing on the first three layers, two-pass hyper-connections
and the DSpark block draft. Both templates serve the model's full 1M-token
context and schedule the verify depth from the draft's confidence
(engine.mtp_schedule): a pass verifies as many of the block's five drafts
as their survival pays for. A sampled request's drafts are the draft head's
most likely tokens (engine.mtp_draft), exact for the output distribution,
and follow the same schedule (engine.mtp_schedule_sampled_scale).
Prefill is bitwise the same for any chunking of a prompt; prompts that
arrive together are read in through one forward and start decoding
together, and a later arrival interleaves with the running decodes.
Measured on four Sparks (2026-10-02): 65–126 tokens/s for one request by
prompt class (44 without the draft), 134–186 tokens/s across six, cold
prefill of a 2K prompt in 1.36 s, and on llama-benchy's sampled pp2048 /
tg128 64 tokens/s for one request and 91 / 121 in total at two and five
concurrent requests; on two Sparks 40–73 tokens/s for one request and
68–101 across four (benchmarks). See the
model card.
The linked templates enable MTP at the depth measured best for that deployment. Plain decode, deeper MTP, alternate cache formats and slot counts are launcher knobs on these templates; the deployment catalogue maps the supported shapes to their flags. The benchmark tables record the modes measured for each deployment.
- Tensor-parallel resident serving on four nodes: each rank holds its slice of the model resident on the GPU and typically boots from a per-rank image cache in 15–30 s, depending on the model.
- Graph-based decode, including collectives and MTP speculative decoding, with scalar or batched graphs selected for the active requests. Greedy MTP produces the same tokens as plain decode; sampled MTP preserves the target distribution.
- Faster DeepSeek decoder selection: parallel scoring and exact radix selection reduce the time spent choosing attention entries, especially at long context. Scores, tie rules and selected entries are preserved; existing recipes use the improvement automatically. See the comparison and validation.
- Row-aware tensor-core execution and grouped prefill: dense kernels select their lowering from the active row count, and queued cold prompts can share a forward pass while retaining request-local attention and state.
- Lossless 12-bit BF16 weights (
engine.bf16_weights): the BF16 matrices a decode step streams are kept in a 12-bit form — the sign and mantissa byte plus a 4-bit exponent code per weight, exact side tables for the rare outliers — and the kernels rebuild the exact BF16 bits in registers. Outputs are bit-identical (greedy transcripts do not change) from 25% fewer bytes: +6–12% single-stream decode on GLM-5.3-Flash, GLM-4.7, the full GLM-5.3 and Qwen3.8-Flash-Next-FP8. The 12-bit form can also be the only resident one, which puts the model below its checkpoint's footprint. - Qwen3.8-Flash-Next AutoRound int4/int8 hybrid on one Spark: the GPTQ-layout checkpoint (int4 group-128 experts, an int8 group-128 head) is served in its packed form with no dequantized copy; the head is read as bit planes with a provable argmax bound, so every greedy step touches 75 % of its bytes for a bit-exact pick, and MTP runs at depth 3. Decode 45–72 tok/s single-stream and 77–119 tok/s at four requests, prefill 0.55 ms per token from 8K to 32K on the exact chain and 0.52–0.53 with the template's two prefill levers, with the decode step flat from 3K context up. See the record.
- Model-specific prefill paths: packed int4/int8 tensor-core prefill for full GLM-5.3, tiled QSA prefill for Qwen (W4A4 expert prefill is opt-in pending the gate in issue #68), and bounded grouped prefill for DeepSeek-V4.1-Flash. Qwen can optionally yield between prefill chunks so active decodes continue making progress.
- Opt-in 512K context for Qwen3.8-Flash-Next with
engine.rope_scaling(YaRN): the two-Spark NVFP4 template supports a 524288-token request ceiling, with 5/5 retrieval probes passing at both 261K and 522K prompt tokens. See the validation record and release-check procedure. - Adaptive DSpark verification: DeepSeek uses confidence-scheduled draft depth, including a batch-aware rule and an adaptive value of decode time.
- Exact prefix caching: matching token prefixes can reuse a stored session snapshot while preserving the cold-prefill result.
- OpenAI-compatible text generation: Chat Completions and legacy
Completions, streaming, constrained tool calls,
response_formatwith supported JSON schemas,reasoning_content, logprobs,stop,n,logit_biasand usage details; plus model, health and metrics endpoints. Assistant history accepts Anthropic-style thinking parts forwarded through LiteLLM. See the API compatibility profile and live prefill metrics for supported options and model-dependent limitations. - Image inputs: PNG/JPEG/WebP data URIs in Chat Completions for GLM-5.3-Flash and Qwen3.8-Flash-Next, using each checkpoint's native vision encoder. Multiple images, streaming and MTP work together, with image-aware prefix caching and resumable prefill; see image inputs for each family's preprocessing geometry, examples and memory requirements.
- Deterministic across ranks: admissions journaled from the head, every tick's operation-stream digest checked on every peer, and all ranks' complete streams compared at shutdown.
- Process-failure detection: a rank process exiting fails the service within seconds. Silent node loss is detected by the bus watchdog.
- Deployment and monitoring: a shared cluster config, versioned
releases, periodic throughput logs and JSON counters at
/metrics(also available at/v1/metrics), including decode graph batch and padding counters and MTP acceptance counters. Startup checks the memory plan before allocation. Cache capacity is configurable, with BF16, FP8 or FP4 latent storage for GLM-5.3. All families size serving logits for decode and MTP verification, projecting only prefill tails. At a 2048-row prefill budget this saves about 301 MiB per rank for GLM-5.3-Flash on four nodes; see the numerical validation.
dgpp's default serving chain computes what the checkpoint specifies, and it computes it the same way every time: a greedy request returns the same tokens run after run, across ranks (every admission is journaled from the head and every tick's operation stream is digest-checked on every peer), and with speculative decoding on or off (greedy MTP produces plain decode's tokens; sampled MTP preserves the target distribution). Performance work holds the cache, rank and speculative-decoding contracts exact. Representation changes such as the packed int4/int8 tensor-core prefill, the 12-bit BF16 weight form, the bit-plane vocabulary head and the n-gram prestage have bitwise regression gates. Accumulation-order changes, including compact BF16 prefill heads, use numerical accuracy checks and matched teacher-forced scoring; benign rounding differences need not reproduce the previous build's bits. Weight storage changes must preserve the source values.
Reduced-precision levers include quantizing activations to FP8 for the tensor
cores, summing expert partials in bf16, and folding scales into the weight
values. dgpp supports those levers, but
never silently. Each is an explicit engine.* key, off by default, that a
deployment turns on in its own configuration; every rank runs the same
setting (the config digest carries it) and the startup log names it. Each
lever's kernel is gated in its test against the exact chain, and its cost
is measured on real prompts — the task evals in benchmarks
and greedy transcript comparisons — before a template may enable it. A
template enables a lever only when that measured quality cost is inside
the run-to-run band of the default chain, and it says so; the
benchmarks list the default and the levers-on numbers as separate rows.
The current keys are engine.prefill_bf16_partials,
engine.prefill_fold_scales and engine.prefill_fp8_gemm (the
configuration table); which levers pay, and
how to measure any change on your own workload, is in
optimizing performance.
See Benchmarks for current serving throughput, cold prefill latency, quality scores and long-context measurements. The document identifies the measured source revision, hardware, workload and timing scope, and links to the raw results and reproduction commands.
As of 2026-10-04, the source tree has fifteen measured deployment templates covering eight model architectures on one, two or four Sparks. The shared engine provides graph decode, transactional MTP, row-batched execution, grouped prefill, prefix caching, deterministic multi-rank scheduling and the OpenAI-compatible service. Current quantized paths cover FP8, NVFP4, MXFP4, full GLM-5.3's packed int4/int8 format and the Qwen3.8 AutoRound int4/int8 hybrid (GPTQ layout, group 128, served as packed). Qwen NVFP4 runs on one or two Sparks by mapping its n-gram table from NVMe and encoding the dense stack to FP8 at load.
The Qwen sixteen-slot decode graphs and prefill continuation are implemented as opt-in controls; the shipped Qwen templates retain four slots and monolithic admission. DeepSeek ships at six slots with DSpark depth 4, confidence scheduling, bounded grouped prefill and the stream-ordered eager collective. Full GLM-5.3 ships at eight slots and uses packed tensor-core prefill from 128 rows. MiMo-V2.6-Flash ships at four slots with MTP depth 1 on two or four Sparks. DeepSeek-V4-Flash (2026-10-02) ships at six slots on four Sparks and four on two, both with the confidence-scheduled DSpark depth (greedy and sampled requests) and the 1M-token context. Qwen3.8-27B (2026-10-03/04) ships the DFlash2 block drafter on one Spark (one block proposal a step, eight request slots, the FP8 head) and MTP depth 3 on two and four (tensor parallel over the DeltaNet and attention heads, the MLP and the head); the drafter runs on every world since 2026-10-04 (the block proposal recorded inside the graph step, the ranks' top-16 lists merged through one fold, the drafter sharded) and is the two- and four-node templates' mode. Version 0.1.0 remains the original GLM-5.3-Flash sign-off release; the current source has advanced beyond that baseline.
PLAN.md summarizes implementation status, CHANGELOG.md records changes, and the remaining work lists current limitations.
Use an internet-connected Spark as rank 0. First get the source:
git clone https://github.com/HawkBearPig/dgpp.git
cd dgppFor an x86 Linux build workstation, use the Docker-based Spark cross-build; run the resulting binaries on a Spark.
Run the guided setup on rank 0:
./scripts/setup.shThe wizard helps you choose a supported model and node count, saves your site
settings, checks SSH and software dependencies on every node, and walks through
RoCE lane selection for multi-node deployments. It then builds the release
server, prepares the downloader, downloads the checkpoint once and syncs peers,
and runs the serving preflight. Existing deployment tuning and unrelated .env
entries are preserved. Reruns reuse the build and complete cached checkpoints.
Python 3.10+ is needed to run setup. On Ubuntu/DGX OS, the wizard can install
standard system packages with sudo; --install-system-deps requests this
up front. NVIDIA drivers/CUDA and physical network setup remain site prerequisites.
Expect substantial checkpoint storage and download time on a fresh machine.
The wizard reports each node's free cache space and explains the lane choices;
it cannot verify cabling or end-to-end RDMA connectivity.
Setup prints the commands to start, inspect and stop the selected deployment.
Add --start to launch after all checks pass:
./scripts/setup.sh --startThe API defaults to localhost and has no authentication or TLS. Wait for READY
before sending requests. For unattended setup, read-only checks, offline cache
sync and the manual walkthrough, see Getting started.
Use ./scripts/setup.sh --help for all options.
Ask the running model a question. This short example requests low reasoning
effort because reasoning tokens also count toward max_tokens:
MODEL=$(curl --fail -s http://127.0.0.1:18080/v1/models | jq -r '.data[0].id')
jq -n --arg model "$MODEL" \
'{model:$model,max_tokens:160,reasoning_effort:"low",messages:[{role:"user",content:"Name three primary colors."}]}' \
| curl --fail http://127.0.0.1:18080/v1/chat/completions \
-H 'Content-Type: application/json' --data-binary @-Stop it when finished, using the same config:
python3 scripts/dgpp-cluster down --config /path/to/the/selected/deployment.jsonA release is a versioned tarball installed once per node. Starting an installed release stages the configuration and uses the installed binary.
CONFIG=/path/to/the/selected/deployment.json # use the path printed by setup
scripts/release.sh # build the release preset, stage, verify, pack
scripts/dgpp-cluster install dist/dgpp-VERSION.tar.zst --config "$CONFIG"
scripts/dgpp-cluster up --release VERSION --config "$CONFIG"
scripts/dgpp-cluster releases --config "$CONFIG"Select a release with the config's release key or --release.
The version is the tree's, stamped at build time by cmake/version.cmake
into the binary (dgpp-serve --version; 0.1.0+g<sha12>, .dirty when
the tree had uncommitted changes) and sent on the journal's settings
record, so a world of mixed versions refuses to form. Inside the tarball:
| path | what |
|---|---|
bin/dgpp-serve |
the server, rpath $ORIGIN/../lib |
lib/libcudart.so.13, lib/libcublasLt.so.13 |
the CUDA runtime it was built against |
scripts/ |
launcher, process/preflight/config helpers, checkpoint downloader and API check |
deploy/*.example.json |
all supported deployment templates |
README.md, docs/ |
package-specific setup, dependency and networking guides |
MANIFEST, MANIFEST.sha256 |
version, git sha, CUDA, build host and date; every other file's checksum |
Beyond the tarball a node needs the NVIDIA driver, rdma-core, libnl and
libstdc++, all part of the DGX OS image, plus the checkpoint and resident
image caches on its local NVMe. install unpacks under paths.release_dir
(~/dgpp/releases/dgpp-<version>/) and checks every file against the
manifest. Versions can be installed side by side; stop the running service
and select an older version to roll back. The launcher starts the service
on demand; no boot-time service units are included. When no release is
selected, up stages build-release/dgpp-serve to the peers. Use
--bin build-ci/dgpp-serve explicitly when deploying a testing build.
Site settings live in .env at the repository root, shared by every model
deployment. Scripts load it automatically through scripts/site_env.py
(shell scripts use scripts/cluster_env.sh). DGPP_ENV_FILE selects another
file; exported variables override file values. Values are literal, optionally
quoted: no shell commands or variable expansion are evaluated. The site helper
loads an allowlist of settings, so credentials such as HF_ACCESS_TOKEN stay out of the
scripts' environment and generated server config.
.env key |
purpose | default |
|---|---|---|
DGPP_NODES |
Space-separated SSH/control hostnames or IPv4 addresses in rank order, using management IPs, fabric IPs, or a mixture. Rank 0 runs locally, must reach peers over SSH, and must be reachable by peers at its listed address. A deployment uses the first world_size entries. See network layouts. |
required |
DGPP_SSH_USER |
Peer login for binary staging, process control, diagnostics and checkpoint sync. Needs SSH-key access and write access to the configured directories. | current login when empty or absent |
DGPP_CLUSTER_CONFIG |
Default deployment filename, saved by guided setup. An explicit --config takes precedence. |
legacy four-node Flash hybrid filename |
DGPP_HTTP_PORT |
Default client-facing API TCP port on rank 0; deployment http.port overrides it. |
18080 |
DGPP_HTTP_BIND |
Default IPv4 listening address on rank 0; deployment http.bind_host overrides it. Keep localhost unless you have arranged access protection. |
127.0.0.1 |
DGPP_FABRIC_PORT |
Rank-0 TCP rendezvous listener used to establish the inter-node transport. Peers must reach it; clients do not use it. Keep it private to the cluster. | 29970 |
DGPP_JOURNAL_PORT |
Rank-0 TCP listener that distributes ordered scheduler operations to peers. Must differ from the fabric/API ports; keep it private to the cluster. | 29971 |
DGPP_LOG_DIR |
Base directory on rank 0 for logs, process records and collected peer logs. The launcher adds a deployment-specific subdirectory. | ~/dgpp/log |
DGPP_STAGE_DIR |
Base directory on peers for staged development binaries, runtime config and logs. The launcher adds a deployment-specific subdirectory; it is not the model cache. | /tmp/bus4 |
DGPP_RELEASE_DIR |
Base directory on each node for versioned installed server releases. Use a persistent, writable location. | ~/dgpp/releases |
DGPP_BUILD_DIR |
Explicit build-directory override. Relative paths use the repository root. Leave unset to keep release and testing builds separate. | build-release for serving/packaging; build-ci for testing (build-<preset> with DGPP_PRESET) |
DGPP_DATA_DIR |
Evaluation-data location. Relative paths use the repository root. | data |
HF_HUB_CACHE, HF_HOME |
Downloaded checkpoints. An explicit hub cache wins; otherwise use HF_HOME/hub. Leave unset for the standard cache or choose a disk with room for the full checkpoint on each node. |
~/.cache/huggingface/hub |
DGPP_RESIDENT_CACHE_DIR |
Per-node disk cache of prepacked weight images for faster reloads, separate from the HF checkpoint. Takes precedence over deployment paths.resident_cache. |
~/.cache/dgpp/resident |
DGPP_ROCE_DEVICES |
One or two ordered local verbs device names, not Linux interface names. Use discover_roce.py; match lane subnets in the same order across nodes. |
discover active Ethernet devices |
DGPP_ROCE_GID_INDICES |
One RoCE-v2 address-table index per explicitly selected device, in the same order. Pin these when a device offers several networks; discover the values on each host. | automatic RoCE-v2 selection |
DGPP_NODE_OVERRIDES |
JSON map keyed by exact DGPP_NODES entries, with per-node device/GID/hub-cache/resident-cache values. Use when peers have different NIC names or disks; see .env.example. |
none |
Model settings live in deployment JSONs under deploy/. Each has a
world_size of 1, 2 or 4 and uses that many nodes from the beginning of
DGPP_NODES. Too few nodes is an error, not a fallback to a smaller world.
Select the deployment explicitly with --config FILE.
Wrappers that accept a config argument pass it through to their children.
Stage and release directories must be absolute or start with ~/, without
spaces or shell syntax.
Log and staging directories are namespaced by deployment-file path;
dgpp-cluster paths prints the effective locations. An explicit --log-dir
is used as-is. up refuses an existing deployment unless --replace is given;
down without --config stops every recorded deployment that is running;
cleanup uses recorded process identity, never a binary-name kill.
The launcher resolves the deployment and site settings into
<log_dir>/cluster.resolved.json when starting the service. Both the head
and peers read that resolved config; the original JSON and .env are not
sent to peers. To inspect or use the runtime config with the native binary:
scripts/dgpp-cluster resolve --config deploy/cluster_glm-5.3-flash_nvfp4-fp8_w4.json
# Save that JSON to a file before passing it to dgpp-serve --config.The native binary reads resolved JSON, not .env or deployment templates.
Unknown keys and invalid types are rejected.
Rank 0 reads the model, ports and engine options, with command-line flags
after --config overriding the file. It sends the effective settings to
peers before model construction. Peers use their own files for bootstrap
addresses and local paths, then verify the shared settings by digest.
The defaults below come from ClusterConfig::Engine; deployment
templates set their serving options explicitly.
| Key | Required | Purpose and when to change it | Default |
|---|---|---|---|
model |
yes | Exact Hugging Face repository ID, such as HawkBearPig/GLM-5.3-Flash-NVFP4-FP8. Selects the checkpoint tensors, tokenizer and chat template. It is not a local filesystem path or a generic quant name. |
— |
world_size |
yes | Number of participating nodes/ranks: 1, 2 or 4. Uses the first N entries of DGPP_NODES; each rank stores its tensor-parallel share in memory. Choose a supported model/world pair, not an arbitrary smaller number to save machines. |
— |
http.bind_host |
no | IPv4 address on rank 0 that accepts API connections. 127.0.0.1 is local-only; a LAN address exposes that interface; 0.0.0.0 exposes all IPv4 interfaces. The server has no authentication/TLS. Does not select the RoCE interface. |
DGPP_HTTP_BIND, otherwise 127.0.0.1 |
http.port |
no | TCP port clients use for the API, in 1–65535. Change it if the default is occupied, then update client URLs. Overrides the site HTTP port and must differ from fabric/journal ports in a multi-node deployment. | DGPP_HTTP_PORT, otherwise 18080 |
http.max_body_bytes |
no | Maximum serialized HTTP request body in bytes, as a positive integer. Allows large document prefills and agent histories; independent of the model's token/KV capacity. Oversized requests receive HTTP 413 based on Content-Length, before tokenization. Buffers grow with received data, so this does not preallocate the limit per connection. The binary flag --http-max-body-bytes overrides it. |
268435456 (256 MiB) |
http.sse_ping_interval |
no | Whole seconds of stream silence before sending an SSE keep-alive comment. Accepts 1–2147483647; -1 disables pings. Covers queued requests, prefill and gaps between output chunks on rank 0. --sse-ping-interval overrides the file; request sse_ping_interval overrides the server setting. Engine deadlines are unchanged. |
30 |
release |
no | Installed software version to run on every node, not a model revision. Select a version under DGPP_RELEASE_DIR/dgpp-VERSION; use it to upgrade or roll back server binaries. --release overrides this field; --bin overrides binary selection. |
Empty: use the development build |
paths.resident_cache |
no | Disk directory for prepacked per-rank weight images, which speed subsequent loads. This is separate from the Hugging Face download cache and consumes additional disk space. Prefer the shared DGPP_RESIDENT_CACHE_DIR site setting unless a deployment needs its own directory. |
Empty: ~/.cache/dgpp/resident; DGPP_RESIDENT_CACHE_DIR takes precedence |
For a first run, retain the template values. Tune kv_capacity for context
space and max_concurrency for simultaneous work; these are different limits.
Startup checks the combined memory plan before loading.
| Key | Required | Purpose and when to change it | Default |
|---|---|---|---|
engine.max_concurrency |
no | Maximum actively executing requests, not TCP connections or queued requests. More slots can improve aggregate throughput but use more state/scratch memory and may increase per-request latency. Allowed range is 1–16, subject to model/MTP row limits below. | 8 |
engine.kv_capacity |
no | Shared context-token pool on each rank, across active requests—not a separate allowance for every request. A prompt and its generated answer must fit. Increase for longer contexts or more simultaneous context; memory use increases and allocation is rounded to model block boundaries. | 8192 tokens |
engine.kv_dtype |
no | GLM-5.3 latent-cache precision: bf16, fp8, or fp4; the MiMo-V2.6-Flash K/V cache takes bf16 or fp8 (e4m3 rows with one scale per head row, half the bytes). Lower precision reduces cache storage at a numerical-accuracy cost; it does not quantize model weights. The GLM index cache stays FP8. Qwen, GLM-4.7 and DeepSeek K/V caches remain BF16. |
bf16 |
engine.embed_sharding |
no | Full GLM-5.3 and DeepSeek embedding/head placement: replicated keeps the full table on every rank; vocab keeps each rank's vocabulary slice and folds token lookups. The full-GLM template uses vocab to save 1.33 GiB/rank; the DeepSeek template retains replicated. Other families ignore it. |
replicated |
engine.default_max_tokens |
no | Answer-token budget for requests that omit max_tokens. Clients may supply their own value; this is not a global hard limit. A larger default also reserves more context space under full admission (kv_capacity must cover the prompt plus this budget, or the request is refused). Every deployment template sets 32768: agent clients such as Hermes send no max_tokens, and a thinking model's reasoning alone exceeds a few hundred tokens. |
256 |
engine.queue_limit |
no | Maximum requests waiting for an execution slot or memory budget. Additional arrivals receive HTTP 503 overloaded. Increase to tolerate bursts, at the cost of longer waits—not higher execution capacity. |
64 |
engine.max_connections |
no | Maximum simultaneously open HTTP connections on rank 0, including idle keep-alive connections and streams. Excess connections receive 503. Size this separately from active request slots. | 64 |
engine.prefix_cache_gib |
no | Memory budget per rank for reusable prefix-state snapshots. Long documents can reuse an earlier snapshot when their question changes. Snapshot slots and the KV token pool are separate limits; see sizing and recipe capacities. Set 0 to disable. | 1.5 GiB |
engine.admission |
no | When to reserve context space. full reserves prompt plus the requested answer budget before admitting a request. grow starts with a smaller reservation and extends it during generation; if space runs out, the youngest request is shed. Use full for predictable reservations, grow to trade that guarantee for denser occupancy. |
full |
engine.admission_window |
no | Answer-token reservation increment used by grow admission. Larger increments reduce growth frequency but reserve more space ahead of use. Has no effect under full. Must be positive. |
256 tokens |
engine.prefill_budget_tokens |
no | Maximum prefill tokens per scheduler tick. -1 selects an aligned budget near 256 on supported Qwen, GLM-5.3-Flash and DeepSeek-V4-Flash graph engines, with cancellation and decode between chunks. Positive values override it; 0 explicitly keeps full-prompt admission. Other engines retain full-prompt admission. DeepSeek-V4-Flash reads every in-flight prompt's next chunk in one forward: there the budget is a quantum per reading prompt (up to the forward's 4096 rows per tick), prompts queued together begin together, and near the end the shares level so they finish — and start decoding — together. Images on an engine without image chunking run monolithically as the tick's only prefill work when they exceed the budget. | -1 (automatic) |
engine.prefill_idle_budget_tokens |
no | Larger prefill budget when no request is actively decoding. Requires an enabled busy budget, must be at least that budget, aligned and within the same prefill limit. Rechecked after each chunk. 0 uses the busy budget for all chunks — except on DeepSeek-V4-Flash under the automatic busy budget, where it defaults to the whole forward (4096). | 0 (disabled) |
engine.admission_gather_ms |
no | How long rank 0 holds the first arrival at an idle engine for the rest of its burst. Requests sent together land a few milliseconds apart; the wait lets them be read in together instead of the first starting one tick ahead. Applies only when nothing is queued or running. 0 ticks at once. | 3 ms |
These settings change how the model runs. Use the matching deployment template
first; change one setting at a time and measure the effect on your workload
(optimizing performance has the protocol).
Every key here keeps the model's outputs exactly except the three
engine.prefill_* levers, which trade prefill arithmetic for speed and are
off unless a deployment turns them on (see
accuracy and correctness). Engine behavior is
set here, in the deployment file; an environment variable never selects a
kernel, a lever or a threshold.
| Key | Required | Purpose and when to change it | Default |
|---|---|---|---|
engine.decode_graph |
no | Use CUDA graph replay for decode to reduce CPU launch overhead. On one node, this selects resident graph serving instead of the eager streaming path. Required for MTP and the single-Spark Qwen templates; leave enabled for the documented serving configurations. | false |
engine.mtp |
no | Enable multi-token prediction: draft candidate tokens, then verify them with the main model. Can reduce decode time when drafts are accepted, but adds draft state and verification work. Requires decode_graph; not a larger request batch. |
false |
engine.mtp_depth |
no | How many tokens to draft per speculative step, 1–5. Greater depth can accept more tokens per pass, but uses more verification rows and can waste work when drafts are rejected. Values above 1 require MTP; supported batching varies by model. DeepSeek's shipped template uses depth 4. | 1 |
engine.mtp_schedule |
no | Enable confidence-scheduled verify depth for a model with a confidence head. Requests verify only the leading drafts whose survival probability justifies another row: greedy requests always, sampled requests when their drafts are the draft head's most likely tokens (engine.mtp_draft greedy; with drawn drafts they keep the configured depth). The depth changes pace only — a greedy transcript and a sampled request's distribution are the same at every depth. Unsupported for Qwen C16/MTP3: leave this false at sixteen slots and depth 3 because additional verification depths exceed the graph-variant limit. |
false |
engine.mtp_schedule_row_ms, engine.mtp_schedule_base_ms |
no | Cost model for scheduled verification: milliseconds for another verify row and fixed work per pass. These values are deployment measurements; retain the DeepSeek template values unless re-profiling that world. | 8.0 / 28.0 |
engine.mtp_schedule_lambda, engine.mtp_schedule_min_depth, engine.mtp_schedule_adapt |
no | Floor for the value of decode time in tokens/ms, minimum verified draft depth, and whether the value adapts from committed tokens and modeled time. Lambda 0 derives the reservation rate. | 0 / 1 / true |
engine.mtp_schedule_sampled_scale |
no | Under engine.mtp_schedule, a sampled request's expected acceptance per drafted position as a fraction of the confidence head's. The head predicts a greedy match; a sampled request accepts a draft with the target's probability of it. Traced on DeepSeek-V4-Flash with the whole block verified every pass (3,140 prose passes, 819 over the five prompt classes), the head over-states a sampled request's first draft (0.75 predicted against 0.67 accepted on prose) and is near calibrated behind it; replayed over those passes, 0.93 commits 1.5–4 % more tokens per second than 0.8 and within 0.2 % of the best first-position-only correction. 0 makes sampled requests verify the whole block. Applies only with engine.mtp_draft greedy. |
0.93 |
engine.mtp_draft |
no | How a sampled request's speculative drafts are chosen: sampled draws them from the draft head's own distribution and verifies with the acceptance ratio; greedy takes the draft head's most likely token and accepts it with the target's probability of that token. Both preserve the target distribution exactly; they differ in tokens accepted per pass. auto uses the rule measured better for the model family (greedy for DeepSeek-V4-Flash's DSpark block, sampled elsewhere). Greedy requests are unaffected. |
auto |
engine.prefill |
no | DeepSeek prefill mode: bounded runs every prompt row through the encoder and only the final window through the decoder; exact runs every layer over every row for parity work. Other families ignore it. |
bounded |
engine.ngram_table |
no | Qwen n-gram embedding-table placement. resident keeps it in device-accessible memory; mmap leaves it on local NVMe and fetches needed rows through the host page cache. The Qwen NVFP4 templates use mmap; the two-Spark campaign measured at most 3.6% lower decode throughput for 23.84 GiB less planned model memory per rank. DeepSeek's Engram tables use their own mapped checkpoint sidecar. |
resident |
engine.dense_weights |
no | Qwen dense-projection storage. checkpoint retains the checkpoint's BF16 form; fp8 converts dense projections at load time to reduce their memory footprint, with quantization error. Does not select another HF repository or change the expert quant; other families ignore it. On Qwen3.8-27B, whose projections ship as FP8, fp8 requantizes the BF16 lm head to block FP8 (half its bytes per decode row: MTP depth-2 C1 151 → 132 ms/step at the same acceptance, HumanEval/GSM8K 39/40 each; greedy transcripts move within the first tokens). The Qwen3.8-27B templates ship it; checkpoint restores the BF16 head. |
checkpoint |
engine.fp8_head |
no | Qwen FP8 vocabulary-head dispatch. mma uses streaming MMA above the dense GEMV threshold and within the configured decode capacity, including short prefills in that interval. Requires dense_weights: "fp8"; the NVFP4 templates select it after matched numerical validation. gemv restores the previous dispatch. |
gemv (NVFP4 templates: mma) |
engine.bf16_weights |
no | Resident form of the BF16 weights that decode streams (attention/linear-attention projections, the LM head; on Qwen3.8-27B the DFlash2 drafter's layers and fc taps, the MTP fc and a BF16 head, whose BF16 bytes stay resident under either packed mode). checkpoint keeps the BF16 bytes alone. bf12 and bf12+bf16 add a lossless 12-bit form of each matrix (sign+mantissa byte plus a 4-bit exponent code; exact side tables for the rare outliers): decode launches of up to eight rows read 0.75 of the bytes and produce bit-identical results, so transcripts do not change. This is a storage format, not quantization. bf12+bf16 keeps both forms resident: prefill is untouched and the 12-bit copies cost about 0.75× those matrices in additional memory (GLM-5.3-Flash: +2.0 GiB/rank at four Sparks, +3.8 at two; GLM-4.7: +4.9). bf12 keeps the 12-bit form alone: each matrix's BF16 bytes are returned to the node as its layer loads, the footprint drops below checkpoint's (GLM-5.3-Flash −0.5 GiB/rank at four Sparks, −1.0 at two; GLM-4.7 −1.4), and a prefill GEMM expands the rows it reads into a small scratch first — the same bits, so the same results, for about 10 ms per prefill chunk on four-node GLM-5.3-Flash (+1–2% on 8K–32K prompts, +25–40 ms to a short prompt's first token). The memory plan includes either. Measured single-stream decode: GLM-5.3-Flash +6–7%, GLM-4.7 +11–12%, full GLM-5.3 +5–6.5%, Qwen3.8-Flash-Next-FP8 +6.4% on four Sparks and +9.0% on two (benchmarks). On Qwen it packs the GDN/QSA projections and the head of the FP8 checkpoint and always keeps both forms resident (every Qwen recipe has the room); under dense_weights: "fp8" those matrices are already FP8 and nothing is packed. DeepSeek accepts the key and currently packs nothing. Templates with memory to spare ship bf12+bf16; the two sized to their nodes' memory (two-node GLM-5.3-Flash, full GLM-5.3) ship bf12. |
checkpoint |
engine.draft_vocab |
no | Qwen3.8 AutoRound hybrid only: path to a one-dimensional int32/int64 .npy of token ids. The MTP draft head then scores only those rows of the head (every other id is -inf for the draft), so each draft step reads that slice instead of the whole head. The target verifies every draft, so outputs are unchanged; only the draft acceptance can move (a token outside the set is never proposed). tools/build_draft_vocab.py TOKENIZER.json OUT.npy --size 65536 builds a set from the tokenizer's alphabet, added tokens and most frequent merges. Off by default; headline numbers are measured without it. |
empty (whole vocabulary) |
engine.prefill_bf16_partials |
no | Qwen3.8 AutoRound hybrid prefill lever: the packed expert chain's down projection is written in bf16 and its per-expert partials summed from bf16 (half the bytes a prefill chunk writes and reads back; the reference stack's form). Not bitwise the default fp32 chain: served transcripts can differ within the quantized model's tolerance. Off by default; the AutoRound template turns it on (−3 to −7 % cold prefill, evals inside the default chain's band; benchmarks list both rows). | false |
engine.prefill_fold_scales |
no | Qwen3.8 AutoRound hybrid prefill lever: the wide packed expert GEMM folds each group's scale into the bf16 weight values and keeps one fp32 accumulator across K (Marlin's form; no per-group fma). Not bitwise the default chain. Off by default and measured as no gain (+3 % at 32K): no template turns it on. | false |
engine.prefill_fp8_gemm |
no | Qwen prefill lever under engine.dense_weights: "fp8": prefill-shaped dense projections run on the fp8 tensor cores from per-token 1×128 e4m3 activations with the checkpoint's 128×128 weight scales (the reference stack's blockwise GEMM) instead of dequantizing each matrix to bf16 for cuBLASLt. Not bitwise the dequantized chain (the activations are quantized). Off by default; requires dense_weights fp8; the AutoRound template turns it on beside prefill_bf16_partials (0 to −3 % alone). |
false |
engine.prefill_fp8_per_tensor |
no | Qwen3.8-27B prefill lever: every FP8 projection's prefill GEMM (rows above the decode GEMV band, and every resumed chunk) runs on cuBLASLt's per-tensor-scale e4m3 kernels from boot-requantized per-tensor weights and per-call per-tensor activations — about 2x the dequantized bf16 GEMM's rate (0.59 vs 1.04 ms/token at 2K), +23 GiB resident at the 27B's shape. Lossy beyond the checkpoint: an 8K-prompt greedy completion differs from the exact path. Opt-in; the default prefill dequantizes each matrix to bf16 for cuBLASLt (exact). | false |
engine.dflash_model |
no | Qwen3.8-27B: a DFlash2 block-drafter checkpoint (z-lab/Qwen3.8-27B-DFlash2, an HF id or a directory) that replaces the MTP draft (mtp off). Five bidirectional draft layers fed by the target's layer taps propose seven tokens per step through the rank-256 selector walk. On one Spark without decode_graph it runs on the eager engine (the batched verify graph, the stacked redrafts); on the graph worlds — the fabric (decode_graph required) or one Spark with the decode graph — the block proposal is recorded inside the graph step on every rank: the drafter's heads and MLP rows sharded across the ranks with two folds a layer, each rank's vocab-slice top-16 merged through one boundary fold so every rank walks the same proposal, the drafts fed to the next replay on the device. The verify is exact: greedy transcripts equal the same world's MTP depth-5 world's (the same verify dispatch class; 4/4 at worlds 1, 2 and 4); sampled requests accept each draft with its exact probability under the request's temperature, top-k / top-p and penalties (the point-mass rule). Requests with logprobs, a logit bias or a grammar run plain. Past 4 slots the 8-row blocks exceed the 32-row decode batch: the engine batches the slots that fit and replays scalar graphs for a live set beyond them. |
none |
engine.dflash_verify_graph |
no | With a drafter: replay the multi-slot verify batch as a captured CUDA graph (one static 16- or 32-row replay; a lone slot stays on the scalar path). The measured best; false runs the packed eager batch. |
true |
engine.dflash_draft_batch |
no | With a drafter: one stacked block forward redrafts every speculating slot (the draft weights read once per step); false redrafts slot by slot. |
true |
engine.dflash_depth |
no | With a drafter: verify only the first N drafts per step (1–7); unverified drafts are re-drafted next step, so transcripts are exact at any value. 0 verifies the whole block. | 0 |
engine.expert_gemm |
no | Packed int4 expert GEMM form for prefill (Qwen3.8 AutoRound hybrid, full GLM-5.3): wide (the 64×128 tensor-core tile, eight warps), wide3 (register decode, three stages), wide4 / wide4r (the four-warp forms), narrow (the 32×64 kernel). All bitwise; wide is the measured best. |
wide |
engine.expert_gemm_prefetch |
no | How many k-steps ahead the expert GEMM prefetches its weight and activation lines into L2, 0–16 (0 disables). Bitwise. | 3 |
engine.expert_tile_list |
no | Launch the expert GEMM over a compact list of the routed segments' tiles instead of a grid over the longest segment. Bitwise; off only for diagnosis. | true |
engine.expert_gemm_pair |
no | Run the expert gate and up projections as one launch. Bitwise; measured level, off by default. | false |
engine.ngram_prestage |
no | Qwen: gather the next prefill chunk's n-gram table rows while the current chunk runs (with ngram_table: "mmap"). Bitwise; off keeps the one-channel staging. |
true |
engine.graph_batch_min_live |
no | Active-request count at which decode switches from scalar to batched graphs. A lower threshold starts batching earlier; batching may improve throughput while doing extra padded-row work. 0 chooses min(2, max_concurrency); explicit values must be 1 through max_concurrency. |
0 (automatic) |
engine.sampling_candidates |
no | Number of candidate tokens gathered per rank on the sampled-token fast path, 1–256. Smaller values reduce routine work but may trigger more full-gather fallbacks. The fallback preserves sampling correctness; this is not the client's top_k parameter. |
128 |
engine.bulk_pace_gbps |
no | Sender pacing rate per queue pair for bulk/prefill communication, in gigabits per second. Negative derives the rate from link speed; 0 disables pacing. Override only when measuring network contention—this is not an API throughput limit. | -1 (automatic) |
engine.bulk_inflight |
no | Maximum in-flight bulk stripes per lane. More can keep the link busy but increase pressure on buffers and competing traffic. -1 uses the transport default (4); normally leave automatic. | -1 (automatic) |
engine.rendezvous_timeout_ms |
no | How long ranks may take to join the transport rendezvous. Increase for slow cold starts or delayed peers; it does not increase an HTTP request's timeout. | 120000 ms |
engine.stats_interval_s |
no | Interval between periodic server throughput/state log lines. Shorter intervals give finer operational visibility and more log output. Set 0 to disable periodic statistics. | 10 seconds |
engine.reasoning_in_content |
no | Put reasoning text in the response's content, separated by the model's </think> marker, instead of a separate reasoning_content field. Use only for clients that need that combined format; it does not disable reasoning. |
false |
engine.no_eos |
no | Ignore the model's end-of-sequence token so measurement runs continue to their token budget. Leave false for normal serving, where a model should be allowed to finish its answer. | false |
The application allows up to sixteen request slots. A speculative request uses
1 + mtp_depth physical decode rows. GLM-5.3-Flash supports eight batched
rows; full GLM-5.3 supports sixteen; GLM-4.7 and DeepSeek support
thirty-two; Qwen supports sixty-four (sixteen requests at MTP depth 3). Graph families capture the slot prefixes that fit, with scalar
fallback for an unsupported active shape. The startup log reports the selected
capacity and rejects a configuration that cannot fit its required rows.
Qwen's 17-64-token decode walks use kernel-only BF16 lowering, including
BF16 projections retained by FP8-dense storage. Smaller walks retain their
existing dispatch even when MTP expands a token into multiple matrix rows.
The shared 8/12-slot graph
families also affect other models above eight slots and can reduce the number
of scheduled verification depths that fit; see operations.
Sampling defaults come from the checkpoint's generation_config.json
and per-request fields. Flags such as --temperature override them for
measurement runs. See operations for memory planning,
MTP depth tradeoffs and deployment checks.
Model implementations share the engine, service, scheduler and text stack:
| directory | namespace | what |
|---|---|---|
src/serve/ |
dgpp::serve |
the HTTP server, the OpenAI-compatible service, the admission journal, the throughput line, the text frontend interface |
src/sched/ |
dgpp::sched |
the scheduler (admission, the tick, cancel and stop), the prefix cache's index, the SchedulerEngine interface every model implements |
src/sample/ |
dgpp::sample |
the sampler's host math: penalties, the logit bias, masks, top-k/p, the exact prefix decision |
src/text/ |
dgpp::text |
the tokenizer (HF tokenizer.json), the chat template (Jinja), the tool grammar and parser, the JSON-schema grammar |
src/kernels/ |
dgpp |
CUDA kernels; pick.cu and sample_pick.cu are the device pick and sampler any model's logits can use, the rest carry their model's name |
src/net/ |
dgpp::net |
the RoCE collective bus, TCP, the roster |
src/engine/ |
dgpp |
shared session interfaces, eager and graph engines, speculative decoding, memory plans and prefix arenas |
src/models/qwen/, src/models/glm4/ |
dgpp |
Qwen3.8-Flash-Next and GLM-4.7 configuration, loaders, layers and sessions |
src/models/glm/ |
dgpp |
GLM-5.3-Flash: the forward pass, loader, MoE and mHC layers, session adapters and MTP draft |
src/models/glm_dsa/ |
dgpp |
full GLM-5.3 configuration, packed loader, MLA/DSA model and session implementation |
src/models/dsv41/ |
dgpp |
DeepSeek-V4.1-Flash configuration, MXFP4/FP8 loader, CSA2, Engram and DSpark session implementation |
src/models/dsv4/ |
dgpp |
DeepSeek-V4-Flash configuration, MXFP4/FP8 loader, sliding-window and compressed-cache attention with the 64-head indexer, token-table routing, two-pass hyper-connections and the DSpark session |
src/models/ |
dgpp |
the KDA and DSA layer libraries and the quantized matrix, shared by any model that uses them |
src/loaders/, src/core/, src/common/ |
dgpp |
safetensors, the HF cache, JSON; arenas and tracing; logging and process memory |
The server binary is dgpp-serve (apps/dgpp_serve.cpp); the GLM
tools keep their names (glm_gen_check, glm_forward_check, …).
tools/ holds the Python checkpoint and reference tooling the tests use;
scripts/ the fabric and serving operations; apps/ the binaries;
benchmarks/ the probes and the engineering record; deploy/ the cluster
config.
| page | what |
|---|---|
docs/operations.md |
booting, stopping and watching the serving world; memory and the image cache on a node; failure semantics |
docs/benchmarks.md |
every measured serving number: per model, world, concurrency and prompt class, with the method to reproduce each |
docs/signoff_v1.md |
the v1 performance and hardening sign-off, with every measurement |
docs/next_steps.md |
what is worth doing next, ranked by cost and benefit |
docs/tools.md |
serving, testing and diagnostic commands |
docs/testing.md |
the test suites and what each proves |
docs/numerics.md |
judging a numerics change; the sampling-width gate |
docs/optimizing_performance.md |
the levers that trade memory, latency, throughput and accuracy, which ones pay on which model, and how to measure a change before trusting it |
docs/mtp.md |
speculative decode with the MTP layer |
docs/measurements.md |
current validated platform, network, kernel and serving measurements, with historical observations separated |
docs/checkpoint_budget.md, docs/qwen38_checkpoint_budget.md, docs/checkpoint_budget_glm53.md, docs/checkpoint_budget_dsv41.md |
generated checkpoint inventories, per-rank residency and decode traffic budgets |
docs/batched_mtp_graph_stall.md |
a worked stall investigation, from symptom to root cause |
docs/qwen38_flash_next_plan.md |
Qwen3.8-Flash-Next's architecture, placement, decisions and status |
docs/glm47_plan.md |
GLM-4.7 NVFP4 architecture, modelopt format contract, decisions, gates and status |
docs/glm53_plan.md |
full GLM-5.3's packed checkpoint, DSA/MLA changes, placement and serving status |
docs/deepseek_v41_flash_plan.md |
DeepSeek-V4.1-Flash's CED, CSA2, Engram, DSpark and serving implementation |
docs/performance_improvement_plan.md |
current performance findings, completed changes and next measured targets |
DESIGN.md |
architecture and implementation contracts |
PLAN.md |
the milestones, their exit gates and status |
benchmarks/README.md, benchmarks/results/ |
benchmark probes, dated results and reproduction commands |
CHANGELOG.md |
the history by milestone |
CONTRIBUTING.md describes the build presets, the test discipline, the
evidence a fabric change needs, the code style and where things go.
Apache License 2.0. See LICENSE.