Skip to content

Vulkan: PTQ1_0 kernel hangs GPU (device lost, xe driver job timeout) on Intel Arc/Battlemage at longer prompts #192

Description

@flatlinebb

Summary

The PTQ1_0 Vulkan kernel hangs and takes down the GPU on Intel Arc/Battlemage once the prompt gets large enough (~1900+ tokens). Short prompts (tens of tokens) work fine; the kernel driver logs a compute-job timeout and resets the GPU, which surfaces to llama.cpp as vk::Queue::submit: ErrorDeviceLost.

Given #185 says the PTQ1_0 Vulkan kernel was only ever validated on an AMD RDNA3.5 iGPU (RADV) and "committed as untested" elsewhere, I suspect Intel/ANV was simply never exercised — filing this as the first Intel Arc data point.

Environment

  • GPU: Intel Arc B580 (Vulkan reports Intel(R) Graphics (BMG G21), deviceID 0xe20b, Battlemage)
  • Driver stack: Mesa mesa-vulkan-drivers 25.0.7-2+deb13u1 (ANV), kernel xe DRM driver
  • Host kernel: 7.0.14-16-pve (Debian, Proxmox LXC w/ GPU passthrough via /dev/dri)
  • llama.cpp fork commit: 1a07bfa5f4144274c8f1c9963821dd9d9a51854b (branch prism, 2026-09-17)
  • Build: cmake -B build-vulkan -DGGML_VULKAN=ON -DGGML_NATIVE=ON -DCMAKE_BUILD_TYPE=Release
  • Model: Ternary-Bonsai-2-27B-Abliterated-PTQ1_0.gguf (BoldingBuilds' abliterated build of PrismML's own PTQ1_0 pack)

Repro

./build-vulkan/bin/llama-server -m Ternary-Bonsai-2-27B-Abliterated-PTQ1_0.gguf \
  --host 127.0.0.1 --port 5800 --jinja --reasoning off -ngl 999 --flash-attn on \
  -ctk q8_0 -ctv q8_0 --parallel 1 --threads 12 --threads-batch 16 --ctx-size 32768

A short chat completion (~20 tokens prompt) succeeds normally. A prompt of ~1900+ tokens reliably fails on the first prompt-processing batch:

1.21.972.052 E ggml_vulkan: device lost on Vulkan0
1.21.972.644 E srv  update_slots: decode() failed: vk::Queue::submit: ErrorDeviceLost

Retrying after the failure sometimes instead hits an assertion during recovery:

/opt/llama-cpp-prismml/tools/server/server-context.cpp:2730: GGML_ASSERT(batch.slot_batched || batch.size() == 0) failed

Host kernel log (dmesg) confirms it's a real GPU hang, not just a driver-side error surface:

xe 0000:05:00.0: [drm] Tile0: GT0: Timedout job: seqno=692, lrc_seqno=692, guc_id=2, flags=0x0 in llama-server [pid]
xe 0000:05:00.0: [drm] Xe device coredump has been created

The devcoredump identifies the failing engine as GT0 (main), GuC firmware xe/bmg_guc_70.bin version 70.72.1 (driver notes it wanted 70.54.0 — flagging in case that mismatch is relevant, though I have no evidence either way).

Notes

  • Other (non-ternary) models on the same GPU/driver/kernel, built from a different llama.cpp fork, run correctly with long prompts — this appears isolated to the PTQ1_0 ternary Vulkan kernel path, not a general Arc/Xe/ANV instability on this box.
  • Short prompts never reproduce it in my testing; it seems tied to batch/prompt size rather than being immediate.
  • Happy to pull more from the devcoredump or test a patch if useful — just say what's needed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions