[Windows 11] AMD Adrenalin driver (32.0.31007.5012) causes R9700 VRAM to offload to system RAM on idle - crash with 32GB RAM #23443
Replies: 15 comments 12 replies
|
Same issue with that version of Adrenalin including up to current version 26.6.4 (32.0.31021.5001) with a 7900xtx. 'Problem' and 'Fix' sections are absolutely correct. |
|
Can confirm too. Rolling back to an earlier version driver is the only fix. |
|
I agree. Need to download driver at link: https://www.amd.com/en/support/downloads/previous-drivers.html/graphics/radeon-ai-pro/radeon-ai-pro-r9000-series/amd-radeon-ai-pro-r9700.html specifically on version: 2026-03-19 |
|
Just brought the R9700 today, tried in Window LMStudio with the latest driver download from Adrenaline. I got the same issue. VRAM was offloaded completely to 0 and my 32GB ram was exploding and caused the system crash. Thank you all for clearly described the issue here and provided a solution of using a rolled backed driver. |
|
Confirming this on a Windows 11 + Radeon AI PRO R9700 (32GB) setup, running Unsloth Studio's bundled llama-server (Vulkan backend, -ngl 99, Vulkan1 device) for chat inference. Driver: 32.0.31035.1003 (AMD Software: Adrenalin Edition 26.7.1, the current latest WHQL release as of 2026-08-18) — bug is still present. The latest driver does not fix this. Three full-system bugchecks within ~2 hours of real usage today, all flagging the same driver:
All three System event log entries name amdkmdag.sys as the likely-related driver. The idle-adjacent timing (crash arriving shortly after load ends, not during) matches the VRAM-offload-on-idle behavior described in this thread rather than a mid-computation hang. Checked Windows' TDR watchdog settings — both TdrDelay/TdrLevel were unset (Windows defaults, 2s timeout). Increased TdrDelay to 10 as a partial mitigation; hasn't been long enough to confirm whether it helps, but doesn't address the underlying idle-VRAM-eviction mechanism this thread describes. |
|
New data point for this thread, from the same 2x R9700 / Windows 11 box as my earlier report (driver matrix: 31007/31019/31035 evict, 23027/22042 don't). The 31xxx eviction is gated on display power state, not idle time. With a dummy plug + live monitor, 31035 kept ~53GB resident through 9+ hours of idle — looked fixed. The moment Windows turned the displays off (30-min timer), the driver evicted the dummy-plugged card's entire 24.8GB in ~6s (~4GB/s, caught on a 1-second fsync'd sampler), started on the second card, hung mid-eviction, failed TDR recovery twice (LiveKernelEvent 141) and bugchecked with 0x116 VIDEO_TDR_FAILURE / amdkmdag.sys. Three nights in a row, then reproduced on demand with a SC_MONITORPOWER broadcast: screen blank to BSOD in ~27 seconds. So: dummy plugs only mask the idle path while screens are on; they do not prevent the display-off eviction, and there seems to be no registry switch for it (consistent with the attempts reported above). On 32.0.22042.14002 the same artificial display-off test shows zero eviction (VRAM immovable 7+ min after blank). That rollback remains the only real fix I'm aware of. (Posting as a data point; I likely won't be able to follow up.) |
|
As a workaround to this, I've just been keeping a tiny workload constantly on the GPU, like so: |
|
Solved (for now) by switching driver branch: AMD Software: PRO Edition 26.Q3 fixes the idle VRAM eviction on the R9700. Short version: the eviction is tied to the Adrenalin driver branch. Moving from Adrenalin to the dedicated PRO Edition driver stops it. On my machine the model now stays fully resident in VRAM even after long idle periods, on both Vulkan and ROCm. What was happeningWith Adrenalin builds from 26.5.1 onward, a fully VRAM-resident model got evicted to system RAM after ~15s idle (WDDM eviction to the allocation backing store). On a 32GB-RAM box this fills RAM and the next request pays a full PCIe re-upload. Adrenalin 26.3.1 (32.0.22042.14002) was the last clean build. The fixAMD Software: PRO Edition 26.Q3 (released 18 Aug 2026) added support for the R9700 SKU. Switching to it resolves the eviction. Steps:
Test result
Notes / caveats
Hopefully this saves others the months of driver roulette. Would still be good to see AMD fix the Adrenalin branch itself, since plenty of R9700 users run it. |
|
Same issue here with 2x r9700s. Went through all troubleshooting steps over the last month. Sleep settings, purchased a dummy plug for the second GPU, etc. I have rolled back and I am monitoring currently. |
|
Update on my earlier "solved by PRO Edition 26.Q3" post — I need to correct the outcome. To be clear about the history: the first PRO Edition driver I installed genuinely worked. The card ran fully loaded — ~90% of the 32GB VRAM in use — with no eviction, exactly as it should. This was never a light-load fluke. It only broke again after a later driver update. So the PRO branch was fixed at one point, and a subsequent update reintroduced the regression. That's the key point: this isn't the card's limit or my setup — AMD had a working driver and a newer build broke it again. Current state: model loads fully into VRAM, then after ~10s idle the driver evicts the whole allocation to system RAM, and the next request crashes — same eviction as on Adrenalin. My current (broken) driver is 32.0.31036.15, driver date 2026-08-12. Worth stressing: on my box the R9700 is a pure compute card with no display attached (video output runs off a separate GPU), so this has nothing to do with monitor/display-off power states — it's straight idle eviction on a headless compute card, triggering in ~10s. Reliable fix remains rolling back to 32.0.22042.14002 (AMD Software 26.3.1) — the build the card ran fine on under full load. I've reopened my case with AMD support and pushed for escalation to the driver team, with the full writeup (WDDM eviction, regression window, the registry tweaks that don't work, and this thread as reference), and asked for a fix timeline on both Adrenalin and PRO. Will post back if I get anything substantive. |
|
AMD has root-caused this on the tracker, ROCm/TheRock#7221. From nkulshre-amd, 2026-09-10:
That fits the ETW traces I posted there (ROCm/TheRock#7221 (comment)): after ~10 s with no GPU work, Windows suspends the headless card to D3, and since the driver does not declare which memory segments survive D3, dxgkrnl purges the whole VRAM first. It also explains why none of the registry tweaks discussed here helped, and why no build before 26.9.2 is expected to fix it: I tried PRO 26.9.1 (32.0.31041.3013) this week and it still evicts. Until 26.9.2 ships, rolling back to 32.0.22042.14002 remains the workaround most of us are using. I'll report back here once I've tested 26.9.2 on my R9700. |
|
I ran into this bug with my dual R9700 setup, here is a workaround for people with two R9700 or more until AMD decides to fix their drivers :/ Dual R9700 on current Adrenalin — a working interim fix, no rollback neededRollback to The dual-card gotcha nobody's mentioned yetA keepalive on one card does not protect the other. On a two-card box, the card you aren't A spare card also makes the bug fire sooner, because it guarantees an idle GPU always exists. Fix: use a layer split. Every inference then runs a full forward pass through all layers, so Check The keepaliveEviction fires at ~8.7 s idle on my system (#7221 measured ~9.5 s). Poke every 5 s for margin: while true; do
curl -s -m 120 http://127.0.0.1:8080/completion \
-H 'Content-Type: application/json' \
-d '{"prompt":"x","n_predict":1,"cache_prompt":false}' >/dev/null 2>&1
sleep 5
done &PowerShell equivalent: Start-Job { while ($true) {
try { Invoke-RestMethod -Uri 'http://127.0.0.1:8080/completion' -Method Post -TimeoutSec 120 `
-ContentType 'application/json' `
-Body '{"prompt":"x","n_predict":1,"cache_prompt":false}' | Out-Null } catch {}
Start-Sleep -Seconds 5 } }
Cost~41 ms per poke — a 0.8 % duty cycle. Measured over a 50-minute run: 28.41 GB held resident Dual-card layer split: 34.20 GB resident (19.03 + 15.17), both cards healthy through 341 s With Caveats
|
|
Same issue here on 2 different setup: W7800 48gb and R7900 xtx 24gb |
|
Hmm i downloaded the newest pro driver yesterday and for now it is working again .... for now XD |


Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Problem
AMD Radeon AI PRO R9700 on Windows offloads entire VRAM content
to system RAM after ~15 seconds idle, causing system crash
(32GB RAM fills up instantly).
Reproduces with both Vulkan and ROCm backends.
Also reproduces in LM Studio (Vulkan backend).
Root Cause
AMD Adrenalin Software (driver 32.0.31007.5012, dated 12.05.2026)
installs aggressive DPM power management that evicts VRAM on idle.
The basic Windows driver without Adrenalin does NOT have this behavior.
System
Fix
Working command
.\llama-server.exe --device Vulkan1
--model "model.Q4_K_M.gguf"
--ctx-size 65536
--threads 8
--parallel 1
--n-gpu-layers -1
--host 127.0.0.1
--port 8080
--cache-type-k q4_0
--cache-type-v q4_0
--cache-ram 0
--no-mmap
What does NOT fix this
Additional Notes
The R9700 is marketed as a dedicated AI workstation card.
The target use case is exactly keeping LLMs permanently loaded
in VRAM for inference workloads.
Ironically, AMD ships the same consumer Adrenalin driver with
aggressive laptop-style power management on this workstation card.
When the LLM is idle between requests, the GPU already draws only
~8W — there is absolutely no reason to evict VRAM contents to
system RAM at that point.
This DPM behavior makes sense on laptops/integrated GPUs to save
battery. On a dedicated 32GB AI accelerator it actively destroys
the primary use case.
AMD should either:
Related Issues
All reactions