Reduce EnsembleGPUKernel CPU host overhead - #560
Merged
ChrisRackauckas merged 10 commits intoOct 4, 2026
Merged
ChrisRackauckas merged 10 commits into
ChrisRackauckas merged 10 commits into
Conversation
Store isbits CPU problems once and construct default ensemble solutions on access. Keep eager output for custom functions and differentiated solves. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Codex <noreply@openai.com> Agent-Harness: Codex CLI 0.157.1 Agent-Model: gpt-6-sol Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-codex/log.txt on amdci2.julia.csail.mit.edu
Inline the shared solution helper so Enzyme retains concrete type information in the differentiated ensemble output comprehension. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Codex <noreply@openai.com> Agent-Harness: Codex CLI 0.157.1 Agent-Model: gpt-6-sol Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-codex/log.txt on amdci2.julia.csail.mit.edu
Drop KernelSolutionVector so sol.u stays a mutable Vector with stable identity, matching master for every batch_size. Keep the CPU isbits prep fast path, and replace the dead isbits(prob) safetycopy skip with a shallow ODEProblem wrapper reconstruct when every field is isbits. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Cursor Agent <noreply@cursor.com> Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe Agent-Model: auto Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Cursor Agent <noreply@cursor.com> Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe Agent-Model: auto Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Cursor Agent <noreply@cursor.com> Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe Agent-Model: auto Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
Member
Author
|
Taking this — Devin CLI 3000.11.3, model swe-2-high (local job 560-devin on amdci2.julia.csail.mit.edu). |
…l-host-fastpath Brings in SciML#563, which resolves the check failures this PR inherited from SciML#555 fallout: the Documentation missing_docs error for internal parameter-packing helpers and the CuArray/ROCArray scalar-indexing error in ensemblegpuarray_scalar_batch.jl. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Devin <noreply@cognition.ai> Agent-Harness: Devin CLI 3000.11.3 Agent-Model: swe-2-high Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
Enzyme v0.13.209 marks the `backend` argument Duplicated instead of Const,
so the Const-annotated augmented_primal/reverse methods no longer dispatch
and every gradient through batch_solve_up_kernel errors with a
custom_rule_method_error. The backend is an empty singleton with no
differentiable fields, so leave it unannotated and accept either activity.
Verified locally: test/enzyme/{imports,transfers,gradients}.jl error on
both master and this branch under Enzyme 0.13.209 and pass with the fix.
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Devin <noreply@cognition.ai>
Agent-Harness: Devin CLI 3000.11.3
Agent-Model: swe-2-high
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
…l-host-fastpath Picks up SciML#566 (dense-output state out of registers) and SciML#568, whose Annotation-typed backend parameter in the _kernel_transfer Enzyme rule supersedes this branch's unannotated-parameter variant; the file is identical to master after resolution. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Devin <noreply@cognition.ai> Agent-Harness: Devin CLI 3000.11.3 Agent-Model: swe-2-high Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
Isbits problems are not always safe to feed to CPU kernels directly: Adapt.adapt_structure on an isbits field type still has to take effect. Check whether adapt leaves every field the kernels observe (rhs, jac, tgrad, jac_prototype, mass_matrix, initialization_data, u0, tspan, p, kwargs) unchanged, and only alias the host vector when it does; otherwise build the adapted problem vector as before. Parameterize safetycopy in the host-overhead benchmark so the default safetycopy=true path is reproducible with the committed script. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Devin <noreply@cognition.ai> Agent-Harness: Devin CLI 3000.11.3 Agent-Model: swe-2-high Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
The CPU kernel host fast path compared adapt(backend, prob) output against the original via isequal and skipped adaptation on a match. Equal-valued parameters with a different representation (e.g. Float32 vs Float64) still changed kernel results, so the isequal check was not a valid proxy for "adaptation is a no-op", and the comparison also called adapt a second time when it detected a change. Call adapt exactly once per trajectory, as master does, and only skip the KernelODEProblem repackaging for the CPU backend outside autodiff, where kernels read fields directly regardless of problem type. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Agent-Harness: Claude Code 2.1.280 Agent-Model: claude-sonnet-5-5 Agent-Session: https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np
ChrisRackauckas
marked this pull request as ready for review
October 4, 2026 14:20
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Please ignore this draft until reviewed by @ChrisRackauckas.
What changed
EnsembleGPUKernel's CPU kernel host fast path now calls
adapt(backend, prob)exactly once per trajectory, with no identity check — matching master's call pattern. Earlier rounds compared the adapted problem against the original withisequaland reused the original (unadapted) problem whenever every kernel-visible field compared equal. A round-7 review found this unsound:isequalis value equality, not representation equality, so a userAdapt.adapt_structurethat changes a parameter's numeric type (e.g.Float32 -> Float64) while leaving its value unchanged was silently discarded, changing kernel results. The fallback path also calledadapta second time on the first trajectory when the check failed, diverging from master's call count/order for stateful adapters.The remaining changes versus master are all in
src/solve.jl: (1) safety-copy construction (src/solve.jl:196-220) builds the per-trajectory copies directly, and this helper also runs for the GPU and autodiff preparation paths; (2) for the CPU backend outside autodiff,_adapt_kernel_problemreturns the adapted problem directly instead of repackaging it intoKernelODEProblemvia_kernel_record, because kernel execution and initialization read only the preserved problem fields; (3)_kernel_solution(src/solve.jl:255) builds the returned solutions. On every path,adaptis called exactly once per trajectory, in master's order.Diff vs
origin/masteris four files:src/solve.jl,test/gpu_kernel_de/host_fastpath.jl,benchmark/kernel_host_overhead.jl,test/runtests.jl.Verification
Two new regression tests pin the round-7 findings, each shown failing on the pre-fix head
f95f62dand passing on this head988d5ed:Kernel host storage honors precision-changing adaptationuses aPrecisionRecipeparameter whoseAdapt.adapt_structurewidensFloat32 -> Float64;p[1] + 1 - p[1]is exactly0in Float32 arithmetic (rate2^24) but1after the adapted widening. Onf95f62d:4 pass / 2 fail,only(sol.u[1].u[end]) == 1.0 ≈ exp(1.0) rtol=1e-10fails (Evaluated: 1.0 ≈ 2.718281828459045). On988d5ed:6/6.Kernel host storage adapts once per trajectoryuses aRateRecipeadapter that increments a counter and returnsrate * calls[], so a double call on any trajectory changes every later trajectory's value. For 3 trajectories: onf95f62d,_ADAPT_CALLS[] == 4(expected3) and endpoints areexp.([0.2, 0.3, 0.4])instead ofexp.([0.1, 0.2, 0.3])— both assertions fail. On988d5ed:4/4pass, counter== 3, endpoints matchexp.([0.1, 0.2, 0.3])tortol=1e-10.Full
test/gpu_kernel_de/host_fastpath.jl(includes both new testsets plus the existing host-storage andadapt_structureregressions):f95f62d→LoadError: Some tests did not pass: 4 passed, 2 failed(stops at the P1 testset);988d5ed→35/35across all four testsets.GROUP=<g> julia +<v> --project -e 'using Pkg; Pkg.test()', on this head — each exit 0 withTesting DiffEqGPU tests passed:GROUP=CPU+lts)35/35, scalar batch55/55, endpoint termination132/132GROUP=CPU35/35, scalar batch55/55, endpoint termination132/132GROUP=Enzyme8/8, transfer-rule/aggregate tests, ensemble gradients Float32/Float64 × adaptive true/false, Lorenz and initial-state/step-size gradients--check --diff(Julia 1.11.9) andtyposon the two changed files: exit 0.Benchmark (this head)
benchmark/kernel_host_overhead.jl,JULIA_NUM_THREADS=8, Julia 1.12.7, masterorigin/master@bacef85vs head988d5ed, sequential runs on a loaded shared host (concurrentGROUP=CPU/GROUP=Enzymetest suites were running during part of this measurement). Endpoint checksums are bit-identical between master and head at every N and bothsafetycopysettings (e.g.-9.95042280896473e6at N=2²⁰), confirming the fix changes no observable result on this workload.Public
solve()median wall seconds, 5 repetitions:Reporting this honestly: removing the identity-check fast path (round 7's required fix) removes most of the speedup earlier rounds measured — that speedup came from skipping
adaptentirely on a match, which is exactly the unsound behavior being removed. What remains is only the_kernel_recordrepackaging elision for the CPU backend, which gives a real but modest reduction forsafetycopy=true(12-29%, consistent in direction) and is within measurement noise forsafetycopy=falseon this shared host (onesafetycopy=falsepoint even came out slower). A quieter machine would narrow the noise but is not expected to change the qualitative picture: the remaining optimization avoids one small struct repackaging per trajectory, not a vector-elision.Not verified
GPU backends were not run on hardware here. Quiet-machine timing was not collected — all numbers above are from a loaded shared host with concurrent test activity.
Reviewer attention / non-blocking notes
isequal-based fast path entirely rather than tightening it to an identity/representation check, per the round-7 review's required change. There is no longer any path whereadaptis skipped for an isbits problem._adapt_kernel_problem(backend::CPU, prob)still special-caseswithin_autodiff()to go through_kernel_recordfor Enzyme; this is unchanged from the autodiff-vs-primal split in earlier rounds and was not part of either finding.ImmutableODEProbleminto the kernel (master usesKernelODEProblem) outside autodiff; downstream_kernel_transfer/convert(ImmutableODEProblem, ...)handle both, and kernel-visible behavior was unchanged in all tests and benchmark checksums above.Risk assessment
src/solve.jl(_prepare_kernel_problems,_adapt_kernel_problem, safety-copy construction,_kernel_solution). No public API change. The safety-copy helper also runs on GPU and autodiff preparation, and on those paths the adapt/record call pattern is unchanged from master. GPU hardware was not run.f95f62dand passing on988d5ed; fullGROUP=CPUon Julia 1.10.12 and 1.12.7 andGROUP=Enzymeon 1.12.7 all pass; benchmark endpoint checksums bit-identical between master and head at every size andsafetycopysetting.safetycopy=truesolves against current master 4f726d5, with lower allocations. Review: ~/sandbox/agent-jobs/DiffEqGPU.jl/reviews/pr-560/round8/REVIEW.md on amdci2.Closes #557
Original work: Cursor Agent CLI 2026.09.26-dd393fe, model auto; transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu.
Merge, CI triage + adapt_structure fix: Devin CLI 3000.11.3, model swe-2-high; local session transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu.
Round-7 fix (remove isequal fast path, call adapt once per trajectory, add P1/P2 regression tests): Claude Code 2.1.280, model claude-sonnet-5-5; session at https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np.
🤖 Generated with Claude Code
https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np