Skip to content

Reduce EnsembleGPUKernel CPU host overhead - #560

Merged
ChrisRackauckas merged 10 commits into
SciML:masterfrom
ChrisRackauckas-Claude:agent/557-egpukernel-host-fastpath
Oct 4, 2026
Merged

ChrisRackauckas merged 10 commits into
SciML:masterfrom
ChrisRackauckas-Claude:agent/557-egpukernel-host-fastpath

Conversation

@ChrisRackauckas-Claude

@ChrisRackauckas-Claude ChrisRackauckas-Claude commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Please ignore this draft until reviewed by @ChrisRackauckas.

What changed

EnsembleGPUKernel's CPU kernel host fast path now calls adapt(backend, prob) exactly once per trajectory, with no identity check — matching master's call pattern. Earlier rounds compared the adapted problem against the original with isequal and reused the original (unadapted) problem whenever every kernel-visible field compared equal. A round-7 review found this unsound: isequal is value equality, not representation equality, so a user Adapt.adapt_structure that changes a parameter's numeric type (e.g. Float32 -> Float64) while leaving its value unchanged was silently discarded, changing kernel results. The fallback path also called adapt a second time on the first trajectory when the check failed, diverging from master's call count/order for stateful adapters.

The remaining changes versus master are all in src/solve.jl: (1) safety-copy construction (src/solve.jl:196-220) builds the per-trajectory copies directly, and this helper also runs for the GPU and autodiff preparation paths; (2) for the CPU backend outside autodiff, _adapt_kernel_problem returns the adapted problem directly instead of repackaging it into KernelODEProblem via _kernel_record, because kernel execution and initialization read only the preserved problem fields; (3) _kernel_solution (src/solve.jl:255) builds the returned solutions. On every path, adapt is called exactly once per trajectory, in master's order.

Diff vs origin/master is four files: src/solve.jl, test/gpu_kernel_de/host_fastpath.jl, benchmark/kernel_host_overhead.jl, test/runtests.jl.

Verification

Two new regression tests pin the round-7 findings, each shown failing on the pre-fix head f95f62d and passing on this head 988d5ed:

  • P1 (precision-changing adaptation silently dropped). Kernel host storage honors precision-changing adaptation uses a PrecisionRecipe parameter whose Adapt.adapt_structure widens Float32 -> Float64; p[1] + 1 - p[1] is exactly 0 in Float32 arithmetic (rate 2^24) but 1 after the adapted widening. On f95f62d: 4 pass / 2 fail, only(sol.u[1].u[end]) == 1.0 ≈ exp(1.0) rtol=1e-10 fails (Evaluated: 1.0 ≈ 2.718281828459045). On 988d5ed: 6/6.
  • P2 (adapter called more than once per trajectory). Kernel host storage adapts once per trajectory uses a RateRecipe adapter that increments a counter and returns rate * calls[], so a double call on any trajectory changes every later trajectory's value. For 3 trajectories: on f95f62d, _ADAPT_CALLS[] == 4 (expected 3) and endpoints are exp.([0.2, 0.3, 0.4]) instead of exp.([0.1, 0.2, 0.3]) — both assertions fail. On 988d5ed: 4/4 pass, counter == 3, endpoints match exp.([0.1, 0.2, 0.3]) to rtol=1e-10.

Full test/gpu_kernel_de/host_fastpath.jl (includes both new testsets plus the existing host-storage and adapt_structure regressions): f95f62d → LoadError: Some tests did not pass: 4 passed, 2 failed (stops at the P1 testset); 988d5ed → 35/35 across all four testsets.

GROUP=<g> julia +<v> --project -e 'using Pkg; Pkg.test()', on this head — each exit 0 with Testing DiffEqGPU tests passed:

suite Julia result
GROUP=CPU 1.10.12 (+lts) passed — host storage 35/35, scalar batch 55/55, endpoint termination 132/132
GROUP=CPU 1.12.7 passed — host storage 35/35, scalar batch 55/55, endpoint termination 132/132
GROUP=Enzyme 1.12.7 passed — imports 8/8, transfer-rule/aggregate tests, ensemble gradients Float32/Float64 × adaptive true/false, Lorenz and initial-state/step-size gradients
  • Runic --check --diff (Julia 1.11.9) and typos on the two changed files: exit 0.

Benchmark (this head)

benchmark/kernel_host_overhead.jl, JULIA_NUM_THREADS=8, Julia 1.12.7, master origin/master@bacef85 vs head 988d5ed, sequential runs on a loaded shared host (concurrent GROUP=CPU/GROUP=Enzyme test suites were running during part of this measurement). Endpoint checksums are bit-identical between master and head at every N and both safetycopy settings (e.g. -9.95042280896473e6 at N=2²⁰), confirming the fix changes no observable result on this workload.

Public solve() median wall seconds, 5 repetitions:

safetycopy N master head change
true 2¹⁶ 0.1309 0.0934 29% faster
true 2¹⁸ 0.5766 0.4398 24% faster
true 2²⁰ 2.3029 2.0300 12% faster
false 2¹⁶ 0.0727 0.0735 within noise
false 2¹⁸ 0.3215 0.3515 within noise (head slower)
false 2²⁰ 1.4820 1.3744 7% faster

Reporting this honestly: removing the identity-check fast path (round 7's required fix) removes most of the speedup earlier rounds measured — that speedup came from skipping adapt entirely on a match, which is exactly the unsound behavior being removed. What remains is only the _kernel_record repackaging elision for the CPU backend, which gives a real but modest reduction for safetycopy=true (12-29%, consistent in direction) and is within measurement noise for safetycopy=false on this shared host (one safetycopy=false point even came out slower). A quieter machine would narrow the noise but is not expected to change the qualitative picture: the remaining optimization avoids one small struct repackaging per trajectory, not a vector-elision.

Not verified

GPU backends were not run on hardware here. Quiet-machine timing was not collected — all numbers above are from a loaded shared host with concurrent test activity.

Reviewer attention / non-blocking notes

  • The fix removes the isequal-based fast path entirely rather than tightening it to an identity/representation check, per the round-7 review's required change. There is no longer any path where adapt is skipped for an isbits problem.
  • _adapt_kernel_problem(backend::CPU, prob) still special-cases within_autodiff() to go through _kernel_record for Enzyme; this is unchanged from the autodiff-vs-primal split in earlier rounds and was not part of either finding.
  • The CPU primal fast path still passes ImmutableODEProblem into the kernel (master uses KernelODEProblem) outside autodiff; downstream _kernel_transfer/convert(ImmutableODEProblem, ...) handle both, and kernel-visible behavior was unchanged in all tests and benchmark checksums above.

Risk assessment

  • Risk: low
  • Blast radius: EnsembleGPUKernel kernel-problem preparation in src/solve.jl (_prepare_kernel_problems, _adapt_kernel_problem, safety-copy construction, _kernel_solution). No public API change. The safety-copy helper also runs on GPU and autodiff preparation, and on those paths the adapt/record call pattern is unchanged from master. GPU hardware was not run.
  • Evidence: two new regression tests (P1 precision-changing adaptation, P2 double-adapt-call) shown failing on f95f62d and passing on 988d5ed; full GROUP=CPU on Julia 1.10.12 and 1.12.7 and GROUP=Enzyme on 1.12.7 all pass; benchmark endpoint checksums bit-identical between master and head at every size and safetycopy setting.
  • Independent review: Codex CLI / gpt-6-astra rated it low risk, high confidence for the exercised CPU paths, MERGE at 988d5ed (round 8). It reran the round-6 and round-7 reproducers, saw the new tests go from 4 pass / 6 fail on f95f62d to 10/10 here, confirmed CPU on 1.10 and 1.12 and Enzyme on 1.12, and measured 25-30% faster default safetycopy=true solves against current master 4f726d5, with lower allocations. Review: ~/sandbox/agent-jobs/DiffEqGPU.jl/reviews/pr-560/round8/REVIEW.md on amdci2.
  • Independent review: Codex CLI (Astra round 7) rated the pre-fix head CHANGES (P1/P2 above); this push implements both required changes and is not yet independently re-reviewed.
  • Merge: low risk. The independent Astra round-8 review is MERGE; the remaining unverified item is GPU hardware. This is a partial improvement for EnsembleGPUKernel solve(): reduce host-side overhead (~0.8 s at 2^20 trajectories) #557: host overhead does not yet approach kernel time, so EnsembleGPUKernel solve(): reduce host-side overhead (~0.8 s at 2^20 trajectories) #557 stays open for the rest.

Closes #557

Original work: Cursor Agent CLI 2026.09.26-dd393fe, model auto; transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu.

Merge, CI triage + adapt_structure fix: Devin CLI 3000.11.3, model swe-2-high; local session transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu.

Round-7 fix (remove isequal fast path, call adapt once per trajectory, add P1/P2 regression tests): Claude Code 2.1.280, model claude-sonnet-5-5; session at https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np.

🤖 Generated with Claude Code

https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np

ChrisRackauckas and others added 5 commits September 27, 2026 13:43
Store isbits CPU problems once and construct default ensemble solutions on access.
Keep eager output for custom functions and differentiated solves.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Codex <noreply@openai.com>
Agent-Harness: Codex CLI 0.157.1
Agent-Model: gpt-6-sol
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-codex/log.txt on amdci2.julia.csail.mit.edu
Inline the shared solution helper so Enzyme retains concrete type information
in the differentiated ensemble output comprehension.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Codex <noreply@openai.com>
Agent-Harness: Codex CLI 0.157.1
Agent-Model: gpt-6-sol
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-codex/log.txt on amdci2.julia.csail.mit.edu
Drop KernelSolutionVector so sol.u stays a mutable Vector with stable
identity, matching master for every batch_size. Keep the CPU isbits
prep fast path, and replace the dead isbits(prob) safetycopy skip with
a shallow ODEProblem wrapper reconstruct when every field is isbits.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Cursor Agent <noreply@cursor.com>
Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe
Agent-Model: auto
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Cursor Agent <noreply@cursor.com>
Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe
Agent-Model: auto
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Cursor Agent <noreply@cursor.com>
Agent-Harness: Cursor Agent CLI 2026.09.26-dd393fe
Agent-Model: auto
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/557-cursor/log.txt on amdci2.julia.csail.mit.edu
@ChrisRackauckas-Claude

Copy link
Copy Markdown
Member Author

Taking this — Devin CLI 3000.11.3, model swe-2-high (local job 560-devin on amdci2.julia.csail.mit.edu).

ChrisRackauckas and others added 5 commits October 2, 2026 12:25
…l-host-fastpath

Brings in SciML#563, which resolves the check failures this PR inherited from
SciML#555 fallout: the Documentation missing_docs error for internal
parameter-packing helpers and the CuArray/ROCArray scalar-indexing error
in ensemblegpuarray_scalar_batch.jl.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Devin <noreply@cognition.ai>
Agent-Harness: Devin CLI 3000.11.3
Agent-Model: swe-2-high
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
Enzyme v0.13.209 marks the `backend` argument Duplicated instead of Const,
so the Const-annotated augmented_primal/reverse methods no longer dispatch
and every gradient through batch_solve_up_kernel errors with a
custom_rule_method_error. The backend is an empty singleton with no
differentiable fields, so leave it unannotated and accept either activity.
Verified locally: test/enzyme/{imports,transfers,gradients}.jl error on
both master and this branch under Enzyme 0.13.209 and pass with the fix.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Devin <noreply@cognition.ai>
Agent-Harness: Devin CLI 3000.11.3
Agent-Model: swe-2-high
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
…l-host-fastpath

Picks up SciML#566 (dense-output state out of registers) and SciML#568, whose
Annotation-typed backend parameter in the _kernel_transfer Enzyme rule
supersedes this branch's unannotated-parameter variant; the file is
identical to master after resolution.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Devin <noreply@cognition.ai>
Agent-Harness: Devin CLI 3000.11.3
Agent-Model: swe-2-high
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
Isbits problems are not always safe to feed to CPU kernels directly:
Adapt.adapt_structure on an isbits field type still has to take effect.
Check whether adapt leaves every field the kernels observe (rhs, jac,
tgrad, jac_prototype, mass_matrix, initialization_data, u0, tspan, p,
kwargs) unchanged, and only alias the host vector when it does; otherwise
build the adapted problem vector as before.

Parameterize safetycopy in the host-overhead benchmark so the default
safetycopy=true path is reproducible with the committed script.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Devin <noreply@cognition.ai>
Agent-Harness: Devin CLI 3000.11.3
Agent-Model: swe-2-high
Agent-Session: local session, transcript at /home/crackauc/sandbox/agent-jobs/DiffEqGPU.jl/jobs/560-devin/log.txt on amdci2.julia.csail.mit.edu
The CPU kernel host fast path compared adapt(backend, prob) output
against the original via isequal and skipped adaptation on a match.
Equal-valued parameters with a different representation (e.g. Float32
vs Float64) still changed kernel results, so the isequal check was not
a valid proxy for "adaptation is a no-op", and the comparison also
called adapt a second time when it detected a change. Call adapt
exactly once per trajectory, as master does, and only skip the
KernelODEProblem repackaging for the CPU backend outside autodiff,
where kernels read fields directly regardless of problem type.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Agent-Harness: Claude Code 2.1.280
Agent-Model: claude-sonnet-5-5
Agent-Session: https://claude.ai/code/session_01SJfy3ez7u4CqvAerKrT5np
@ChrisRackauckas
ChrisRackauckas marked this pull request as ready for review October 4, 2026 14:20
@ChrisRackauckas
ChrisRackauckas merged commit b6ae953 into SciML:master Oct 4, 2026
27 of 31 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

EnsembleGPUKernel solve(): reduce host-side overhead (~0.8 s at 2^20 trajectories)

2 participants