fix: per-trajectory error norm for EnsembleGPUArray - #571
Open
SebastianM-C wants to merge 1 commit into
Open
SebastianM-C wants to merge 1 commit into
SebastianM-C wants to merge 1 commit into
Conversation
The batched solve used `diffeqgpunorm`, an RMS over all N·ntraj components, as `internalnorm`, and passed it after `kwargs...` so it could not be overridden. A trajectory that needs small steps is outvoted by the easy ones and silently misses its tolerance, by more the larger the batch. For u' = ω cos(ω t) on [0, 10] with Tsit5 at abstol = reltol = 1e-8, trajectory 1 at ω = 50 and the rest at ω = 1, its error at t = 10 is 1.6e-9 solved alone and 5.0e-7 inside a batch of 1000. `TrajectoryNorm(len)` takes the RMS of each trajectory's own components and returns the largest, for N×ntraj state arrays and their vectorized form, on host and device arrays, and for Dual-valued arrays. With it the error above is 1.6e-9 for any batch size. The default is set before `kwargs...`, so `internalnorm` passed to `solve` replaces it. The shared step is now the one the hardest trajectory needs, so batches take more steps. Lorenz, the EnsembleGPUArray docstring example (Float32, random p, saveat = 1, CUDA): - Tsit5, 10_000 trajectories on [0, 100]: naccept 1076 -> 2092, nreject 0 -> 6 - Rosenbrock23, 1_000 trajectories on [0, 10]: naccept 265 -> 797, nreject 15 -> 128 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
EnsembleGPUArraysolves all trajectories as one batched problem, and useddiffeqgpunormas itsinternalnorm: an RMS over allN·ntrajcomponents. A trajectory that needs small steps is outvoted by the easy ones and silently misses its tolerance, by more the larger the batch. The norm was also passed afterkwargs..., so a user-suppliedinternalnormcouldn't override it.Example: u' = ω cos(ω t) on [0, 10], Tsit5, abstol = reltol = 1e-8, trajectory 1 at ω = 50 and the rest at ω = 1. Trajectory 1's error at t = 10 is 1.6e-9 solved alone and 5.0e-7 inside a batch of 1000, with no warning.
This PR adds
TrajectoryNorm(len), which takes the RMS of each trajectory's own components and returns the largest. It handlesN×ntrajstate arrays and their vectorized form, host and device arrays, andDual-valued arrays. With it the error above is 1.6e-9 at any batch size. The default is now set beforekwargs..., sointernalnormpassed tosolvereplaces it. TheEnsembleGPUArraydocstring gains a Step-size control section that describes this and the override.Behaviour change: the shared step is now the one the hardest trajectory needs, so heterogeneous batches take more steps. Lorenz, the
EnsembleGPUArraydocstring example (Float32, randomp,saveat = 1, CUDA):Those extra steps are the cost of every trajectory actually meeting its tolerance. To get the old behaviour back, pass
internalnorm = DiffEqGPU.diffeqgpunorm.Tests:
test/ensemblegpuarray_trajectory_norm.jl: the norm on host, device andDualarrays, and the ω = 50 trajectory meeting its tolerance in batches of 2 and 1000. On an RTX 4080 SUPER (CUDA group): 12/12, and 12/12 in the CPU group. The existingEnsembleGPUArraytest files still pass with the new default:ensemblegpuarray.jl6/6,ensemblegpuarray_scalar_batch.jl55/55,ensemblegpuarray_sde.jl1/1,reduction.jl3/3, andensemblegpuarray_oop.jl,ensemblegpuarray_inputtypes.jl,lower_level_api.jlrun without errors.🤖 Generated with Claude Code · model: claude-opus-5-5[1m]
Review: unreviewed