The way you work around this (apart from doing what you can to make the system as predictable as possible) is to accept that the data is noisy, and then working around it by doing multiple trials and following up with a statistical analysis on the results.
You can use the desired confidence to inform the warm-up, number of trials, benchmark duration, and so on.
If you end up with a multimodal distribution it can be worth tracking percentiles.
I did warm-up rounds, to prime caches and so on. I fiddled with statistics: IIRC simple average seemed to give the best results overall. I thought keeping the lowest n results would be the best, but that turned out to be wrong.
The results were indeed multimodal.
All I wanted was to have a reproducible benchmark to see how the code changes affected performance over time.