Intel published some guidance on doing precise benchmarks. Basically run your code inside the kernel, turn off interrupts, use the cpuid instruction to prevent out-of-order execution, use rdtscp instruction instead of rdtsc, etc.
> prevent out-of-order execution
Won't this make the results inaccurate?
> prevent out-of-order execution
Won't this make the results inaccurate?