This just makes me think about optimizing model weights to run as 'close to the metal as possible' (i.e. within or 'just above' a UEFI boot environment, like the NightRun project), so that there isn't any OS-level overhead either. If we're optimizing, let's optimize! Curious though, maybe the OS doesn't impose much of a burden here? Open to hearing what other tinkerers think...
Once you get down to the level of writing GPU kernels the OS isn't very much in the way, aside from some host-side memory logistics.
So it's more about making those GPU kernels perform the required memory move and arithmetic operations as close to the theoretical optimum as possible.