Once you get down to the level of writing GPU kernels the OS isn't very much in the way, aside from some host-side memory logistics.
So it's more about making those GPU kernels perform the required memory move and arithmetic operations as close to the theoretical optimum as possible.