The goal of the inference engine is to make the best use of whatever hardware you have to run models performantly and let people run bigger models. At first this will include using all the hardware on a given machine optimally. Eventually we also want to support interconnect between multiple machines to enable running bigger models!
Let me know if youre interested in a collaboration then. I am working on a custom mlx sharding system.