Do you support splitting models across devices so larger models can run on clusters?
I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.
The goal of the inference engine is to make the best use of whatever hardware you have to run models performantly and let people run bigger models. At first this will include using all the hardware on a given machine optimally. Eventually we also want to support interconnect between multiple machines to enable running bigger models!