It's not just inference, some things done in data centers like simulations, testing, are complementary to inference.
And in these types of hardware, the time between a successful prototype and a fully deployed product is pretty long. Maybe they count on that to know when to stop?