No, what they are doing is trying to optimize inference to increase margins which leads to degradations. Model deployment is not like websites, you can continuously tune performance based on usage, new memory optimizations, etc.