I think it's been pretty much proven by now that there are no cases where local inferencing is better than remote inferencing, unless absolute privacy is a hard requirement. The efficiencies that come with datacenter scale and hw can't be beaten.
Yeah, but data centers don't usually host abliterated models, hence the point of the article.
Yeah, but data centers don't usually host abliterated models, hence the point of the article.