And hire 2 or 3 dev ops to keep it running ?
That another 400 to 700k.
It becomes your problem and not someone else’s. However, I don’t trust hosted LLMs for anything that needs to be private.
> However, I don’t trust hosted LLMs for anything that needs to be private.
Why not? Do you trust AWS with things that need to be private?
As mentioned, it is a "large enough company" already; they have full time sysadmins running things.
Adding another rack beside the VMWare cluster, managing any storage/networking issues, etc. will be incremental costs; they already have a pager (probably not a pager not anymore just an app on their phone) like rotation schedule etc.
There will be cloud/SaaS vendors who have lower cost of labor/capital due to automation and financing terms.
Having these models in the open caps the inference margin.
> I don’t trust hosted LLMs for anything that needs to be private
I'd update this to
'I don’t LLMs for anything that needs to be private'
What's to prevent the LLM from sliding a heavily obfuscated binary blob into the application that does nefarious things? If you aren't creating the LLM itself from scratch, I don't feel it can be trusted.
> And hire 2 or 3 dev ops to keep it running
Not a devops but I'd say one full time is already too many.
Just like that new jobs created by AI! Localized model maintainer/technician.
You’ll slap some training on existing technologists/infra/sysadmin folks and perhaps have a support contract for the edge cases (hardware troubleshooting and advanced replacement).
(managed an entire data center building with thousands of servers a lifetime ago with ~2-3 other people, it’s only gotten easier over the last two decades imho)
Why would a singe system require 3 full-time dev ops?
Just let it manage itself, what could go wrong! :)
Where do I sign up to get 200k/yr to keep one rack running? Sounds like an incredibly chill job