So YES if you're using Cowork or Chat
No, because they are cached, the inference cost is paid once per model, does not scale linearly per user or use.
No, because they are cached, the inference cost is paid once per model, does not scale linearly per user or use.