I've been thinking about buying a system to run LLMs locally but the price for one that'll run Qwen3.8-27B well is quite offputting to say the least.
What I've been looking at instead is inference providers that use TEE and E2EE to provide cryptographic guarantees that my prompts and responses are only visible to me and the GPU itself.
Despite their docs and assurances of what their guarantees mean, I'm having trouble getting to a point where I'm actually comfortable trusting them with secrets though. Phala for example seems to be E2EE only to the gateway and will then forward prompts to (potentially third party) providers.
Has anyone been down this path and found a provider they feel safe with?
Well, yes me. I was on the same path. Researching hardware and coming to the same conclusion. I found tinfoil.sh which looks promising?
My current solution is a private ChatGPT-like interface using OpenRouter’s API with Zero Data Retention enabled. Not perfect or verifiable but I think it’s acceptable for now.
https://news.ycombinator.com/item?id=43996555
https://openrouter.ai/docs/guides/features/zdr