logoalt Hacker News

ttultoday at 4:53 AM0 repliesview on HN

Most people running local models would probably love to run larger models if only they had access to big enough hardware. I'm curious: to those of you running models locally, if there was a way to inference the model of your choice at a reasonable cost by effectively time-sharing a B300 rack through some privacy-protecting intermediary, would you consider that?

If there was a "Mullvad of GPU clouds", would that solve the privacy concerns?