logoalt Hacker News

walrus01yesterday at 8:38 AM9 repliesview on HN

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month.


Replies

mdasenyesterday at 2:08 PM

That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.

There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.

Average electric rates by region:

    New England            28.1 cents
    Mid Atlantic           25.1 cents
    East North Central     20.8 cents
    West North Central     14.8 cents
    South Atlantic         16.1 cents
    East South Central     15.5 cents
    Mountain               14.6 cents
    Pacific Contiguous     26.1 cents
    Pacific Noncontiguous  42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph...
show 4 replies
ljlolelyesterday at 2:51 PM

can send safely context if there’s confidential computing ala my site https://trustedrouter.com/

show 1 reply
ComputerGuruyesterday at 8:56 PM

> I pay about $0.075 USD per kWh

Around here electricity companies quote prices like yours but that is supply only while transmission, taxes, and fees are again as much on top. Is that really all inclusive?

pizza234yesterday at 9:26 PM

Which are those use cases, considering that the hypothetical server, as described by the GP, is also extremely slow (4/6 hours for a response)?

fookeryesterday at 8:51 AM

Great, so the other member of the set matters for you more than cost.

Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?

show 3 replies
light_hue_1yesterday at 9:51 AM

As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.

show 4 replies
polycancelyesterday at 7:27 PM

>and people will compromise speed for data sovereignty

People should always compromise speed for data sovereignty! Who said: that in this digital day and age, information about money is more important than money!

doctorpanglossyesterday at 5:12 PM

It's completely academic. At 5tok/s you can process 13 MTok per month at concurrency 1. I use 5 BILLION tokens per week when coding.

show 1 reply
tiahurayesterday at 4:06 PM

There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty.

Are there? At the highest levels of defense and law, AWS and Azure are used.

Having tried selling some of these entities on doing things in-house, there seems to be little interest.

show 1 reply