logoalt Hacker News

One month coding with GLM 5.3 Flash

76 points • by ThibWeb • today at 3:29 PM • 57 comments • view on HN

Comments

epistasis • today at 7:38 PM

One thing about these numbers that's absolutely shocking to me is how low the energy use is:

> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.

With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.

My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.

➕ show 6 replies
aktenlage • today at 7:22 PM

> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?

➕ show 1 reply
consumer451 • today at 10:05 PM

What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.

I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.

fallinditch • today at 9:45 PM

> For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models

GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.

I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.

➕ show 1 reply
monksy • today at 7:56 PM

GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.

➕ show 1 reply
disiplus • today at 7:53 PM

Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.

➕ show 2 replies
sheepscreek • today at 9:10 PM

> 450M tokens / $150 / 5kWh

Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.

➕ show 2 replies
ThibWeb • today at 3:29 PM

It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control

➕ show 2 replies
gpugreg • today at 7:26 PM

How did you measure energy usage?

Edit: I found a linked article that mentions the inference provider who does the measurements.

➕ show 1 reply
throw930rmdkdk • today at 6:50 PM

Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.

➕ show 2 replies
benjiro29 • today at 8:46 PM

There are some issue point not properly mentioned ...

* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).

* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.

* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.

Edit: That are some interesting downvotes ...

To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.

Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).

Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.

➕ show 2 replies