logoalt Hacker News

a11r • today at 4:51 PM • 9 replies • view on HN

I'm a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I'm running 4-bit quants on an RTX Pro 6000 rented for approximately $1/hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: https://www.reddit.com/r/BlackwellPerformance/s/FrKwk3GoDK


Replies

Winfred-zz • today at 9:34 PM

I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

│------------------- │ Ninfer-3090 │ Strata

│ Code generation │ 52/78 (66.7%) │ 70/78 (89.7%)

│ Code completion │ 40/50 (80.0%) │ 44/50 (88.0%)

│ Total------------- │ 92/128 (71.9%) │ 114/128 (89.1%)

│ API failures------ │ 10 │ 5

- Ninfer generation: ~122 min total.

- Strata generation: ~142 min total.

So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t/s, this modified version can do 50 t/s (but it's extremely long in it's thinking, it just goes on and on.

This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t/s.

That strata has been optimized on my Oh My Pi conversations. So when I'm using it, it's probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

➕ show 1 reply
segmondy • today at 10:30 PM

The larger the model, the more you can go down. K3 in Q1 will match and likely beat Qwen3.8-Flash-Next.

sfifs • today at 10:08 PM

I've run DeepSeek V4 Flash on DS4 on single DGX and standard model weights on aDGX cluster. There was some degradation going to the hybrid 2 but quant but really not much.

nialv7 • today at 8:02 PM

ISTA-DASLab's IQ3_S quant is really good https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RC...

➕ show 1 reply
robot_jesus • today at 4:55 PM

Can you say more about where you're renting the RTX Pro 6000 for $1/hour?

➕ show 2 replies
NinjaTrance • today at 6:16 PM

Just as curiosity, how long does it take to set the environment up and running?

Is it viable to start/stop it multiple times per day?

➕ show 1 reply
sail0rm00n • today at 4:55 PM

$1/hr sounds great. Where are you getting it for those prices?

➕ show 1 reply
aatd86 • today at 5:09 PM

$1/hour ? I need that deal as well.

➕ show 1 reply
nickpsecurity • today at 8:18 PM

That's cheap, too. Which hosting service are you using?