How useful is the second 3090 in this setup? I run the 5-bit quantized model on a single 3090. Does ...

drnick1 • yesterday at 5:54 PM • 1 reply • view on HN

How useful is the second 3090 in this setup? I run the 5-bit quantized model on a single 3090. Does the second 3090 allow you to use the full precision model instead or a less aggressive quantization by splitting the layers? What about running the 35B model instead?

Replies

tedivm • yesterday at 11:40 PM

More memory means less aggressive quantization, more concurrent requests, and larger context windows. I also get a boost in tokens per second (not double, about 1.5x compared to a single GPU).

The 35B model is an MoE (mixture of experts), which uses only a subset of parameters at a time. The 27b one is slower but has way better performance.

alt Hacker News

Replies