Yeah that doesn't sound quite right. Given that the 5080 has 16 GB of VRAM I would have expected a few more options, for example Gemma 12B, to at least show up as available. Do any bigger models show as available or was that specifically the recommendation?
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.
Yeah that doesn't sound quite right. Given that the 5080 has 16 GB of VRAM I would have expected a few more options, for example Gemma 12B, to at least show up as available. Do any bigger models show as available or was that specifically the recommendation?
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.