I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S):
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe because of quantization.
If I read correctly, that's based on a 1-bit quantization, and can we really expect that to produce any useful output at all?
You should find an excuse to offer 3D printed extruded pelicans from various models as awards for something. I have no idea for what, but the idea captivates and I'd love to win one somehow. They'd be collector's items in a few decades
The spark can easily run UD-Q4_K_XL on this model... using IQ1_S doesn't make much sense.
Doing this on a 1-bit quant is unfair
Tried again with a different quant, UD-Q2_K_XL:
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
The Pelican Brief
Why did you use 1-bit quantization vs 3-bit quantization?
It looks like the 3-bit requires 90 GB[1] which, I imagine, would fit within the DGX Spark's 128GB of unified memory.
[1] https://unsloth.ai/docs/models/qwen3.8-next