Clef is based on Qwen3.8-27B and Clef-flash is based on Qwen3.8-9B (edit: actually Qwen3.5-9B). So, similar in spirit to Kev by my understanding, but based on a newer model.
Atom is 60M Param (around 133x to 400x smaller).
16ms latency. And locally run.
Why go big when you can go small ?
> and Clef-flash is based on Qwen3.8-9B
There is no official qwen 3.8 9b
From the model card:
> Clef-Flash is post-trained from Qwen/Qwen3.5-9B. See Clef for the larger variant.