logoalt Hacker News

tesnorindiantoday at 7:07 AM0 repliesview on HN

I gave owao/Nanbeige4.2-3B-GGUF (Q8 quant) a try to understand how loop transformers work and compare it with other models especially with Ling 3 Tiny MoE model. As reported in the article, it is compute intensive (due to looped layers) and made a mistake during tool call just like how Ling 3 Tiny MoE did for exactly the same prompt.