They normally release a 35b dense and an 27b moe (4B active per token)
For context 35B on my m4 runs at 10 tokens a second, 27B moe runs 50-60 tokens a second.
You have your numbers switched. 27B is the dense model and runs slowly on unified memory. 35B (A3B active) runs great on unified memory.
You have your numbers switched. 27B is the dense model and runs slowly on unified memory. 35B (A3B active) runs great on unified memory.