I wish they would release the quantized versions in a safetensor format. Many frameworks can't load PTE and GGUF.
Try quantizing yourself, just a suggestion.
https://github.com/pytorch/executorch/tree/main/examples/mod...
Try quantizing yourself, just a suggestion.
https://github.com/pytorch/executorch/tree/main/examples/mod...