Rapid-MLX was my first thought too. It’s optimized for self-hosted agents on Apple silicon. It’s my current choice for self-hosted models.
https://github.com/raullenchai/Rapid-MLX