Has anyone ever done a comparison between the smaller models like Luna, against the previous GPT 5 frontier models? Have we gotten to the point where the small models are as good as the frontier models of the past, or is there still a way to go?
I ran a prompt with ChatGPT since I'm also curious. The price reduction is crazy.
- GPT-5 high: score 35, approximately $0.37/task
- Luna medium: score 38, approximately $0.01/task
- Luna max: score 51, approximately $0.042/task
So Luna medium is:
- slightly more capable than GPT-5 high;
- approximately 35–40× cheaper per benchmark task.
And Luna max is:
- 16 Intelligence Index points better;
- still roughly 9× cheaper per task.
This reduction was possible within 1 year.
I ran a prompt with ChatGPT since I'm also curious. The price reduction is crazy.
- GPT-5 high: score 35, approximately $0.37/task
- Luna medium: score 38, approximately $0.01/task
- Luna max: score 51, approximately $0.042/task
So Luna medium is:
- slightly more capable than GPT-5 high;
- approximately 35–40× cheaper per benchmark task.
And Luna max is:
- 16 Intelligence Index points better;
- still roughly 9× cheaper per task.
This reduction was possible within 1 year.