3.5 Flash was always too expensive for a "flash" model. They marketed it as "near frontier" level, but there are several order-of-magnitude cheaper open models that compete with it.
In my tests, 3.6 Flash is NOT more token efficient, so it actually ends up costing more than 3.5 Flash, even with the output price reduction.
EDIT:
It less less verbose in final output though, but it reasons more.
I assume the optimization comes when you have long-running tasks with many tool calls, and by reasoning more, it reduces the number of tool calls needed.
In my tests, 3.6 Flash is NOT more token efficient, so it actually ends up costing more than 3.5 Flash, even with the output price reduction.
EDIT: It less less verbose in final output though, but it reasons more.
I assume the optimization comes when you have long-running tasks with many tool calls, and by reasoning more, it reduces the number of tool calls needed.