Presumably it drops back to 4.8 in those cases so it's not really worse
If it switches mid conversation, this is a massive increase in token consumption because it has to re-read your conversation into cache, right?
yes. At the bottom of the release post it says that they are releasing two new features, one of which is customizing fallback behavior instead of blocking for restricted models
If it switches mid conversation, this is a massive increase in token consumption because it has to re-read your conversation into cache, right?