No, they're expecting to see a failure rate consistent with previous failure rates, not periods of low failure rates and other periods of high failure rates, with the same model.
And you can expect more consistency from SOTA models than you can from an old model like GPT3--you agree, right?
GP expects the same level of consistency throughout their time using the same model. Not request-to-request, more like day-to-day.