> I think running a model unsupervised in anything above medium is a sucker move to burn more tokens.
What about errors compounding and any on-the-go decisions being made being the wrong ones? That kills anything long form, no matter how good and detailed your initial plan is - there will always be something along the way.
The models would need a way to identify a difficult problem and apply max reasoning there themselves and cruise through everything else at a lower level.