Pangram is witchcraft to me. The way they can correctly detect AI writing from small samples, with so little statistical signal, is crazy. I've seen it correctly detect AI writing even when people use those "humanizer" rewriting skills that remove the hallmarks of AI writing and make the text indistinguishable from human text. But not to Pangram.
I agree. I ran it on a small sample last night, basically three paragraphs of AI text that I thought read 100% human, except for em dashes, which I use in writing anyway. It flagged the first and last paras as "high" AI and the middle para as "high" human. I have no idea what it is seeing that sets it off.
Consider that they are in effect detecting whether text has been generated by one of a few dozen entities that has produced more text for them to analyze than any given human author.
Now imagine if they put the same effort into detecting if a given text was produced by one of a few dozen super-prolific human writers. I'd imagine they'd get pretty good at that too.
The main limitation of those "humanizer" rewriters is that most of them focus on making the next less detectable to humans, by making them read better. There's likely to be plenty of signal left that isn't affected by trying to make the text read better the same way human writers have plenty of idiosyncrasies despite being human.