logoalt Hacker News

altmanaltmantoday at 8:13 AM0 repliesview on HN

> In preliminary testing, Mantzarlis found Pangram was more likely to misclassify AI-generated text as human-authored when it rhymed, repeated itself, and when it used archaic language. He then built an adversarial set of 588 AI-generated text samples tailored to these weaknesses. When he used Pangram to evaluate them, the tool falsely labelled AI text as human 86% of the time.

> “I don't think that Pangram is bad,” Mantzarlis said. “I think actually Pangram at scale is probably a pretty solid tool. That said, I am extremely worried about it being used in individual cases.”

https://reutersinstitute.politics.ox.ac.uk/news/human-wrote-...

Using it for an individual article to fully determine if its AI or not is "impossible" because you're not even using the tool properly.