logoalt Hacker News

alansabertoday at 9:34 AM0 repliesview on HN

For sure. Even if it wasn't a measure to avoid copyright, you pre-process LLM training data to remove errors, characters that can't be tokenized, etc etc. Doing so with another LLM has been standard for a while.