logoalt Hacker News

hyperbovineyesterday at 6:46 PM4 repliesview on HN

Try instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act.

These models might be smart but they're not close to being able to savor irony.


Replies

jbmyesterday at 7:14 PM

I was a little radicalized when ChatGPT literally refused to translate parts of 1000+ year old religious texts and told me it was due to copyright concerns.

show 6 replies
bijowo1676yesterday at 7:37 PM

in case of chatgpt/anthropic, the LLM model simply represents the hypocrisy of their owners

show 1 reply
citizenpaulyesterday at 7:50 PM

I live in SV. When I was at the grocery store last year I overheard a group of lawyers talking about their progress on litigation against AI companies and how they need more SWE help to progress.

I'd say that they have valid concerns about being cagey on the copyright stuff despite the obvious hypocrisy of it.

Stealing IP is effectively legal in China so they don't really have the same concerns.

show 2 replies
Onavoyesterday at 7:02 PM

This behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.

show 3 replies