529
‘Impossible’ to create AI tools like ChatGPT without copyrighted material, OpenAI says
(www.theguardian.com)
This is a most excellent place for technology news and articles.
¿Porque no los dos?
I don't understand why people are defending AI companies sucking up all human knowledge by saying "well, yeah, copyrights are too long anyway".
Even if we went back to the pre-1976 term of 28 years, renewable once for a total of 56 years, there's still a ton of recent works that AI are using without any compensation to their creators.
I think it's because people are taking this "intelligence" metaphor a bit too far and think if we restrict how the AI uses copyrighted works, that would restrict how humans use them too. But AI isn't human, it's just a glorified search engine. At least all standard search engines do is return a link to the actual content. These AI models chew up the content and spit out something based on it. It simply makes sense that this new process should be licensed separately, and I don't care if it makes some AI companies go bankrupt. Maybe they can work adequate payment for content into their business model going forward.
Because it's not just big companies that are affected; it's the technology itself. People saying you can't train a model on copyrighted works are essentially saying nobody can develop those kinds of models at all. A lot of people here are naturally opposed to the idea that the development of any useful technology should be effectively illegal.
I am not saying you can't train on copyrighted works at all, I am saying you can't train on copyrighted works without permission. There are fair use exemptions for copyright, but training AI shouldn't apply. AI companies will have to acknowledge this and get permission (probably by paying money) before incorporating content into their models. They'll be able to afford it.
What if I do it myself? Do I still need to get permission? And if so, why should I?
I don't believe the legality of doing something should depend on who's doing it.
Yes you would need permission. Just because you’re a hobbyist doesn’t mean you’re exempt from needing to follow the rules.
As soon as it goes beyond a completely offline, personal, non-replicatible project, it should be subject to the same copyright laws.
If you purely create a data agnostic AI model and share the code, there’s no problem, as you’re not profiting off of the training data. If you create an AI model that’s available for others to use, then you’d need to have the licensing rights to all of the training data.