1) "Dear artists, the model cannot infringe upon your copyright because it's merely learning like a human does. If it accidentally outputs parts of your book, you know, it just accidentally plagiarized. We all do it haha! Our attorneys remind you that plagiarism is not illegal in the US."
2) "Dear engineers, the output of our model is copyrighted and thus if you use it to train your own model, we own it."
I am not sure how both of those can be true at the same time.
2) doesn't line up with the US court's current stance that only a human can hold copyright, and thus anything created by a not-human cannot have copyright applied. This applies to animals, inanimate objects, and presumably, AI.
I have no idea how this impacts the encodability of the license from FB which may rely on things other than copyright, but as of right now, the output absolutely cannot be copyrighted.
Adobe doesn't hold copyright on images produced using Photoshop. Assuming prompt guidance can be used to claim copyright (unclear, see https://arstechnica.com/information-technology/2023/02/us-co... ), that copyright would presumably be held by the person doing the guidance and not the company that trained the AI.
We all truly do "accidentally plagiarize", especially artists. Many guitarists realize they accidentally copied a riff they thought they'd come up with on their own for example.
I added the "haha" in there because the probability of a human doing this kind of goes way down as the length of the text increases. Can you type, verbatim, an entire chapter of a book? I can't. But, I bet the AI can be convinced in rare cases to do that.
The whole thing is very interesting to me. There was an article on here a couple days ago about using gzip as a language model. Of course, gzipping a book doesn't remove the copyright. So how low does the probability of outputting the input verbatim have to be before copyright is lost?
Reading the book and benefitting from what you learned? Obviously not copyright infringement. Putting the book into gzip and sending your friend the result? Obviously copyright infringement. Now we're in the grey area and ... nobody knows what the law is, or honestly, even how to reason about what the law wants here. Fun times.
(Personally, I lean towards "not copyright infringement", but I'm not a big believer in copyright myself. In the case of AI training, it just makes it impossible for small actors to compete. Google can just buy a license from every book distributor. SmolStartup can't. So if we want to make AI that is only for the rich and powerful, copyright is the perfect tool to enable that. I don't think we want that, though.
My take is that the rest of society kind of hates Tech right now ("I don't really like my Facebook friends, so someone should take away Mark Zuckerberg's money."), so it's likely that protectionist laws will soon be created that ruin it for everyone. The net effect of that is that Europe and the US will simply flat-out lose to China, which doesn't care about IP.)
China currently has the most stringent limits for LLMs available to end users because of concerns about their political alignment. So if you believe that it's the whole market competition part that's most important in getting the best results long term, they have shot themselves in the foot first.
Of course, the models that are developed for internal use by the Chinese government won't be so limited, regardless of what the law says. But then neither be the ones developed by Western three-letter agencies. So don't worry about the "Great Game"; they'll do just fine one-upping each other and screwing over all of us in the process.
The overwhelming majority of all human advancement is in the form of interpolation. Real extrapolation is extremely rare and most don't even know when it's happening. This is why it's extremely hypocritical for artists of any sort to be upset about Generative AI. Their own minds are doing the same exact thing they get upset about the model doing.
This is why fundamental "interpolative" techniques like ChatGPT (whose weights are in theory frozen) is still basically super-intelligent.
Wow you appear to know a great deal about how human minds work: "doing the same exact thing they get upset about the model doing"...
May I query you put up a list of publications on the subject of how minds work?
My insights are widely accepted theories from various fields, all available in the public domain.
It's a well-understood concept that our minds function by making sense of the world through patterns. This is the essence of interpolation - taking two known points and making an educated guess about what lies in between. Ever caught yourself finishing someone's sentence in your mind before they do? That's your brain extrapolating based on previous patterns of speech and context. These processes are at the heart of human creativity.
The field of Cognitive Science has extensively documented our tendency for interpolation and pattern recognition. Works like The Handbook of Imagination and Mental Simulation by Markman and Klein, or even "How Creativity Works in the Brain" by the National Endowment for the Arts all attest to this.
When artists create, they draw from their experiences, their knowledge, their understanding of the world - a process overwhelmingly of interpolation.
Now, I can see how you might be confused about my reference to ChatGPT being "super-intelligent". Perhaps "hyper-competent" would be more appropriate? It has the ability to generate text that appears intelligent because it's interpolating from a massive amount of data - far more than any human could consciously process. It's the ultimate pattern finder.
And that, my friend, is my version of "publications on the subject of how minds work." I may not be an illustrious scholar, but hey, even a clock is right twice a day! And who knows, maybe I'm on to something after all.
There was a famous case where John Fogerty (formerly of Creedence Clearwater Revivial) ended up getting sued by CCR's record label, claiming a later solo song he did with a different label was too similar to a CCR song that he wrote, and they won. So legally speaking, you can even get in trouble for coming up with the same thing twice if don't own the copyright of the first one.
The copyright situation with music is kinda broken, different parts of the performance get quite different priority when it comes to copyright (many core elements of a performance get basically no protection, whereas the threshold for what counds as a protectable melody is absurdly low). Especially this means its less than worthless for some genres/traditions: for jazz and blues, especially, a huge part of the genre and culture is adapting and playing with a shared language of common riffs.
1) "Dear artists, the model cannot infringe upon your copyright because it's merely learning like a human does. If it accidentally outputs parts of your book, you know, it just accidentally plagiarized. We all do it haha! Our attorneys remind you that plagiarism is not illegal in the US."
2) "Dear engineers, the output of our model is copyrighted and thus if you use it to train your own model, we own it."
I am not sure how both of those can be true at the same time.