Not necessarily, because the models have an element of randomness. Also, I was under the impression that ChatGPT has more "safeguards" (manifesting as a refusal to answer questions) than the raw API.
I don’t doubt the poster was telling the truth when they said they asked for a summary of the book and didn’t get one.
It refutes the idea that chatgpt’s inability to provide a summary means it didn’t scan the original text: since it can provide a summary, the argument is entirely spurious.
The part that's interesting is whether the summary is correct, though. Of course, depending on how you prompt it, you might or might not get an outright refusal.
How is it different from asking to me to summarize anything? I could have bought the book, or read the Wikipedia page, or listened people talking about it, or downloaded the torrent. In all those cases my summary could be right or could be wrong.
If the rights holders know that I dowloaded the torrent they could sue me. In the other cases they can't.
What if it turns out that OpenAI bought a copy of every book ingested be ChatGPT?
> In all those cases my summary could be right or could be wrong.
Well that's incredibly nihilistic. Whether the summary is correct or not matters a great deal! And if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct.
But ChatGPT? Who the hell knows? You can't trust a thing it says, especially about obscure topics. The summary is useless if you have to do a bunch of verification to see if any of it is even true, a problem that summaries even by moderately competent human writers don't have!
if someone I knew said they read a book, even a very obscure one, and then summarized it to me, I'd have great confidence that they would get such simple facts as "who are the characters" and "what are the major plot points" correct.
People, especially people you know, have reputations, based on history and experience that others have dealing with them. People can be known as liars, and anything they say is colored by such a reputation. Humans have language idioms for communicating about and dealing with such people too, phrases like "take anything that person says with a grain of salt". Look at how George Santos' history of lying about his own experience is being dealt with.
ChatGPT can be (is?) the same, and it has a bad reputation for truth telling. And LLMs' reputation is not necessarily getting better in this regard.
The problem is that many people attribute output that came from a machine to be of higher quality (on whatever axis) than output that came from a human, even a human they personally know and have experience dealing with. This is the same kind of prejudice as any other, or perhaps a more insidious prejudice.
Agree 100% with all of this. LLMs have a huge reputation problem; you simply cannot trust what they say because they've been proven time and time again to hallucinate fictional answers. Until that problem is solved I'm struggling to see how they're as useful as people are claiming they are.
You know what would be a fun test of integrity -- look up an obscure novel (potentially even the aforementioned one) that you know LLMs consistently hallucinate about because the details aren't in its training set, and then assign an essay about it as an academic assignment. It'll be pretty obvious who's read the book and who merely consulted an LLM because the latter will just be complete gibberish to anyone who actually knows what happens in that novel.
I feel like thats one of the many questions regulators and law makers are going to be asked long term. I'm sure buying the book for "commercial purposes" like that would't be appropriate, but then again, does that mean if I read it and then summarize it in my work, or regurgitate its info as part of my job...I'm violating a license?
A world where humans have special permissions but LLMs don't seems pretty interesting to consider, especially if they're both doing the same kind of things with the data.
There's nothing illegal about reading a book and then circulating your summary/review of it. This doesn't even get into issues of fair use because you aren't redistributing any of the copyrighted material in the first place, merely facts and your own opinions about it.
The separate issue that's concerning is that GPT can't be trusted to accurately summarize anything obscure at all, but it'll sure throw text at you nonetheless.
> There’s nothing illegal about reading a book and then circulating your summary/review of it. This doesn’t even get into issues of fair use because you aren’t redistributing any of the copyrighted material in the first place, merely facts and your own opinions about it.
A summary may or may not be a derivative work before considering Fair Use; “redistributing copyrighted material” isn’t the only exclusive right of copyright: producing copies is, but more to the point so is producing derivative works.
If the rights holders know that you downloaded the torrent they can sue you - but the fact that you produced that summary is weak evidence for that claim.
Producing the summary is absolutely not an infringing act. Downloading the torrent might be.
What if that turns out to be completely irrelevant?
Let's say, for the sake of argument, that I knew absolutely nothing about contract law and was then filmed stealing a book you wrote on the subject from a book store. I then started a business where I would answer questions about contract law, based solely on what I learned from the book. Of course, my memory isn't perfect, but I don't like to admit when I'm wrong, so sometimes I just make stuff up. People line up to pay me anyway.
Now, the owner of the book store you might be able to get me arrested for petty theft. Do you think there is any possibility you, as the author, could successfully litigate a copyright claim against me? I'd argue not. Do you think you could get an injunction enjoining me from engaging in my contract law Q&A business? Again, I think that would be highly unlikely.
It isn't clear to me that any court is going to hold that LLMs are being used to create derivative works, any more than someone who reads a book, whether they paid for it or not, and then speaks or writes about a topic covered by a book they've read has done so. It is entirely possible the IP laws, as they currently exist simply do not cover what LLMs are doing. The laws certainly were not written with this kind of scenario in mind.
Replace "answer questions about contract law, based solely on what I learned from the book" with "generate cartoon images, based solely on what ML learned from Disney IP" and see how badly that will go.
Replace it with any subject. The point stands - it isn't at all clear how the courts will treat this.
Take your example: I'm a self-taught artist, and I learned everything I know about art by studying cartoons made by Disney. Maybe I paid for these cartoons, maybe I didn't. I then make a website where I draw my own cartoons, which, since I've never seen any other art, look a lot like Disney's. Unless I'm straight-up copying their characters, they would have no claim against me.
> I then make a website where I draw my own cartoons, which, since I've never seen any other art, look a lot like Disney's. Unless I'm straight-up copying their characters, they would have no claim against me.
The law pertains ultimately to the actions of humans. We don't allow non-human animals or machines access to legal system. Even in the specious only-Disney-inspired-artist scenario presented (courts don't use unrealistic hypotheticals like that), there would have to be consideration given to the fact that you somehow never got any access to other art, so you were severely disadvantaged.
But most of all, you the disadvantaged Disneyesque-drawing artist not a generative AI, so you should have more legal latitude to create works inspired than others work than the person who creates the LLM has.
The LLM creator instead has just created a very good style plagiarism machine, one that lacks the ability to be inspired, much less attribute the styles that it plagiarizes.