Someone already suggested bringing it to the cops earlier in this thread
Technoguyfication
joined 1 year ago
It’s wild to see people in the piracy community of all places have an issue with someone benefiting from data they got online for free.
Those are the best projects. There’s no bugs, all unit tests passed, no tickets to look at. Pure bliss.
People are acting like ChatGPT is storing the entire Harry Potter series in its neural net somewhere. It’s not storing or reproducing text in a 1:1 manner from the original material. Certain material, like very popular books, has likely been interpreted tens of thousands of times due to how many times it was reposted online (and therefore how many times it appeared in the training data).
Just because it can recite certain passages almost perfectly doesn’t mean it’s redistributing copyrighted books. How many quotes do you know perfectly from books you’ve read before? I would guess quite a few. LLMs are doing the same thing, but on mega steroids with a nearly limitless capacity for information retention.