this post was submitted on 29 Sep 2023
439 points (93.5% liked)

Technology

59207 readers
3007 users here now

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


founded 1 year ago
MODERATORS
 

Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

you are viewing a single comment's thread
view the rest of the comments
[–] [email protected] 4 points 1 year ago

If an AI "reproduces" a work it was trained on it is a failure of an AI. Why would anyone want to spend millions of dollars and devote oodles of computing power to build something that just does what a simple copy/paste operation can accomplish?

When an AI spits out something that's too close to one of the original training set that's called "overfitting" and it is considered an error to be corrected. Most overfitting that's been detected has been a result of duplication in the training set - when you hammer an AI image generator in training with thousands of copies of the Mona Lisa it eventually goes "alright, I get it already, when you say 'Mona Lisa' you want that exact pattern!" And will try its best to replicate that pattern when you ask it to later. That's why training sets need to be de-duplicated.

AIs are meant to produce new things.