• mac@lemm.ee
    link
    fedilink
    English
    arrow-up
    0
    ·
    1 day ago

    is it not relatively trivial to pre-vet content before they train it? at least with aigen text it should be.

    • General_Effort@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      arrow-down
      1
      ·
      1 day ago

      It depends on what you are looking for. Identifying AI generated data is generally hard, though it can be done in specific cases. There is no mathematical difference between the 1s and 0s that encoded AI generated data and any other data. Which is why these model collapse ideas are just fantasy. There is nothing magical about any data that makes it “poisonous” to AI. The kernel of truth behind these ideas is not likely to matter in practice.