FaceDeer

FaceDeer@fedia.io · 21 hours ago

It’s not wrong for either to draw inspiration from the other. It’s the hypocrisy that’s wrong.

FaceDeer@fedia.io · 2 days ago

They’re not both true, though. It’s actually perfectly fine for a new dataset to contain AI generated content. Especially when it’s mixed in with non-AI-generated content. It can even be better in some circumstances, that’s what “synthetic data” is all about.

The various experiments demonstrating model collapse have to go out of their way to make it happen, by deliberately recycling model outputs over and over without using any of the methods that real-world AI trainers use to ensure that it doesn’t happen. As I said, real-world AI trainers are actually quite knowledgeable about this stuff, model collapse isn’t some surprising new development that they’re helpless in the face of. It’s just another factor to include in the criteria for curating training data sets. It’s already a “solved” problem.

The reason these articles keep coming around is that there are a lot of people that don’t want it to be a solved problem, and love clicking on headlines that say it isn’t. I guess if it makes them feel better they can go ahead and keep doing that, but supposedly this is a technology community and I would expect there to be some interest in the underlying truth of the matter.

FaceDeer@fedia.io · 2 days ago

No, researchers in the field knew about this potential problem ages ago. It’s easy enough to work around and prevent.

People who are just on the lookout for the latest “aha, AI bad!” Headline, on the other hand, discover this every couple of months.

FaceDeer@fedia.io · 2 days ago

AI already long ago stopped being trained on any old random stuff that came along off the web. Training data is carefully curated and processed these days. Much of it is synthetic, in fact.

These breathless articles about model collapse dooming AI are like discovering that the sun sets at night and declaring solar power to be doomed. The people working on this stuff know about it already and long ago worked around it.