Stack Overflow bans users en masse for rebelling against OpenAI partnership — users banned for deleting answers to prevent them being used to train ChatGPT

misk@sopuli.xyz · 6 months ago

Stack Overflow bans users en masse for rebelling against OpenAI partnership — users banned for deleting answers to prevent them being used to train ChatGPT

Hypx@fedia.io · 6 months ago

Eventually, we will need a fediverse version of StackOverflow, Quora, etc.

Thomas@discuss.tchncs.de · 6 months ago

Those would be harvested to train LLMs even without asking first. 😐

sramder@lemmy.world · 6 months ago

At this point I’m assuming most if not all of these content deals are essentially retroactive. They already scrapped the content and found it useful enough to try and secure future use, or at least exclude competitors.

Ricky Rigatoni@lemm.ee · 6 months ago

They scraped the content, liked the results, and are only making these deals because it’s cheaper than getting sued.

AeroLemming@lemm.ee · edit-2 2 months ago

deleted by creator

linearchaos@lemmy.world · 6 months ago

Honestly? I’m down with that. And when the LLM’s end up pricing themselves out of usefulness, we’ll still have the fediverse version. Having free sites on the net with solid crowd-sourced information is never a bad thing even if other people pick up the data and use it.

It’s when private sites like Duolingo and Reddit crowd source the information and then slowly crank down the free aspect that we have the problems.

The Ad sponsored web model is not viable forever.

bort@sopuli.xyz · 6 months ago

The Ad sponsored web model is not viable forever.

a thousand times this

danc4498@lemmy.world · 6 months ago

I’d rather the harvesting be open to all than only the company hosting it.

mox@lemmy.sdf.org · 6 months ago

Assuming the federated version allowed contributor-chosen licenses (similar to GitHub), any harvesting in violation of the license would be subject to legal action.

Contrast that with Stack Exchange, where I assume the terms dictated by Stack Exchange deprive contributors of recourse.

chameleon@kbin.social · 6 months ago

SO already was. Not even harvested as much as handed to them. Periodic data dumps and a general forced commitment to open information were a big part of the reason they won out over other sites that used to compete with them. SO most likely wouldn’t have existed if Experts Exchange didn’t paywall their entire site.

As with everything else, AI companies believe their training data operates under fair use, so they will discard the CC-SA-4.0 license requirements regardless of whether this deal exists. (And if a court ever finds it’s not fair use, they are so many layers of fucked that this situation won’t even register.)

Rolando@lemmy.world · 6 months ago

But users and instances would be able to state that they do not want their content commercialized. On StackOverflow you have no control over that.

ArbitraryValue@sh.itjust.works · 6 months ago

You can state what you don’t want, but no one will be paying attention. Except maybe the LLM reading your posts…

pivot_root@lemmy.world · 6 months ago

Yup. Laws are only suggestions until you get caught.

ArbitraryValue@sh.itjust.works · edit-2 6 months ago

I suspect it isn’t even illegal, but I’m not an expert.

thejml@lemm.ee · 6 months ago

Not fediverse, but open-source and community run: https://codidact.com

Avid Amoeba@lemmy.ca · edit-2 6 months ago

Oh this looks decent. British non-profit, I like it. Registering.

linearchaos@lemmy.world · 6 months ago

Smells too much like duo-lingo. Here, everyone jump in and answers all the questions. 5 years later, ohh look at this gold mine of community data we own…

residentmarchant@lemmy.world · 6 months ago

This was actually the whole original point of Duolingo. The founder previously created Recaptcha to crowd source machine vision of scanned books.

His whole thing is crowd sourcing difficult tasks that machines struggle with by providing some sort of reason to do it (prevent spam at first and learn a language now)

From what I understand Duolingo just got too popular and the subscription service they offer made them enough money to be happy with.

BraveLittleToaster@lemmy.world · 6 months ago

Everything you write on here is public. There’s nothing stopping anyone from using that data for training

linearchaos@lemmy.world · 6 months ago

We needed it a few years ago.

NoIWontPickAName@kbin.earth · 6 months ago

Can we pass on quora?

Ricky Rigatoni@lemm.ee · 6 months ago

Federated yahoo answers.

brbposting@sh.itjust.works · 6 months ago

how is feddi formed

HowManyNimons@lemmy.world · 6 months ago

Arguably, they need to do way instain mother> who kill thier babbys. becuse these babby cant frigth back?

It’s important to remember that it was on the news this mroing a mother in ar who had kill her three kids.

NoIWontPickAName@kbin.earth · 6 months ago

Too much, can’t figure it out

Avid Amoeba@lemmy.ca · 6 months ago

We already have the SO data. We could populate such a tool with it and start from there.