Bluesky users debate plans around user data and AI training

Date:

Share post:


Social network Bluesky recently published a proposal on GitHub outlining new options it could give users to indicate whether they want their posts and data to be scraped for things like generative AI training and public archiving.

CEO Jay Graber discussed the proposal earlier this week, while on-stage at South by Southwest, but it attracted fresh attention on Friday night, after she posted about it on Bluesky. Some users reacted with alarm to the company’s plans, which they saw as a reversal of Bluesky’s previous insistence that it won’t sell user data to advertisers and won’t train AI on user posts.

“Oh, hell no!” the user Sketchette wrote. “The beauty of this platform was the NOT sharing of information. Especially gen AI. Don’t you cave now.”

Graber replied that generative AI companies are “already scraping public data from across the web,” including from Bluesky, since “everything on Bluesky is public like a website is public.” So she said Bluesky is trying to create a “new standard” to govern that scraping, similar to the robots.txt file that websites use to communicate their permissions to web crawlers.

Debates about AI training and copyright have dragged robots.txt into the spotlight, among other things highlighting the fact that it’s not legally enforceable. Bluesky frames its proposed standard as one that would have a similar “mechanism and expectations,” providing “a machine-readable format, which good actors are expected to abide, and does carry ethical weight, but is not legally enforceable.”

Under the proposal, users of the Bluesky app, or other apps that use the underlying ATProtocol, could go into their settings and allow or disallow the usage of their Bluesky data across four categories: generative AI, protocol bridging (i.e., connecting different social ecosystems), bulk datasets, and web archiving (such as the Internet Archive’s Wayback Machine).

If a user indicates that they don’t want their data used to train generative AI, the proposal says, “Companies and research teams building AI training sets are expected to respect this intent when they see it, either when scraping websites, or doing bulk transfers using the protocol itself.”

Molly White, who writes the Citation Needed newsletter and Web3 is Going Just Great blog, described this as “a good proposal,” and said it was “weird to see people flaming BlueSky for it,” since it’s not so much “welcoming in AI scraping” but rather “trying to add a consent signal to allow users to communicate preferences for the scraping that is already happening.”

“I think the weakness with this and [Creative Commons’] similar proposal for ‘preference signals’ is that they rely on scrapers to respect these signals out of some desire to be good actors,” White continued. “We’ve already seen some of these companies blow right past robots.txt or pirate material to scrape.”



Source link

Lisa Holden
Lisa Holden
Lisa Holden is a news writer for LinkDaddy News. She writes health, sport, tech, and more. Some of her favorite topics include the latest trends in fitness and wellness, the best ways to use technology to improve your life, and the latest developments in medical research.

Recent posts

Related articles

Amazon’s Echo will send all voice recordings to the cloud, starting March 28

Amazon Echo users will no longer have the option to process their Alexa requests locally, which means...

Week in Review: SXSW week comes to a close

Welcome back to Week in Review! I’m Karyne Levy, TechCrunch’s deputy managing editor, and I’ll be writing...

SpaceX launches astronauts for long-awaited International Space Station crew swap

SpaceX successfully launched four people into space on Friday, beginning a mission that will give the International...

Skype is shutting down in May — these are the best alternatives

After 23 years of connecting people around the world, Skype, the popular video-calling service, is shutting down....

Republican Congressman Jim Jordan asks Big Tech if Biden tried to censor AI

On Thursday, House Judiciary Chair Jim Jordan (R-OH) sent letters to 16 American technology firms, including Google...

Bench is charging people for services they already paid for, some customers say

After Employer.com acquired bankrupt accounting startup Bench in a fire-sale late last year, CEO Jesse Tinsley pledged...

AI coding assistant Cursor reportedly tells a ‘vibe coder’ to write his own damn code

As businesses race to replace humans with AI “agents,” coding assistant Cursor may have given us a...

Profitable Klarna files for a potentially blockbuster IPO

Swedish fintech Klarna took the next step in its highly anticipated U.S. IPO on Friday when it...