How to Market to Expensive Keywords
How to Market to Expensive Keywords
12 Steps to Create Videos

One million public Bluesky posts scraped for AI training [Video]

Categories
AI Behavioral Targeting

Bluesky is already facing its first major AI scrape, despite the stance of its owners that it will never train generative AI on user data.

Reported by 404Media on Nov. 26, one million public Bluesky posts — complete with identifying user information — were crawled and then uploaded to AI company Hugging Face. The dataset was created by machine learning librarian Daniel van Strien, intended to be used in the development of language models and natural language processing, as well as general analysis of social media trends, content moderation, and posting patterns. It contains users’ decentralized identifiers (DIDs) and even has a search function to find content from specific users.

According to the dataset’s description, the set “contains 1 million public posts collected from Bluesky Social’s firehose API (Application Programming Interface), intended for machine learning research and experimentation with social media data. Each post contains text content, metadata, and information about media attachments and reply relationships.”

Mashable Light Speed

7 Invisible Obstacles to Digital Marketing Success
7 Invisible Obstacles to Digital Marketing Success
5 Steps to Creating Successful Ads