Archived
  • NetworkX
  • Gephi
  • Python
  • Social mining

Mastodon AI discourse

608 posts, 200 users, and a network too flat to hold a community

A social mining pipeline over Mastodon's AI conversation, measuring how information spreads across a decentralised network with almost no clustering.

Mastodon is federated, so nobody owns the graph. That makes it a better place than most to ask how a conversation actually travels.

Collecting two networks

608 AI-related posts came in through hashtag crawls — #ai, #machinelearning, #ChatGPT and neighbours — alongside a 200-user network, both through the Mastodon API with deduplication and JSON storage.

Those became two graphs. An information-diffusion network, where posts are nodes and replies and boosts are edges. And a friendship network, where users are nodes and mutual follows are edges.

The number that mattered

Average path length came out at 1.97. Clustering coefficient came out at roughly zero.

Together those say something specific: this is a network where anything can reach anything in about two hops, but where almost no one's contacts know each other. It is not a community. It is an audience arranged around a handful of hubs — PageRank put accounts like @Techcrunch and @machinelearning at the top, and their reach explains the short paths on its own.

What people were saying

Every post went through an LLM keyword extraction pass — DeepSeek R1 through OpenRouter — and the resulting frequency distribution followed a power law, mirroring the degree distribution of the network itself. ChatGPT, machine learning, ethics and automation dominated; everything else formed a long tail.

Analyzing Mastodon Social Media Data for Patterns about AI Technology Conversations — 3 pagesOpen the PDF