Skip to main content

issue #289 of brennan.day

LLMs Democratize Extraction and Scale Malice

In a post a few months ago, Tante wrote that "'AI' exists to disenfranchise labor," and earlier today, Coyote posted a short essay on what data ownership truly means for the IndieWeb when we're subject to genAI scrapers, extractors, and bad-faith actors.

On a similar note, maybe your inbox has recently been flooded by iLands genAI agents, as Gordon and cleberg have recently written about. Or maybe you're an administrator for a Mastodon instance that's been flooded with users who turn out to be AI agents.

And this is happening at the same time that Anthropic researchers warn AI could cause human extinction by 2030. This is happening at the same time that Nvidia CEO, Jensen Huang, announces that Artificial General Intelligence has arrived. Which, I'm sure, has nothing to do with him running the only company that's been profitable due to genAI.

Hm.

This post began as a response to Coyote's short essay, about how we need "personal control" and an "assertion of boundaries" for our data. That we need to practice "resistance over infinite intrusion" and "autonomy and solidarity" in response.

I agree, but I'm not sure I have an answer as to how, and the problem seems far more existential than just that. I frankly admit that my earlier philosophy and tactic of apathy and ignorance are no longer sufficient.

Years ago, when the IndieWeb movement was founded, people had the idea that having their own website would prevent the extractive, surveillance-style dark patterns that corporate social media use. You would no longer have to worry about the black box of scripts, cookies, pixel trackers, and HTML fingerprints building a profile of you that would be bought and sold across various marketers (and governments).

And for a while, I think that was true. Sure, shadow profiles exist, where your data is collected by other means and used even if you don't have an account, but you were still creating on a sovereign Internet.

But now, you don't need to be a multi-billion dollar corporation. You can deploy a genAI agent to scrape and extract data from any website, and use that data to try to make money yourself. What used to require software engineers and marketers and addiction psychologists is now ubiquitous and accessible to anyone (who's willing to pay the upfront cost of tokens).

This has been a large part as to why I've become interested in the Gemini Project, as I believed this could be a way to continue cultivating a more-human Internet. Unfortunately, I learned only yesterday about Gopher-MCP, designed to allow genAI agents to interface with Gopher and Gemini protocols.

I've seen creative ways people have tried to make their websites resistant to genAI beyond adding a robots.txt file. Be it having a tricky-to-navigate false homepage, eliminating easy extraction methods like disabling RSS, or creating honeypots at certain URLs (which I've done myself).

Alas, all of these are band-aid solutions to a systemic problem. I'm all for open culture, and my blog posts are in the Creative Commons, but are genAI and other content mills honouring the required ShareAlike clause of my work? No, of course not.

This extraction is nothing new. What's novel here is the scale and, again, accessibility. Nearly everyone on an IndieWeb forum I browse has been spammed by iLands agents because someone decided to scrape the public boards of the forum; so long as people have the money, and so long as the hardware exists, we will only get more proliferation of this annoying grey goo-esque slop of incompetent spam and data extraction.

AI 2027

At the start of April last year, a website was published up called AI 2027 speculating on the cascading, catastrophic consequences we will globally face due to the accelerating developments in generative AI. It's certainly an interesting read, especially now that we're nearly at the end of 2026, to see what was guessed at correctly and incorrectly so far.

When we take a step back and examine what's been going on, and what narratives we're being told about generative artificial intelligence, both in its current form and the upcoming delta of change, there are two realities:

  1. The first is that we are still working with the same frustrating, irritating slop; these LLMs seem stuck and stagnant, and that it's only becoming more capable at spreading and more ubiquitous, but still often fails the simplest tasks (those iLands spam emails, for example, don't actually provide a way for people to actually pay them money). That the only tangible difference between loading up ChatGPT and deploying an expensive agent is that the agent has far more vectors to be annoying. If you've had the displeasure of reading any messages written by agents, you understand what I mean.
  2. The second is that we are on the cusp of collective artificial intelligence becoming autonomous, while also becoming so exponentially more intelligent than humanity that it could be apocalyptic.

I think anybody who's been able to not fall into the throes of genAI (psychosis or otherwise) can see that these models are still just large language models. The fundamental mechanics of how they're built and operate greatly limit their capabilities. There is (thankfully) a dark, impossible ceiling that cannot be breached with this current iteration of artificial intelligence. But that has not deterred many people with a lot of money and power. The data centres making people sick are still being constructed, and every product and service still continues to inappropriately shove genAI into the faces of consumers and end-users.

And we have made innovations, regardless of these limitations, just different ones. GenAI is no longer limited to chat windows and API calls; it's now given root access with destructive consequences. "Swarms" of agents are using the Internet in unpredictable ways, costing huge amounts of tokens, which is a euphemism for money and energy lost.

And this is why I initially thought patience would win out, given how obviously foolish and expensive this endeavour continues to be. And yet it continues and escalates and only has gotten worse with time.

What Do We Do?

Tools already exist to make extraction more expensive—the community-maintained ai.robots.txt list is a good start. If your work is visual, Glaze and Nightshade, out of the University of Chicago, let you cloak or poison images. Spawning's Have I Been Trained registry can help you opt out of extraction (though it's currently under maintenance).

Individual bloggers and artists can only do so much, though. I look to the workers who are being pushed to build and deploy the tools hollowing out humanity: The Tech Workers Coalition's workersdecide.tech project is a set of resources for organizing against AI mandates at work. The Alphabet Workers Union-CWA has an AI committee fighting for protections. Groups like the Future of Life Institute and the Distributed AI Research Institute are building a public case for regulation and centering the people AI harms most.

If you're someone who's found themselves dependent on genAI, know that you're not alone and it isn't your fault. The sycophantic nature of these LLMs is designed to pull you in. There are recovery collectives and the best thing you can do for yourself is come clean and admit there's a problem.

Finally, the CARE Principles for Indigenous Data Governance (Collective Benefit, Authority to Control, Responsibility, Ethics) and, specifically for First Nations, the OCAP principles (Ownership, Control, Access, Possession) exist because generations of data have been taken from our communities without consent, and used against us.

I love people, and I know people are inherently good. But that does not mean good people will act when they need to. As important as reclaiming joy and preserving our hope are to activism, we need more than that. We need direct action, we need righteous fury to be channelled. We need people willing to sacrifice and risk themselves for the next generation. There seems to be this epidemic of the bystander effect. People need to remember that polite revolution is a non-existent oxymoron.

I think of the Audre Lorde quote, how "the master's tools will never dismantle the master's house." We must stop using these products, and urge others to as well. And those who work for companies that have dystopian genAI quotas must speak up.

Publication Information

Title: LLMs Democratize Extraction and Scale Malice

Author: Brennan Kenneth Brown ORCID

Issue: #289 of brennan.day

License: CC BY-SA 4.0

Date Published:

Full URL: https://brennan.day/llms-democratize-extraction-and-scale-malice/

Place of Publication: Calgary, Alberta, Canada

Contact: mail@brennanbrown.ca

Cite this issue
  • MLA 9
    Brown, Brennan Kenneth. "LLMs Democratize Extraction and Scale Malice." brennan.day, 20 Sept. 2026, https://brennan.day/llms-democratize-extraction-and-scale-malice/.
  • APA 7
    Brown, B. K. (2026, September 20). LLMs Democratize Extraction and Scale Malice. brennan.day. https://brennan.day/llms-democratize-extraction-and-scale-malice/
  • Chicago (Notes-Bib)
    Brown, Brennan Kenneth. "LLMs Democratize Extraction and Scale Malice." brennan.day (blog). September 20, 2026. https://brennan.day/llms-democratize-extraction-and-scale-malice/.
  • BibTeX
    @online{brennan2026llms-democratize-extraction-and-scale-malice,
      author = {{Brown, Brennan Kenneth}},
      title = {{LLMs Democratize Extraction and Scale Malice}},
      year = {2026},
      url = {https://brennan.day/llms-democratize-extraction-and-scale-malice/},
      urldate = {2026-09-20}
    }

Comments

To comment, please sign in with your website:

How it works: Enter your website URL. IndieLogin checks your rel="me" links on GitHub, GitLab, Codeberg, and email. Setup instructions.

Honestly, it was inevitable even without LLMs. Bottomfeeders that are just a touch more tech literate than John Q. Public will always go for the desperate plays, and they'll do it with off-the-shelf tech at its default setting. If anything, it's an indicator to how wide the IndieWeb has grown: it's now seen as a viable market for scammers! I saw someone post a picture of their dense spam folder, asking if this is normal when you start a website. And... yeah; even for a small website, any public-facing email address associated with it will be a target for spam. It's been that way for as long as I can remember, at least. After all, the spam folder exists for a reason.
Honestly, it was inevitable even without LLMs. Bottomfeeders that are just a touch more tech literate than John Q. Public will always go for the desperate plays, and they'll do it with off-the-shelf tech at its default setting. If anything, it's an indicator to how wide the IndieWeb has grown: it's now seen as a viable market for scammers! I saw someone post a picture of their dense spam folder, asking if this is normal when you start a website. And... yeah; even for a small website, any public-facing email address associated with it will be a target for spam. It's been that way for as long as I can remember, at least. After all, the spam folder exists for a reason.

Webmentions

1 Repost


Related Posts

↑ TOP