Home›Blog›Reddit can sue AI scrapers under copyright law. Google could not.

Reddit can sue AI scrapers under copyright law. Google could not.

Illustration of two courthouse outlines with opposite verdict marks, connected by a single line to one shared scraper icon between them
Two federal courts considered near-identical claims against the same scraper eleven days apart and reached opposite results. (Illustrative)

On 20 July 2026 a federal court in California threw out Google's attempt to use copyright law to stop a company scraping its search results. On 31 July a federal court in New York allowed Reddit to press an almost identical claim against the same scraper, and against the AI search company that allegedly bought the data. The anti-scraping system at the centre of both cases was the same: Google's own SearchGuard. The deciding difference is that Reddit owns, or claims to own, the content sitting inside those search results and Google does not. For anyone publishing material that ends up feeding AI answers, that distinction matters more than either headline.

This piece reflects reporting as of August 2026. Both rulings are early-stage decisions on motions to dismiss and neither has been tested on appeal.


Diagram showing search results flowing from a publisher through a search engine to a scraper, with the copyright ownership marker sitting on the publisher
The same anti-scraping system sat at the centre of both cases. The difference is that the publisher owns the underlying content and the search engine does not. (Illustrative)

The two cases, briefly

Both revolve around a technique that has become central to AI data collection: rather than scraping a website directly, where the site can block you, you scrape that site's content out of Google's search results instead.

Google v SerpApi. Google sued the scraping service SerpApi, alleging it had bypassed SearchGuard, Google's anti-scraping protection, to collect and resell search results. The legal theory was the Digital Millennium Copyright Act's anti-circumvention rules, which make it unlawful to get around a technical measure that controls access to a copyrighted work. Google filed in December 2025. On 20 July, Chief Judge Yvonne Gonzalez Rogers in the US District Court for the Northern District of California granted SerpApi's motion to dismiss.

Reddit v Perplexity, SerpApi and others. Reddit sued in October 2025, naming SerpApi, the proxy providers Oxylabs and AWMProxy, and Perplexity. It alleges that scraping firms harvested its content at industrial scale by pulling it out of Google search results, disguising their identity to get around the protections in their way, and selling the resulting data on to the AI answer engine Perplexity. On 31 July, Judge Paul A. Engelmayer in the Southern District of New York largely denied the motions to dismiss, letting the core anti-circumvention claims proceed.

Why Google lost

Google's problem was ownership, in two layers.

First, where a set of search results contains no copyrighted content at all, there is nothing for an anti-scraping system to be controlling access to in the copyright sense. URLs, short snippets and factual index data are, in substance, facts about where things are on the internet, and copyright does not protect facts. Google's own complaint described its results as a mixture of copyrighted and uncopyrighted material, and that concession sank this part of the case. Those claims were dismissed with prejudice, meaning Google cannot refile them.

Second, and more consequentially, for the results that do contain someone's copyrighted material, Google was in the wrong position to assert the protection. The provisions only reach a measure that controls access to a copyrighted work with the authority of the copyright owner. Google does not own most of the web pages it indexes, nor the licensed images that appear in its Knowledge Panels. The court found that Google had pleaded no facts showing those copyright holders had authorised it to deploy SearchGuard on their behalf, and it declined to infer that authority from display licences whose terms Google never set out. That narrower Knowledge Panel claim was dismissed with 21 days to amend.

One thing the court did not do is throw Google out for lack of standing. It expressly rejected SerpApi's argument that only copyright owners may sue, holding that the statute reaches any person injured by a violation. Google's problem was not that it had no right to be in court. It was that it had not pleaded the facts this particular claim requires.

Commentators were blunt about the wider principle. Google was, in effect, trying to use copyright law to enforce a business preference about who gets to read its pages.


Diagram of a legal case progressing from filing through motion to dismiss to trial, with only the first two stages marked complete
Surviving a motion to dismiss means the allegations are plausible enough to continue. It is an early stage, not a finding that Reddit is right. (Illustrative)

Why Reddit did not

Reddit is in the position Google was not. The material at issue is content posted on Reddit, and Reddit points both to its terms and to the substantial body of posts its own staff wrote. That claim is contested: SerpApi argued Reddit has no copyright interest in what its users post. But at this stage the judge only had to decide whether the allegation was plausible, and he found it was.

Judge Engelmayer's reasoning tracked that directly. He found Reddit's injuries fell inside what the DMCA's anti-circumvention rules were written to address, describing the platform as an example of the global digital online marketplace for copyrighted works the statute set out to promote. He also found that Google's SearchGuard does qualify as a measure controlling access to a work, which is what a claim under section 1201(a) requires.

That last point is the strange, elegant part. The same technical system Google could not successfully sue over became the basis of a viable claim for Reddit, because Reddit can point to a copyright interest that Google was missing.

Reddit did not get everything. The judge dismissed a separate trafficking claim under section 1201(b) against SerpApi, along with unfair competition and unjust enrichment claims against both defendants. What survived is the core allegation.

What this does and does not settle

It is worth being precise, because a lot of coverage is not.

  • Nothing has been decided on the merits. Denying a motion to dismiss means the court accepts that, if the facts are as alleged, there is a case to answer. Reddit still has to prove it.
  • These rulings do not formally conflict. They come from different district courts, neither of which binds the other, and they turned on different plaintiffs in different postures. There is no circuit split, and no appellate ruling.
  • The rule that emerges is narrow but useful. Read together, the two decisions suggest the DMCA's anti-circumvention provisions can be a live tool against scraping, but principally in the hands of whoever owns the content. Intermediaries who merely display other people's material are on much weaker ground.

Why this matters if you publish anything

The practical question for most people is simple: if an AI company is ingesting your work, do you have any lever at all?

Most of the copyright litigation against AI companies has been fought over training and output. This is a different theory. It is not about whether ingesting your writing to train a model is fair use. It is about whether someone broke a lock to get at it. That is a narrower, more technical question, and narrower questions are often easier to win.

The practical catch is that the theory needs a lock. Reddit's claim works partly because there were technical measures being deliberately evaded, with identity-masking alleged as part of the evasion. A site with no access controls, no blocks and no terms being circumvented has much less to point at. That is an argument for the unglamorous housekeeping: setting your robots directives deliberately, using a crawler-blocking service if scraping is a real cost to you, and keeping a record of what you blocked and when.

It is also worth noticing what these cases say about the shape of the AI data market. Nobody in the Reddit case is accused of scraping Reddit directly. The alleged route runs through Google, via intermediaries, to an AI company. Blocking the crawlers you can see does not necessarily stop your content arriving somewhere else by a longer road.

Frequently asked questions

Has Reddit won its case against Perplexity?

No. It has survived an attempt to have the case thrown out at the earliest stage. The core claims proceed to litigation, where Reddit will have to prove what it has alleged.

Why could Google not use the same argument?

Because Google does not own most of the content in its search results, and it did not allege that the copyright owners had authorised it to deploy an access control on their behalf. The court also held that where results contain no copyrighted content at all, there is no protected work for the DMCA to attach to. It did not, however, rule that Google lacked the right to sue.

Does this mean scraping search results is illegal now?

No. One court found that Google could not stop it on these grounds, and another found that a content owner could plausibly bring such a claim. Neither is a general rule, and neither has been through an appeal.

Is this the same as the AI training copyright cases?

No. Those ask whether using copyrighted work to train a model is infringement or fair use. This asks whether someone unlawfully got around a technical access control. They can succeed or fail independently.

Does any of this apply in the UK?

Not directly. The DMCA is US law. The UK has its own provisions on circumventing technological protection measures under the Copyright, Designs and Patents Act 1988. On text and data mining, the government's consultation is finished: it reported in March 2026 that it was dropping the broad exception with opt-out it had originally favoured, leaving existing copyright law in place for now. The commercial reality still reaches UK publishers, because the AI services and the scrapers largely operate globally.

The honest takeaway

The interesting thing here is not that Reddit had a good day in court. It is that the same anti-scraping system produced a losing claim for the company that built it and a surviving claim for the company whose content it protects. If that reasoning holds up, the practical lesson for publishers is that the strongest position is being the copyright owner with a documented technical measure that someone deliberately worked around.

It is early. A motion to dismiss is the lowest bar in litigation, no appellate court has looked at any of this, and the eventual answer may come from legislation rather than judges. But after a couple of years in which content owners have had very little that works against AI scraping, this is the first ruling that points at something that might.

Sources

Enjoyed this? Get the weekly roundup:
← Back to blog