Skip to main content
Back to News
news/AI Policy

Reddit's Copyright Case Against Perplexity Clears Key Hurdle

A federal judge rejected Perplexity's bid to dismiss Reddit's copyright lawsuit, clearing a path to trial that could reshape how AI firms collect training data.

Stefan Trbojevic

Stefan Trbojevic

2 August 20262 min read
LinkedIn
Editorial illustration for Reddit copyright lawsuit against Perplexity AI - courtroom meets tech aesthetic with Reddit orange branding

The takeaway

Courts are drawing a line between legitimate web scraping and coordinated circumvention of platform protections. For AI builders, the message is clear: how you collect data matters as much as what you collect.

Why it matters for builders

The ruling signals that courts will scrutinize not just what data AI companies use, but how they obtain it. Companies that build data pipelines around platform anti-scraping measures face growing legal risk. The line between enterprise data tools and adversarial infrastructure is blurring fast.

Reddit's Copyright Case Against Perplexity Clears Key Hurdle

A federal judge has rejected Perplexity's motion to dismiss Reddit's copyright infringement lawsuit, allowing the closely watched case to move forward and setting up a potential landmark ruling on how AI companies can use publicly available web content for training.

The ruling, delivered on July 31 in New York federal court, keeps alive Reddit's allegations that Perplexity and three data-scraping services conspired to vacuum up millions of user posts without permission. The named co-defendants include Lithuanian data scraper Oxylabs, proxy service AWMProxy, and Texas-based SerpApi, which sells search engine scraping tools.

What happened

Reddit originally filed the lawsuit in October 2025, accusing Perplexity of bypassing the platform's anti-scraping protections by routing data collection through Google search results. According to the complaint, when direct scraping of Reddit proved too difficult, the defendants pivoted to extracting Reddit content from Google's search engine results pages — a workaround Reddit argues violates the Digital Millennium Copyright Act and laws against unfair trade practices.

The judge's decision to let the case proceed is significant because it rejects Perplexity's argument that scraping publicly accessible web content is fundamentally lawful. The ruling signals that courts are willing to examine not just what data AI companies collect, but how they collect it.

Why it matters

Reddit has positioned itself as one of the most aggressive defenders of user-generated content in the AI era. The company has signed lucrative data licensing deals with Google and others while simultaneously blocking crawlers from OpenAI, Microsoft, and any company that hasn't paid for access. This dual strategy — monetize through partnerships, litigate against free-riders — is becoming a template for other content platforms.

Ben Lee, Reddit's chief legal officer, told The Verge the ruling "brings us one step closer to holding bad actors accountable," adding that the company "supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission."

The case arrives amid a broader reckoning over training data. The New York Times is suing OpenAI and Microsoft. Getty Images sued Stability AI. Record labels have targeted Suno and Udio. What makes the Reddit case different is its focus on the method of collection — the alleged conspiracy to route around technical barriers — rather than just the fact of ingestion.

For AI builders, the implications are clear: the era of "scrape everything and sort out the legalities later" is ending. Courts are drawing lines between responsible data access and industrial-scale circumvention of platform protections.

Data flow diagram showing how Perplexity allegedly routed around Reddit anti-scraping protections

The three data-scraping co-defendants make the case particularly instructive. Oxylabs operates one of the largest proxy networks for web data extraction. AWMProxy has been described in court filings as a "former Russian botnet" turned commercial service. SerpApi is a legitimate Texas company selling search engine scraping APIs. The mix of actors shows how AI data supply chains can blur the line between enterprise tools and outright adversarial infrastructure.

The case now enters discovery, where both sides will exchange evidence about exactly how Perplexity's data pipeline operated. If Reddit can prove a coordinated scheme to circumvent its technical protections, the precedent would reach far beyond this single lawsuit — potentially reshaping how every AI company approaches web-scale data collection.

Share𝕏

The Automation Brief

Read 5 AI stories instead of 50.

The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.

No noise. Unsubscribe anytime.

Editorial notes

Reported by

Stefan Trbojevic

Edited by

n8n Lab Editorial

Published

2 August 2026

Updated

2 August 2026

AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.

n8n Lab is an independent service provider. We are not affiliated with, endorsed by, or sponsored by n8n GmbH. “n8n” is a trademark of n8n GmbH and is used here only to describe the platform-specific implementation and automation services we provide.