The takeaway
Courts are drawing a line between legitimate web scraping and coordinated circumvention of platform protections. For AI builders, the message is clear: how you collect data matters as much as what you collect.
Why it matters for builders
The ruling signals that courts will scrutinize not just what data AI companies use, but how they obtain it. Companies that build data pipelines around platform anti-scraping measures face growing legal risk. The line between enterprise data tools and adversarial infrastructure is blurring fast.
Reddit's Copyright Case Against Perplexity Clears Key Hurdle
A federal judge has rejected Perplexity's motion to dismiss Reddit's copyright infringement lawsuit, allowing the closely watched case to move forward and setting up a potential landmark ruling on how AI companies can use publicly available web content for training.
The ruling, delivered on July 31 in New York federal court, keeps alive Reddit's allegations that Perplexity and three data-scraping services conspired to vacuum up millions of user posts without permission. The named co-defendants include Lithuanian data scraper Oxylabs, proxy service AWMProxy, and Texas-based SerpApi, which sells search engine scraping tools.
What happened
Reddit originally filed the lawsuit in October 2025, accusing Perplexity of bypassing the platform's anti-scraping protections by routing data collection through Google search results. According to the complaint, when direct scraping of Reddit proved too difficult, the defendants pivoted to extracting Reddit content from Google's search engine results pages — a workaround Reddit argues violates the Digital Millennium Copyright Act and laws against unfair trade practices.
The judge's decision to let the case proceed is significant because it rejects Perplexity's argument that scraping publicly accessible web content is fundamentally lawful. The ruling signals that courts are willing to examine not just what data AI companies collect, but how they collect it.
Why it matters
Reddit has positioned itself as one of the most aggressive defenders of user-generated content in the AI era. The company has signed lucrative data licensing deals with Google and others while simultaneously blocking crawlers from OpenAI, Microsoft, and any company that hasn't paid for access. This dual strategy — monetize through partnerships, litigate against free-riders — is becoming a template for other content platforms.
Ben Lee, Reddit's chief legal officer, told The Verge the ruling "brings us one step closer to holding bad actors accountable," adding that the company "supports responsible access to public content, but we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission."
The case arrives amid a broader reckoning over training data. The New York Times is suing OpenAI and Microsoft. Getty Images sued Stability AI. Record labels have targeted Suno and Udio. What makes the Reddit case different is its focus on the method of collection — the alleged conspiracy to route around technical barriers — rather than just the fact of ingestion.
For AI builders, the implications are clear: the era of "scrape everything and sort out the legalities later" is ending. Courts are drawing lines between responsible data access and industrial-scale circumvention of platform protections.

The three data-scraping co-defendants make the case particularly instructive. Oxylabs operates one of the largest proxy networks for web data extraction. AWMProxy has been described in court filings as a "former Russian botnet" turned commercial service. SerpApi is a legitimate Texas company selling search engine scraping APIs. The mix of actors shows how AI data supply chains can blur the line between enterprise tools and outright adversarial infrastructure.
The case now enters discovery, where both sides will exchange evidence about exactly how Perplexity's data pipeline operated. If Reddit can prove a coordinated scheme to circumvent its technical protections, the precedent would reach far beyond this single lawsuit — potentially reshaping how every AI company approaches web-scale data collection.
The Automation Brief
Read 5 AI stories instead of 50.
The essential moves in AI agents, models, automation and infrastructure — filtered for builders and operators, with the part that actually matters.
No noise. Unsubscribe anytime.
Editorial notes
Stefan Trbojevic
n8n Lab Editorial
2 August 2026
2 August 2026
Sources
AI disclosure: AI assisted with research and drafting. Factual claims are reviewed by an editor.




