In short
Cara, the artist-focused platform built to resist AI training abuse, was scraped repeatedly in August 2026. In a surprising turn, the person behind the first scrape is now helping creator Jingna Zhang build Lantern, an open-source tool to detect artist scraping in public AI datasets.
- Cara was hit by three major scraping incidents in August 2026, including one archive of about 12 million artworks.
- The first scraper later apologized, deleted his dataset and began helping develop Lantern with Cara’s team.
- Lantern is designed to notify artists when their work appears in public AI datasets and support takedown efforts.
- The case highlights how hard it is to make public creative work truly unscriptable on the open web.
- Zhang says the broader policy debate must go beyond copyright and address consent, security and AI-era data extraction.
Cara, the artist-focused social platform built by photographer Jingna Zhang, was hit by a wave of scraping in August 2026 that exposed millions of works and renewed fears about how easily creative portfolios can be harvested for AI training. In an unexpected twist, the man behind the first and most disruptive scrape has apologized, deleted his dataset, and begun collaborating with Zhang on Lantern, an open-source tool designed to alert artists when their work appears in public AI datasets.
The episode has become a sharp example of the collision between generative AI, copyright disputes, and the limits of online security. It also shows how, even in communities created to avoid exploitation, artists can still be targeted by data harvesters who believe the law has not caught up with the technology.
What happened to Cara in August?
Three separate scraping incidents hit Cara over a short period beginning on August 13, overwhelming the platform’s small team and driving up server costs while unsettling users who had joined precisely to escape AI training pipelines tied to larger social platforms.
Cara was founded in early 2023 as a portfolio and image-sharing service for artists who wanted a space that would not feed their work into AI systems. It now has roughly 1.5 million artists on the platform, many of whom were drawn by its anti-scraping ethos and features such as AI-image filtering and Glaze, a protective tool intended to make style imitation harder for models that ingest web imagery.
But the platform’s safeguards were never designed to make scraping impossible. Like most sites on the open web, Cara can block some kinds of abuse, yet it cannot fully stop someone determined to copy public content at scale.
How severe was the first scrape?
The first incident was the most dramatic. A user posted a 12-terabyte archive containing about 12 million artworks from Cara’s public pages to the Reddit community r/DefendingAIArt, describing it as a low-cost project that reportedly cost less than $10 to assemble.
That archive represented nearly the entire public library of images available on the site, according to Zhang. She said the company learned about the scrape because users flagged the Reddit post after the scrapers themselves appeared to be boasting about the dataset and inviting others to use it.
Zhang said the episode felt targeted and deeply personal, especially because the legal protections against this kind of harvesting have not kept pace with the technology and can leave platforms and artists with limited options.
For artists who had moved to Cara to avoid exactly this outcome, the event landed as a betrayal. Some users reported panic, deleted their portfolios, or reconsidered whether posting online was worth the risk.
Why did the scraping spark such a backlash?
The backlash came from the fact that Cara was not a generic image board. It was built as a deliberate alternative to mainstream services where creators worry that posting art can also mean surrendering it to AI companies.
That context made the scraping feel like more than a technical violation. To many artists, it confirmed a broader fear: if a platform explicitly designed to resist AI harvesting can still be scraped, then no online venue can offer complete protection.
The incident also fed into an already heated legal and cultural debate. Zhang is involved in two class actions brought by visual artists. One targets Stability AI, Midjourney and other firms over alleged training on copyrighted images; the other names Google and focuses on image generator tools. Against that backdrop, the Cara scrapes became a live, on-the-ground version of the same conflict.
What did the second and third scrapes involve?
The second scrape reportedly pulled about 8.5 million links from Cara along with metadata including usernames, titles and tags, and those records were uploaded to Hugging Face, a platform widely used by AI developers to share datasets and models.
Hugging Face later said it would ask the user who uploaded the data, known as CaptiveDreamer, to remove personal metadata, but it would not remove the URLs themselves. Its reasoning was that the platform was not hosting copies of the art, only links to content already published on Cara.
A third scrape, on August 22, captured roughly 123,000 images along with text posts and user bios that included personal information. That material was then shared on Academic Torrents, another site used for file distribution.
| Event | Date | What was taken | Why it mattered |
|---|---|---|---|
| First major scrape | August 13, 2026 | 12 TB archive; about 12 million artworks | Captured nearly all public images on Cara and triggered the widest alarm |
| Second scrape | Mid-August 2026 | About 8.5 million links plus metadata | Raised questions about whether links and metadata alone can be used for training or targeting |
| Third scrape | August 22, 2026 | 123,000 images, text posts and user bios | Expanded concerns about exposure of personal information |
How did the scraper end up working with Cara?
Heft, a student in North America with a background in software and an interest in digital preservation, says the initial scrape began as a technical project and was never intended to be publicly weaponized. He asked to be identified by one of his screen names because he says he has received doxing attempts and death threats since the episode.
According to Heft, he made a poor choice when he posted the dataset on Reddit to provoke reaction. He says he knew the move would anger artists, but he did not appreciate the scale of the distress it would cause.
In later conversations with Zhang and Cara users, he says he was confronted with accounts of panic attacks, deleted portfolios and genuine fear about being exploited. That experience changed his view of the scrape, which he now calls cruel and thoughtless in hindsight.
Heft told WIRED that he understood the harm only after seeing how personally artists reacted, and that he had underestimated the emotional impact beyond simple anger at the scraping itself.
He also says he never believed the data would be useful for training a major AI model. In his view, 12 million images are not enough to train a modern image system, and the companies that build such models typically rely on enormous web-scale datasets already assembled by organizations such as LAION.
That may explain why he sees the bigger threat as not just model training, but the culture of casual harvesting itself: the notion that public posting means unconditional reuse.
Why did Zhang accept his help?
Zhang says Heft apologized, removed the dataset and then began helping the Cara team analyze weaknesses that make scraping possible. His technical advice has been valuable because he can show, often within minutes, how proposed protections can be bypassed.
For Zhang, the collaboration is not an endorsement of what happened. It is a pragmatic attempt to build something useful out of a harmful incident. She describes herself as an accidental tech founder, not someone who set out to run infrastructure for a security battle against scrapers.
The arrangement also highlights a practical reality in AI policy: some of the people best positioned to understand platform vulnerabilities are the ones who know how to exploit them.
What is Lantern and how would it work?
Lantern is an open-source tool being developed by Zhang and Heft to help artists discover whether their work has been absorbed into public AI datasets after the fact. Rather than trying to make scraping impossible, the tool aims to create a warning system and a paper trail.
Artists would create a one-way fingerprint of their images without uploading or storing the actual files on the platform. Lantern would then regularly scan newly released public AI datasets. If it found a match, it would notify the artist and provide a link to the dataset so the artist could request removal or pursue a takedown.
The project reflects a shift in strategy from prevention to detection. That distinction matters because total prevention is likely unrealistic on a public internet where any visible image can be copied with a few lines of code and a modest server bill.
What makes Lantern different from other protections?
Lantern is meant to work after scraping has already occurred, which may sound modest but is actually significant. Many existing protections are aimed at making training data less useful, disguising style, or erecting barriers that can be bypassed. Lantern instead focuses on visibility: knowing when a work appears in a dataset and having evidence to act on.
Because it is open source, other developers and advocacy groups can inspect, improve or adapt it. Zhang hopes that openness will make the tool more resilient and more widely adopted across art communities.
- One-way fingerprinting: Artists can tag images without storing copies on the platform.
- Dataset scanning: The system checks public AI training datasets as they appear.
- Match alerts: Artists are notified when their work is detected.
- Takedown support: Users get links and documentation to pursue removal requests.
- Open-source access: Anyone can contribute to the code or adapt it for new uses.
Why can’t Cara simply stop all scraping?
Cara can reduce abuse, but it cannot fully stop scraping because anything publicly accessible on the internet can be copied by a determined actor. The challenge is structural, not just technical.
Some measures may slow down automated collection, raise costs or make abuse easier to detect. But they do not change the basic reality that a public site must remain accessible to legitimate users, which leaves openings for malicious ones.
Zhang says Cara has already added temporary measures, including login gates, but she does not believe those fixes solve the underlying problem. A platform can make data collection more expensive, but it cannot make public content truly inaccessible without also degrading the user experience.
According to Zhang, Cara has taken reasonable steps within the limits of a usable social platform, but those steps cannot fully eliminate the risk of scrapers operating across the wider internet.
What do artists fear most now?
Many artists fear that moving to a different platform will not solve the problem. Zhang warns that bigger sites may actually attract more scraping because they contain more valuable data and offer larger returns to anyone building training sets.
That is one reason the controversy around Cara matters beyond one app. It shows that the incentive to scrape art is broad, persistent and not confined to a single community or platform.
For creators, the emotional injury is compounded by uncertainty. They do not always know where their work has gone, how it will be used or whether there is any realistic path to recovery once it has been collected.
How are companies and data platforms responding?
The responses from major platforms have been shaped by legal distinctions that often frustrate artists. Hugging Face’s reply to takedown requests is a good example: if a site hosts only links and not the artworks themselves, it may argue that copyright removal requests do not apply in the same way.
This creates a gray zone where metadata, URLs and references can still be useful to AI developers or data aggregators, even if the original files are elsewhere. For artists, the practical effect can feel the same as direct hosting because the data is still moving through systems that facilitate AI development.
Meanwhile, AI labs have little incentive to publicly discuss every dataset they scan or every source they might ingest. That opacity leaves creators with limited visibility into how far their work travels.
What role does law play here?
Law is central, but it is also lagging. Zhang argues that current rules do not adequately address the scraping of public creative work for AI training, especially when the copying is technically possible but ethically disputed.
That gap helps explain why the debate has spilled into lawsuits. Artists are testing whether copyright law can be used to challenge training practices, but those cases move slowly and cannot always keep pace with new scraping methods.
As the legal process unfolds, platforms like Cara are left to improvise protective tools, community policies and reactive enforcement. That is an unstable way to govern a problem that operates at internet scale.
What does this episode say about the future of AI and art?
The Cara episode suggests that the conflict between artists and AI developers is entering a phase where damage control may matter as much as prevention. If scraping is commonplace and large-scale model training continues to depend on web data, then creators will increasingly need tools that document exposure, support takedowns and make abuse visible.
It also suggests that some of the most productive responses may come from unexpected places. A person who once scraped art and posted it for shock value is now helping design a defense system. That does not erase the harm, but it does show how technical knowledge can be redirected.
For Zhang, the larger lesson is that AI-related attacks will only become more common. She believes the policy conversation needs to move beyond copyright alone and into broader questions of security, consent, platform design and the social cost of mass data extraction.
Zhang’s view is that the industry and policymakers must prepare for a world in which AI scraping is routine, and in which apologies or reversals are likely to remain the exception rather than the norm.
Timeline of the Cara scraping dispute
The story moved quickly from technical abuse to public controversy to uneasy collaboration. Here is a simplified timeline of the key events.
| Approximate date | Development | Significance |
|---|---|---|
| Early 2023 | Cara launches as an artist-first social and portfolio app | Offers an alternative to platforms tied to AI training data use |
| August 13, 2026 | First major scrape becomes public | 12 million works are exposed in a large archive posted online |
| Mid-August 2026 | Second scrape and metadata dump emerge | Questions grow around links, metadata and dataset sharing |
| August 22, 2026 | Third scrape reveals images and personal data | Heightens concern over safety and privacy |
| Late August 2026 | GoFundMe launched for legal fees; Lantern begins development | Defense shifts toward legal action and open-source detection |
What comes next for Cara and its users?
Cara is now trying to balance three things at once: preserving a usable platform, responding to hostile scraping and reassuring artists that their work will not be abandoned. The legal fund Zhang launched has already raised more than $100,000 toward a $120,000 target, signaling that the community sees the fight as both urgent and expensive.
But money alone will not solve the deeper issue. The internet’s architecture still favors copying, and AI development has increased the incentive to collect vast image libraries from everywhere they can be found.
That means the most realistic near-term outcome may be a combination of better alerts, more aggressive legal pressure, technical friction and collective pressure on the companies and platforms that benefit from scraped data.
For now, the Cara saga stands as a cautionary tale: even a platform created to protect artists can be mined at scale, and even the person who caused the harm can end up helping build the next line of defense.
Frequently asked questions
What happened to Cara in August 2026?
Cara was targeted by three major scraping incidents in August 2026, including one that exposed about 12 million public artworks and another that collected links and metadata. The attacks raised server costs, alarmed users and intensified worries about AI training abuse.
Who is helping Cara build Lantern now?
The first scraper, a student known as Heft, is helping develop Lantern after apologizing and deleting the dataset he posted. He has been advising Cara’s team on technical weaknesses and how to spot or document scraping after it happens.
What is Lantern supposed to do?
Lantern is an open-source tool that lets artists create a one-way fingerprint of their images and then scans public AI datasets for matches. If a match appears, the artist gets an alert and a link to the dataset so they can request removal or file a takedown.
Can Cara stop scraping completely?
No, not completely. Cara can add barriers such as login gates and protective features, but any public site remains vulnerable to determined scrapers. The platform can make abuse harder and more costly, but it cannot guarantee perfect protection.
Why are artists so worried about scraping?
Artists worry that scraping turns their publicly posted work into training material without consent or compensation. Many fear that once an image is collected, they lose control over where it appears, how it is used and whether it helps build AI systems they never agreed to support.









