Meta AI Web Crawlers Are Reading Your Site for Free: Now What?

Meta AI web crawlers now read more of the open web than any other AI company, and most site owners have no idea it is happening. While publishers spend their energy negotiating with Google over how their content trains AI, Meta’s bots quietly pull page after page without a licensing conversation, a payment, or a promise of traffic in return. That imbalance matters for anyone who publishes online, from a national newsroom to a local service business trying to earn visibility in AI answers.

This guide explains what Meta’s crawlers are, why they escape the scrutiny Google gets, how much of the web they consume, and the practical choices you have. You will leave with a clear framework for deciding whether to open the door, close it, or simply watch who walks through.

Key takeaways

  • Meta operates two fast growing AI crawlers, Meta-ExternalAgent and Meta-WebIndexer, that together carry the majority of AI agent traffic on the web.
  • Publishers negotiate with Google because Google historically traded rankings for traffic. Meta never made that promise, so there is nothing to bargain away.
  • DataDome recorded 17.7 billion AI agent requests in the second quarter of 2026, up 45 percent from the previous quarter.
  • Only large media brands sign licensing deals. Meta’s reported News Corp agreement reaches up to 50 million dollars a year, but small sites get no such offer.
  • You can allow, block, or monitor these crawlers. Each choice carries a trade off between AI visibility and control over your content.

What are Meta AI web crawlers, and why do they matter now?

Meta runs several automated bots, but two do the heavy lifting for artificial intelligence. Meta-ExternalAgent collects text to train and improve Meta’s AI models. Meta-WebIndexer builds an index of web content so Meta AI can surface and cite pages inside its answer experiences. A third bot you have seen for years, facebookexternalhit, only fetches link previews when someone shares your page on Facebook or Instagram, so it plays no part in AI training.

The distinction matters because these crawlers do different jobs. Blocking the AI agents does not change how your links look when people share them socially. It also does not touch your position in Google or Bing. Understanding which bot does what is the first step before you decide how to treat any of them.

Why is everyone negotiating with Google while Meta reads the web for free?

The answer comes down to history. For two decades Google offered a simple exchange: let us crawl your pages, rank them, and we will send you clicks. Publishers accepted that trade because the traffic paid the bills. When Google started feeding that same content into AI Overviews and AI Mode, publishers felt the deal had changed, so they pushed back and demanded new terms. There was leverage because Google had something they wanted, namely visitors.

Meta never offered that bargain. Its crawlers take content to sharpen models and answers, yet Meta never promised to send traffic back. You cannot renegotiate a deal that never existed. That is why Meta reads the web with far less resistance, even as its footprint grows past everyone else. If you want the deeper backstory on how AI answers are reshaping the traffic economy, our breakdown of what Google’s earnings reveal about the future of AI search connects the dots.

How much of the web are Meta’s crawlers actually reading?

The scale surprises most people. Bot security firm DataDome recorded 17.7 billion AI agent requests in the second quarter of 2026, a 45 percent jump from the first quarter. Inside that surge, Meta’s bots stood out. Meta-ExternalAgent traffic grew 74 percent quarter over quarter, and Meta-WebIndexer grew a striking 163 percent. Together the two now carry the majority of AI agent traffic across the sites DataDome protects.

Here is the irony. OpenAI’s GPTBot remains the most blocked AI crawler on the web, even though it is no longer the biggest reader. Site owners aim their defenses at the name they recognize while the larger visitor slips past unnoticed.

Crawler Primary job Recent growth Safe to block?
Meta-ExternalAgent Trains Meta’s AI models Up 74% quarter over quarter Yes, no search impact
Meta-WebIndexer Indexes pages for Meta AI answers Up 163% quarter over quarter Yes, but you lose Meta AI citations
facebookexternalhit Builds social link previews Stable No, it breaks Facebook and Instagram shares
GPTBot Trains OpenAI models Most blocked, not the largest Yes, no search impact

What licensing deals exist, and who actually gets paid?

A licensing market has formed, but it has a velvet rope. Meta signed a reported agreement with News Corp worth up to 50 million dollars a year, and it has struck arrangements with a handful of other major outlets including CNN, Fox News, USA Today, and People Inc. Google has done the same elsewhere, with a reported 60 million dollar a year deal with Reddit for access to user posts.

Notice the pattern. Every name on that list is a large brand with lawyers, leverage, and an audience big enough to matter. The average business, blog, or local service site never gets that call. You supply the training data, and you receive nothing in exchange, which is exactly why understanding your options matters more for smaller publishers than for the giants. Our look at why LLM referral traffic converts differently shows where the real upside sits for sites without a licensing check.

How should your business respond to Meta AI web crawlers?

You have three honest choices, and the right one depends on your goals.

  • Allow them. If you want your pages cited inside Meta AI answers, let Meta-WebIndexer through. Visibility in AI answers is becoming its own form of reach, and pulling out of the index removes you from that surface entirely.
  • Block them. If you object to training AI for free, add the AI agents to your robots.txt. Remember that robots.txt works on an honor system, and reports suggest Meta’s crawlers do not always respect it, so a firewall or server level rule enforces the block far more reliably.
  • Monitor first. Before you decide anything, read your server logs. Knowing how often each bot visits, and which pages it favors, turns a guess into an informed call.

Whatever you choose, treat AI visibility as a strategy rather than an accident. Building content that earns citations is a discipline of its own, and our guide to generative engine optimization for revenue lays out how to do it without chasing vanity metrics.

Should you block Meta’s crawlers or let them index your site?

Blocking feels satisfying, yet it carries a cost. Shut out Meta-WebIndexer and your pages vanish from Meta AI answers, which trims a growing referral channel. Shut out Meta-ExternalAgent and you stop feeding the training model, but you gain no visibility either way, so that block is close to pure principle.

For most service businesses, a middle path works best. Allow indexing so you stay eligible for citations, block pure training crawlers if the principle bothers you, and keep watching your logs. If you are weighing whether a control file even helps, our data driven take on whether you need an llms.txt file in 2026 is worth a read before you write a single rule.

Frequently asked questions about Meta AI web crawlers

Which AI crawler reads the most web content?

Meta’s two AI agents, Meta-ExternalAgent and Meta-WebIndexer, now carry the majority of AI agent traffic, according to DataDome. That puts Meta ahead of OpenAI’s GPTBot, which is read less but blocked more.

Does blocking Meta-WebIndexer hurt my Google rankings?

No. Meta-WebIndexer is separate from Googlebot and Bingbot. Blocking it removes you from Meta AI answers but has no effect on how you rank in traditional search engines.

How do I block Meta AI web crawlers?

Add Meta-ExternalAgent and Meta-WebIndexer to your robots.txt with a disallow rule. Because robots.txt is voluntary and not always honored, back it up with a server level or firewall block for stronger enforcement.

Will blocking Meta’s AI bots stop my Facebook link previews?

No. Link previews come from facebookexternalhit, a different bot. Leave that one allowed so your shared links still display images and titles on Facebook and Instagram.

Can small websites get a Meta licensing deal?

In practice, no. Meta’s licensing agreements go to large media brands like News Corp and CNN. Smaller sites supply training data without payment, so their leverage sits in visibility rather than a contract.

The bottom line for site owners

Meta AI web crawlers have become the quiet giants of the web, reading more content than any rival while publishers point their attention at Google. You cannot force Meta to the negotiating table, but you can decide, on purpose, how your content participates. Read your logs, weigh visibility against principle, and set your rules with intent instead of leaving the choice to a default no one at Meta will ever explain to you. The businesses that treat AI crawlers as a strategy, not a nuisance, will be the ones still visible when the next reader arrives.

Scroll to Top