Skip to main content
ExplainerSocial Media ForensicsExplainer· 5 min read· in Content Types

How AI Agents Are Replacing CrowdTangle to Track Cross-Platform Misinformation

Two years after Meta shut down the industry's primary social monitoring tool, a new generation of AI-powered infrastructure is helping researchers track coordinated narratives across the internet.

By Wei Zhang

Independent Investigators 60%Platform Operators 40%
Independent Investigators
Argue that open, cross-platform data access is essential for democratic transparency and holding networks accountable.
Platform Operators
Prioritize user privacy and regulatory compliance, arguing that data access must be strictly vetted and controlled.

Perspectives this story doesn't cover

  • Privacy Advocates
  • Commercial Data Brokers

Summary

  1. Meta shut down CrowdTangle in 2024, removing the primary tool researchers used to track viral misinformation.
  2. Meta's replacement, the Meta Content Library, is restricted to academics and operates in a secure cleanroom.
  3. New AI tools like Arbiter use language models to scrape and analyze narratives across nine different platforms.
  4. These tools cluster similar posts to extract core claims, then cross-reference them against Wikipedia to flag manipulation.
  5. Independent monitoring tools face ongoing challenges from rate limits and restricted platform APIs.

Independent investigators and newsrooms can once again track coordinated manipulation campaigns across the social web, two years after the industry's primary monitoring tool went dark. A new generation of artificial intelligence infrastructure has begun mapping how narratives jump between platforms, restoring visibility into the forces shaping public feeds before they influence real-world decisions. The shift marks a fundamental change in how digital forensics are conducted, moving away from single-platform dashboards toward autonomous agents that scrape and synthesize data across the entire internet.

The transition follows the August 2024 shutdown of CrowdTangle, a Meta-owned analytics dashboard that served as the central nervous system for tracking viral content. For a decade, researchers relied on it to watch claims spread across Facebook and Instagram in real time. When Meta retired the tool, it severed the primary line of sight into the world's largest information ecosystem. The company cited the need to comply with tightening global privacy regulations and to meet new standards for data-sharing, but the closure left watchdog groups blind during a year of major global elections.[1]

Meta did not leave the space entirely empty. The company shipped the Meta Content Library, a highly controlled replacement dataset containing public posts from Facebook, Instagram, and Threads. But the new system fundamentally changed who gets to look. Access is strictly gatekept by the Inter-university Consortium for Political and Social Research at the University of Michigan. The review process takes between two and six weeks, and approval is explicitly restricted to academic and non-profit researchers. Commercial journalists and independent investigators are barred from entry.

For those who do clear the institutional hurdle, the Meta Content Library operates inside a digital cleanroom. Researchers must query the data through a Virtual Data Enclave or Meta's Secure Research Environment. They cannot download the raw data locally, and they are required to delete any cached records every 180 days. It is a secure, compliant system designed for rigorous academic study, but its architecture is entirely incompatible with the speed of a daily news cycle or the demands of breaking-news verification.

Access to Meta's replacement dataset requires institutional approval and operates inside a secure digital cleanroom.

"If you're looking at a specific post, and if you want to look at a specific narrative, then MCL is okay," said Duuya Baatar, cofounder of the Mongolia-based Nest Center, which retained access through a fact-checking partnership. "If you want to analyze how the information ecosystem is changing, then MCL is nowhere near CrowdTangle." That operational gap triggered the development of independent, cross-platform alternatives designed specifically for the pace of modern journalism, prioritizing trend discovery over deep historical archiving.[1]

The most prominent of these new systems is Arbiter, built by the nonprofit research lab SimPPL. Rather than relying on a single company's internal dashboard, Arbiter deploys large language models and autonomous agents to pull public posts from nine different networks simultaneously, including X, TikTok, Reddit, YouTube, and Bluesky. The platform is designed to automate the manual labor of digital forensics, allowing a single reporter to track a coordinated campaign as it migrates from fringe message boards to mainstream algorithmic feeds.[1][2]

The most prominent of these new systems is Arbiter, built by the nonprofit research lab SimPPL.

When a researcher inputs a plain-language query about a developing topic, Arbiter's agents fan out across platform application programming interfaces and third-party data dumps. They retrieve thousands of relevant posts and feed them into a language model, which clusters similar phrasing and extracts the core underlying claims. This clustering mechanism strips away the noise of individual user variations, revealing the structural skeleton of a talking point and identifying which accounts are injecting identical language into different communities to manufacture artificial consensus.[2]

The tool then attempts to verify those extracted claims automatically. Arbiter periodically downloads the entirety of English Wikipedia using the Wikimedia Enterprise Snapshot API. It cross-references the social media claims against this structured encyclopedic data, highlighting deviations from established facts. By automating the initial verification step, the system allows human investigators to focus their energy on tracing the origin of the network manipulation rather than spending hours manually debunking the individual claims themselves.[2][3]

Arbiter uses language models to cluster thousands of posts into distinct claims, then cross-references them against encyclopedic data.

The ultimate goal of this architecture is predictive rather than reactive. "Social listening systems tell you where conversations are happening as they blow up," SimPPL co-founder Swapneel Mehta told Nieman Lab. "The problem is, as journalists, you want to be ahead of the alert." By mapping the velocity of a narrative across multiple platforms, the software generates trend lines that indicate when a specific falsehood is gaining enough momentum to break into the broader public consciousness, allowing newsrooms to prepare counter-reporting before the peak.[1]

The capability has already been tested in the field with measurable results. Earlier this year, Baatar's team at the Nest Center used Arbiter to monitor discourse around a Mongolian Supreme Court ruling that struck down a law banning false information. The tool surfaced a cluster of small Facebook accounts suggesting lawmakers would replace the statute with new libel provisions. Baatar initially dismissed the chatter, but weeks later, the parliament formally announced a working group to do exactly that, validating the system's early-warning signals.[1]

Yet the new monitoring ecosystem remains highly fragile. Arbiter and tools like it are navigating what researchers call the API apocalypse—an era where platforms are aggressively locking down programmatic access to their data to prevent unauthorized artificial intelligence training. While SimPPL currently indexes posts receiving roughly 100 billion views a month, it still faces severe rate limits, relies heavily on data donated by academics, and requires constant legal review to ensure its scraping methods do not violate shifting platform terms of service.[1][2]

By mapping narrative velocity across multiple platforms, predictive tools aim to alert newsrooms before a falsehood peaks.

The current landscape represents a permanent fracture in how the internet is monitored. The era of a single, unified dashboard provided by the platform itself is conclusively over. In its place is a bifurcated system: pristine, slow-moving academic cleanrooms on one side, and scrappy, AI-driven cross-platform scrapers fighting for API access on the other. For the journalists tasked with explaining the internet to the public, visibility now depends entirely on their ability to stitch these fragmented views back together.[3]

Definitions

Social Listening
The process of monitoring digital conversations to understand what users are saying about a specific topic, narrative, or brand.
Virtual Data Enclave
A secure, controlled computing environment where researchers can analyze sensitive datasets without downloading the raw data to their own devices.
Large Language Model (LLM)
An artificial intelligence system trained on vast amounts of text, used in this context to automatically cluster similar social media posts and extract their core claims.
API Rate Limit
A restriction enforced by a platform on how many times a software tool can request data within a specific timeframe, designed to prevent server overload or mass scraping.

Sources

Source coverage

3 outlets

2 viewpoints surfaced

Independent Investigators 60%Platform Operators 40%
  1. [1]Nieman LabIndependent Investigators

    Two years ago, Meta killed CrowdTangle. Can a new AI tool fill the void?

    Read on Nieman Lab
  2. [2]SimPPLIndependent Investigators

    Arbiter investigates who is shaping your social media feed

    Read on SimPPL
  3. [3]Factlen Editorial Team

    Synthesis by Factlen editorial team

    Read on Factlen Editorial Team

Comments

Stay informed

Every angle. Every day.

Get Content Types stories with full source coverage and perspective breakdowns delivered to your inbox.