Somewhere on your website, right now, there is probably a tracker nobody remembers adding, whether it’s an embed in an old article, a pixel on a campaign page from three years ago, or a tag a vendor loaded, which loaded another one. Your privacy team has never seen it, your consent banner doesn't cover it, and it fires anyway, with obvious potential consequences from a compliance and performance standpoint.
This article is about how to find that tracker and why we believe our approach to website scanning is best suited to help you do exactly that.
First things first: Why compliance monitoring matters
Consent management doesn't end when your banner goes live. What you disclosed yesterday drifts away from what your site actually does today, as your business evolves:
- The marketing team adds a new tool
- Vendor update scripts
- Ad tech partners sync with new partners behind the scenes
- You stop working with a vendor, but no one updates the website
- Landing pages tied to a new campaign require specific tracking
- IT makes changes to your tag manager that update how and when tracking tags fire
- Developers copy old scripts to a new page or site section
A website is a living thing, and its tracking footprint changes without anyone deciding it should.
The regulatory reality is unforgiving about this, and enforcement doesn't happen at the level of your average page but at the level of a specific URL, since complaints generally cite a single page where a tracker fired before consent.
When organizations run their first compliance scan with us, the average score is 39.6% before any remediation, a number explained by the fact that across more than 2,000 organizations, an average of 67% of the vendors active on a site were never declared in the consent banner.
Regulators are paying close attention to this exact gap. In 2025, the California Privacy Protection Agency fined retailer Todd Snyder $345,000 after finding its consent banner was misconfigured and unmonitored, so opt-out requests never actually took effect. In Europe, France’s CNIL fined Google €325M and Shein €150M for cookie violations. And the list goes on.
What we can learn from this is that you can't rely on the tool if you're not checking that it works. Which raises the question every organization must ask itself:
How do you confidently monitor a website with thousands (sometimes millions) of pages when the tracker that makes you non-compliant might exist on exactly one of them?
Let’s examine how most organizations attempt to do so, and where our approach differs.
How most organizations under-scan their websites

Most scanning setups fall short in one of two directions. They look opposite, but they fail the same way, by leaving parts of your real tracking footprint invisible.
1. Scanning wide but shallow
The first failure mode is treating scanning as a page-counting exercise, loading as many URLs as possible, recording what appears, and calling it a day.
Why is this approach flawed?
A page that just loads is not a page that behaves, and many trackers fire only when a visitor does something, whether it’s scrolling past the fold, opening a chat widget, adding to cart, or logging in.
A scanner that never scrolls, never clicks, and never authenticates sees a version of your website that no real visitor experiences. In effect, volume without behavior gives you a big number and false confidence, and page count alone is a vanity metric.
2. Scanning smart but narrow
The second failure mode is to sample intelligently, an approach that is seemingly more sophisticated. Because most large sites are template-driven, it is tempting to group similar pages, scan a few representatives from each group, and declare the group covered. Why scan your 40,000th product page when it behaves like the first 100?
Why is this approach flawed?
The core logic holds right up until it meets a real website. Let’s take a large publisher as an example, with ten years' worth of articles (and pages to scan). The templates are identical, but the content isn't:

Two articles on the same template can have completely different third-party footprints, because the trackers live in the content, not the template.
Sampling assumes pages that look alike behave alike, but the pages that make you non-compliant are precisely the ones that break that assumption, and you can't know a group of pages is homogeneous without scanning enough of them to verify it. Otherwise, you're only assuming coverage instead of measuring it.
Our approach at Didomi: Depth and breadth, not a choice.
Our core belief is that depth and breadth aren't competing scanning strategies, but two halves of one answer, and a vendor that makes you choose between them is showing its limitations rather than best addressing your needs.
Depth without breadth means you understand a corner of your site perfectly while the archive stays dark. Conversely, breadth without depth means you take a static photograph of everything on your site, without accounting for interactions.
Real coverage behaves like a real visitor across as much of the site as possible, repeatedly and with memory, so coverage compounds over time rather than resetting with every scan.

Of course, that's expensive to build, but we think comprehensive compliance monitoring requires it. The risk lies in the exceptions, and you find exceptions by both looking broadly and acting deeply.
How we've built our compliance monitoring solutions at Didomi

Our compliance monitoring and scanning solutions are articulated around four pieces designed to work together:
- XL scan: The breadth part of our product offers up to 20,000 pages in a single scan. That’s enough to sweep the archive, the campaign pages, the corners nobody remembers… not just the homepage and its template neighbors.
- Behavioral simulation: The depth aspect is handled by our bots, which can scroll, click, fill forms, and follow custom scenarios you define, including logged-in journeys where an authenticated user sees a genuinely different site. This allows you to surface trackers that only fire when a visitor behaves like a visitor.
- AI layer: Compliance depends on what happens after the consent choice, so our bots have to make that choice like a human would. Our AI can read your consent banner (whatever the language and design), identify the correct accept and refuse options, and interact with them reliably. A tracker that fires after refusal, or before any choice, is exactly the kind of violation that only shows up when the bot gets the banner interaction right. Learn more about how we integrate AI in our scanning and monitoring solutions.
- Rotating coverage: Classic scans rotate through randomly selected pages on every run. Each cycle looks at a different slice, so coverage accumulates across scans instead of photographing the same corner of the site forever.
As we mentioned, scanning at this scale costs real compute, and we deliberately chose to invest in the fleet rather than in reasons not to need one. Spread across a customer base, the infrastructure that catches a forgotten tracker costs a fraction of a single enforcement action. We think that's the cheapest insurance in privacy tech, and we priced our approach on that belief.
What to ask any monitoring vendor you're considering
Whenever you’re shopping for a scanning and monitoring solution, consider bringing up these questions (even to us):
- What is your actual maximum page count per scan? Please provide a number.
- Do you simulate real behavior, like scrolling, clicking, logged-in sessions, and custom journeys?
- How would you find a tracker that exists on exactly one forgotten page of my site?
- How does coverage evolve between scans? Does each scan start from zero, or does knowledge accumulate?
If the answers are clear and concrete, you're in good hands. If you’re provided with a long explanation of why the question is wrong, that should also tell you something.
To ask our team these questions and discuss your specific challenges, book a call with one of our experts, and to continue reading about our solutions, head to our guide to compliance scans:
{{compliance-scan-types}}














