← Blog

fig. network
Trackingplan

Same models, different answers.

Inside the new agentic Trackingplan: how it works, and why it isn’t comparable to an LLM with an MCP.

Alexandros Andre ChaaraouiCo-founder & CTO
13 min read · 2764 words

TL;DR An agent can only be as right as what it can see. Claude or ChatGPT with an MCP sees the numbers your analytics tool reports. Trackingplan’s agent sees the data behind them, and it can correlate that data and verify what it finds before it answers:

  • The raw data. Every request your site sends to every destination, alongside the dataLayer, the consent state, your GTM releases and your ad spend.
  • Built-in expertise. A semantic layer that brings every tracking and martech tool into one schema, plus years of digital analytics know-how written down as rules, metric definitions and methods.
  • Tools to check its work. SQL, Python and a real browser that can reproduce a problem on your live site.
  • A standard of proof. It names a cause only when two independent pieces of evidence agree, and it flags any number distorted by a tracking issue our monitoring has caught.

We just launched the new Trackingplan, your digital analytics agent. You ask, and your data answers:

Two example questions to Trackingplan, and its answers.
Two example questions to Trackingplan, and its answers.

The answers come from the raw traffic, not from a summary, and they show the evidence behind them.

Everyone is shipping an agent right now, so it’s fair to ask why we didn’t just connect Claude or ChatGPT to analytics data through an MCP server and call it a day.

We asked ourselves the same thing. Here’s the long answer.

The model is not the hard part

Connect a general-purpose agent to your analytics through an MCP and it sees the numbers the tool reports, after the tool has processed, filtered and modeled the data. It never sees the request your site actually sent, and it has no idea what correct tracking looks like. It doesn’t know your purchase event lost its transaction_id on Tuesday, so it will happily compute a conversion rate from it. It can’t see what changed in your tag manager, and it can’t open your website to check.

We all have access to the same frontier models, and they’re good enough at reasoning. What decides whether an analytics agent gets it right is everything around the model:

We built each of those layers for one job: answering digital analytics questions with numbers you can trust, and knowing when the data underneath is wrong.

1. The data: what your site actually sent

It starts with hits: every request your website, app or server sends to its analytics, marketing and advertising destinations. Trackingplan captures them at the source and decodes each one with a parser built for its destination.

That’s the difference from a data agent sitting on your warehouse or your GA4 export. Those see what a tool received and processed. Hits are what your site sent, before any destination touched it. Here’s what a single one carries:

Anatomy of one hit: a GA4 purchase and everything Trackingplan captures around it.
Anatomy of one hit: a GA4 purchase and everything Trackingplan captures around it.

Alongside the hits are Core Web Vitals and JavaScript errors, the pixels that load on each page, and consent acceptance rates for every destination, cookie and domain.

The agent also reads the output of Trackingplan’s monitoring, which watches your events, your dataLayer implementation, your schemas and business rules, and your traffic. It sees every issue that monitoring has caught, with its history, and the settings behind each check.

With this launch we added two new sources:

The hard tracking bugs live at the seams between these sources: a tag removed in a container release, a campaign whose UTMs don’t match what lands on the site, a consent banner that blocks one pixel but not the next. An agent that sees only one side can’t see the seam.

2. A semantic layer designed for the digital analyst

Every destination has its own format. GA4, Meta, TikTok, Adobe and the rest name the same order, page and campaign differently, and consent shows up as a CMP cookie or a Google Consent Mode signal. Left raw, all of that has to be untangled before every single question.

The semantic layer reconciles it into one nomenclature and one schema: around 30 tables covering hits, detected issues, schemas, daily stats, consent, tag manager and ads. It’s designed for the digital analyst, and for the agent writing queries on their behalf.

Every tool’s data flows into one semantic layer, and a single query reaches all of it.
Every tool’s data flows into one semantic layer, and a single query reaches all of it.

In practice:

3. The expertise, and how the agent uses it

This is where years of digital analytics work (implementation, measurement, consent, attribution and data quality) become instructions a model can follow. It’s organized the way an expert works: a few rules always in mind, and the manual opened only when it’s needed.

What the agent always has in mind, and what it opens for one question.
What the agent always has in mind, and what it opens for one question.

An always-on core. The rules every answer follows, plus a glossary and a one-line index of every column, so the agent knows what exists before its first query. Among the rules:

Manuals on demand. One per table, one metrics file per area, and a set of tested example queries. The agent reads a table’s manual before it first queries it, so a chat’s context grows with what the question touches, not with everything that exists. It also keeps the agent sharp: models follow instructions measurably worse when the prompt is padded with things the question doesn’t need.

Skills. The methods an experienced analyst follows, written down so the agent picks the right one for the question. Some are named investigations, like Root Cause Analysis or tracing a detected issue back to the change that caused it. Some produce a health report for an event, a property, a destination or your consent setup. And some are the checks analysts repeat every time the numbers don’t add up: duplicate order IDs, events lost between the dataLayer and your destinations, consent choices that don’t change what fires, UTMs that change mid-session, sessions with no page view, and more. A full sweep runs whichever of those fit your setup.

What your team teaches it. Facts and standing rules, like “the Meta pixel was retired” or “exclude staging hostnames.” They’re saved from the conversation, shared by everyone working on that site or app, reviewable in settings under Plan context, and applied in every chat and every audit.

4. A harness built to check its own claims

The model is the reasoning engine. The harness decides what it can touch and how it proves things, and ours is custom-built for this job:

And it’s not a single model. In the chat you choose between three tiers, all from the GPT-5.6 family:

Behind the scenes, a small model writes chat titles and summaries.

5. Evidence before conclusions

The piece we’re proudest of is a rule about when the agent is allowed to name a cause.

A cause becomes a conclusion only when two independent pieces of evidence agree: for example, a container change detected just before the problem started, plus traffic or a live check confirming how it broke. Anything less is presented as a hypothesis and labeled as one. If it can’t find the cause, it says so; it never claims there isn’t one. And when a check it ran on its own initiative comes back empty, it doesn’t pad the answer with the story of that check.

The rule exists because a language model, handed a symptom, will always find a plausible story for it. Left alone, it may invent a cause the data doesn’t support, or take a minor factor and present it as the main driver. Asking for two independent pieces of evidence successfully mitigates both, and keeps the agent from sending you off to fix the wrong thing.

Every answer shows the queries behind it. Most readers won’t audit SQL, though, so the discipline has to be built into the agent.

Beyond monitoring: what you can do now

Trackingplan has always told you when your tracking breaks. The agent takes on the analyst work around it.

Audits. You describe a recurring check in plain words: what to look at, what to ignore, and how to judge the result.

A library of ready-made audits comes with it. This is the first set we’re launching, and it will keep growing:

The Audit Library’s first set.
The Audit Library’s first set.

Live debugging on any website. Ask it to check a page and it shows you which pixels load and fire, what the dataLayer carries and how consent behaves, with every captured request there to inspect.

Tag manager history. Ask what changed in a GTM container and when, or which release lines up with a broken event or a drop in traffic.

Reports. Turn an investigation into something you can hand over. Reports mix narrative, charts and tables, keep a version on every save, and can be edited by hand or by asking. Keep one private, share it with your team, or publish it by public link. If an admin allows it, readers of the link can ask follow-up questions, answered only from the data behind that report.

Tracking issues debugged end to end. When monitoring catches a broken event, a dataLayer change, a schema or business-rule violation or a traffic drop, click Analyze and the agent works through it:

  1. It reads the issue: what kind it is, the affected event, when it started, its impact and the failing values.
  2. It narrows the problem down in the hits: which pages, devices, consent states and campaigns it affects.
  3. It finds the tag, trigger and variable that send the event, and checks what changed in the container around the time the problem started.
  4. When a live check could settle it, it offers one. If you say yes, it reproduces the flow in a real browser, with the same consent state the affected traffic had.

The answer reads like this:

An issue traced back to the GTM release that caused it, confirmed live, with the fix.
An issue traced back to the GTM release that caused it, confirmed live, with the fix.

How we keep it honest

And if you live in Claude or ChatGPT?

We’re not against MCP: we ship one. Connect Claude, ChatGPT or Cursor to Trackingplan and you get an Ask Trackingplan tool. Behind that tool is the whole agent described here, with its data, its context and its standard of evidence, working from inside the assistant you already use. The same agent also answers in Slack.

Behind the scenes

Behind the scenes, in numbers.
Behind the scenes, in numbers.

The part that doesn’t come with the model

Models will keep getting better, and we’ll keep switching when they do. What doesn’t come with any model is the data your site actually sends, a semantic layer that makes it readable, the expertise of people who’ve spent years in digital analytics, and the discipline to call something a cause only when the evidence agrees.

That’s the part we built. Already a customer? It’s in your Trackingplan app now. If not, try it on sample data.

Alexandros Andre ChaaraouiCo-founder & CTO

Meet Alexandros, our CTO at Trackingplan, driving innovation with a Ph.D. in Machine Learning to develop reliable analytics solutions.

Read next

All Blog →

Trackingplan

See everything. Miss nothing.

Your implementations audited around the clock with real-time, real-user data. Real-time alerts about errors or changes in your data, campaigns, pixels, privacy and consent. Let AI flag issues before they cost you.