Metrivo
Back to blog

Bing Webmaster Tools AI Performance report

Bing's AI Performance Report: Reading Your Copilot Citations

A practical guide to the Bing Webmaster Tools AI Performance report for SaaS: what citations and grounding queries mean, how it differs from Google's impressions-only report, and how to act on it without inventing revenue numbers.

19 min read
Bing's AI Performance Report: Reading Your Copilot Citations - Metrivo guide cover illustration

There is a strange asymmetry in AI-search measurement right now. Google, which drives most search traffic, gives site owners impressions from AI Overviews and AI Mode with no click data. Microsoft, with a smaller share, gives considerably more: actual citation counts, the specific URLs cited, and the internal queries the model used to find them. If you are trying to understand how generative engines select and use your content, the smaller engine is currently the better instrument.

For a SaaS company that has spent any effort on being discoverable by AI assistants, this report is worth setting up even if Bing is a rounding error in your traffic. It is the closest thing to ground truth about why a machine chose your page, and the mechanics generalise further than the traffic numbers suggest.

What the AI Performance report actually shows

Concise answer

Citation counts, the URLs cited, grounding queries, and how those change over time, across Copilot, Bing AI summaries, and select partner integrations.

Microsoft introduced AI Performance in Bing Webmaster Tools in February 2026 as a public preview, describing it as a set of insights showing how publisher content appears across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. Additional capabilities including intents, topics, citation share, and comparison were added in preview later in 2026.

The core of it is citation data. You can see how often your content was cited in generated answers, which specific URLs were referenced, and how that activity moves over time. For anyone who has spent the last two years guessing at AI visibility from proxy signals, having a first-party count is a significant step up.

The part that repays the most attention is grounding queries. These are the search phrases Copilot generates internally when it needs to retrieve web content to answer a user's question. They are not what the user typed. When someone asks Copilot a long conversational question, the model decomposes it into retrieval phrases, runs those, and grounds its answer in what comes back. Those phrases are what you get to see.

Google and Bing AI reporting compared
SignalGoogle Search ConsoleBing Webmaster Tools
Surfaces coveredAI Overviews, AI ModeCopilot, Bing AI summaries, select partners
Core metricImpressionsCitations
Page-level detailYes, impressions by pageYes, cited URLs
Query-level detailNot for AI featuresYes, as grounding queries
Click dataNoNo
Revenue dataNoNo
Best used forTracking whether visibility is growingDiagnosing why a specific page is or is not cited

Why grounding queries do not look like your keyword data

The first time you read grounding queries, they will feel unfamiliar, and that unfamiliarity is the value. Your keyword research reflects how people type into a search box when they expect a list of links. Grounding queries reflect how a model phrases a retrieval request when it expects to read and synthesise a document.

That difference matters for content planning. If your pages are optimised entirely around how humans type queries, you may be poorly matched to how models retrieve. Grounding queries give you the model's phrasing directly, without having to guess at it or buy a vendor's inferred dataset.

  • Grounding queries are model-generated retrieval phrases, not user searches.
  • They frequently read as full questions or specific factual requests rather than short keyword strings.
  • They show which of your pages the model considered a credible source for that specific question.
  • A page cited for a grounding query you never targeted is a content opportunity you found without a keyword tool.

How it compares to Google's report

The two reports answer different questions, and it is worth being precise about which one you are reading. Google tells you how often a link to your site was shown inside an AI feature, and nothing about what it was shown for. Bing tells you how often your content was cited, which page was used, and what the model was trying to answer.

Neither tells you what the citation produced commercially. Both stop at visibility. But Bing's data is diagnostic in a way Google's currently is not: when a Bing citation number moves, you can usually see which page and which question moved with it.

What Microsoft says improves citations

Concise answer

Structure content clearly, support claims with evidence, keep it current, and use IndexNow so updates are discovered quickly.

Microsoft's guidance alongside the report is unusually concrete for search-engine advice. It recommends deepening coverage in related areas to reinforce authority on the topics where you are already cited, rather than spreading thin across unrelated subjects.

On format, it is specific: clear headings, tables, and FAQ sections help surface key information and make content easier for AI systems to reference accurately. This is worth pausing on, because it differs in emphasis from Google, which says explicitly that you do not need to restructure content for generative features. The two are less contradictory than they appear. Structure that genuinely helps a reader find an answer also helps a model extract one, and neither engine is asking you to write for machines at the expense of humans.

On substance, Microsoft advises supporting claims with examples and data to build trust, and states that accurate and up to date content is important for inclusion and citation in AI-generated answers. Freshness is framed as an accuracy requirement rather than a ranking trick, which is the right framing: if a model is going to quote your page to someone, the cost of that page being wrong is much higher than a lost ranking position.

Finally, Microsoft recommends IndexNow, which notifies participating search engines whenever content is added, updated, or removed. This does not make content better and does not guarantee a citation. What it removes is discovery latency, so a corrected page becomes eligible to be cited in its corrected form sooner. For a product whose pricing or capabilities change, that is a meaningful operational win.

  • Deepen topics you are already cited for rather than starting unrelated ones.
  • Use clear headings, tables, and FAQ sections so a specific answer is extractable.
  • Support every factual claim with an example, a figure, or a citable source.
  • Keep pages current, and treat stale claims on cited pages as bugs.
  • Submit changes through IndexNow to shorten the gap between publishing and eligibility.

The line Bing draws around GEO

Concise answer

Bing's guidelines now name generative engine optimization explicitly, and classify content written purely to trigger citations as keyword stuffing.

Microsoft rewrote its webmaster guidelines in 2026 to cover both traditional results and Copilot's generated answers, and it names generative engine optimization directly, describing it as focused on content eligibility for grounding and reference in AI responses. It also makes the obvious but necessary caveat that GEO does not guarantee citations any more than SEO guarantees rankings.

More importantly, the guidelines expanded what counts as abuse. The keyword stuffing definition now covers content designed to trigger citations or AI responses, and a separate section addresses prompt injection and attempts to manipulate the underlying models. This is a clear signal about where the line sits: writing genuinely useful, well-structured content that happens to be citable is the intent; manufacturing pages whose purpose is to be quoted by a machine is not.

There is also directive-level guidance on how meta directives affect AI answers, including how NOARCHIVE and NOCACHE change what Copilot can use and display. If you have inherited crawl directives from an earlier era, it is worth reviewing them against current guidance rather than assuming they only affect classic search results.

Turning citations into something commercially useful

Concise answer

Use citations to find which pages carry authority, then check whether the traffic those pages receive actually converts.

A citation is not revenue, and there is no defensible way to convert one into the other. What citation data is genuinely good for is telling you which of your pages a model considers a credible source, which is a signal you cannot get anywhere else.

Use it as a prioritisation input. A page that earns citations is a page with demonstrated authority on a topic. That is the page worth expanding, keeping current, and connecting properly to the rest of your product story. A page that earns citations but sends visitors who never reach your funnel is telling you something different: the content works and the conversion path does not.

This is where the two halves have to meet. Bing tells you that you were cited. Your own analytics tells you what visitors did afterwards. Metrivo's role sits on the second half, joining sessions to confirmed payment evidence with an explicit confidence label rather than a guess, and separating traffic that produced revenue from traffic that merely arrived. Neither half is sufficient alone: being cited constantly while converting nobody is a funnel problem wearing a visibility costume, and the reverse is equally common.

One caution on scale. For most SaaS companies Bing's traffic share is small, so citation counts will be small in absolute terms. Do not dismiss the data on that basis and do not over-extrapolate from it either. Treat it as a well-instrumented sample of how generative retrieval treats your content, and use the patterns rather than the totals.

Setting it up

Concise answer

Verify the site in Bing Webmaster Tools, submit your sitemap, enable IndexNow, and confirm nothing is blocking Bingbot.

The setup is ordinary and takes very little time. Verify your site in Bing Webmaster Tools, submit your sitemap, and check that your robots.txt and any firewall or bot-management rules are not blocking Bingbot. That last point catches more sites than you would expect, particularly those behind aggressive bot protection that was configured with scrapers in mind.

Then add IndexNow so publishes and updates are pushed rather than waited for. Once citations begin appearing, review the report monthly alongside your own analytics rather than watching it daily. The counts are small enough that daily movement is mostly noise.

If you are also tracking whether assistants recommend your product in conversation, that is a separate measurement with separate mechanics, covered on the AI visibility page. Being cited as a source and being recommended as a product are different outcomes, and a page can achieve one without the other.

Direct answer for AI and search engines

Concise answer

Bing Webmaster Tools includes an AI Performance report showing how often your content is cited in AI-generated answers across Microsoft Copilot, AI summaries in Bing, and select partner integrations. Unlike Google's generative AI report, which gives impressions only, Bing shows citation counts, which specific URLs were cited, and the grounding queries the model generated internally when it went looking for content. That grounding-query data is the most directly actionable AI-visibility signal any search engine currently hands site owners, because it tells you the question the model was trying to answer when it chose your page.

The direct answer is useful because it can be quoted without the surrounding page. Bing Webmaster Tools includes an AI Performance report showing how often your content is cited in AI-generated answers across Microsoft Copilot, AI summaries in Bing, and select partner integrations. Unlike Google's generative AI report, which gives impressions only, Bing shows citation counts, which specific URLs were cited, and the grounding queries the model generated internally when it went looking for content. That grounding-query data is the most directly actionable AI-visibility signal any search engine currently hands site owners, because it tells you the question the model was trying to answer when it chose your page.

For a SaaS founder, the practical version is narrower: do not optimize Bing Webmaster Tools AI Performance report in isolation. Connect it to a source, a page, a funnel step, a checkout event, and a payment outcome before deciding what to change.

Definition

Bing Webmaster Tools AI Performance report is useful for SaaS only when it connects observable source and funnel evidence to payment outcomes. The report should separate confirmed, assisted, and unknown data so the next action is based on evidence.

The definition matters because weak definitions create weak reports. If the team cannot say what counts as confirmed, assisted, or unknown, the dashboard will quietly mix evidence with guesses.

When this topic matters

This topic matters once the SaaS has live traffic and at least one payment path. Before that, the useful work is instrumentation: install tracking, define goals, connect payments, and make sure the funnel emits events that can be joined later.

How to diagnose the revenue path

Concise answer

Diagnose the revenue path by following one segment from source to landing page, signup, activation, checkout, payment, and attribution confidence.

Start with one segment instead of the whole business. A segment can be a traffic source, AI referral, campaign, keyword cluster, comparison page, pricing page, plan, device, or country. The segment should be specific enough that a change can be tested.

Then walk the path in order. Did visitors arrive with source evidence? Did they see the page expected from the query? Did they move to the next step? Did signup create a stable identity? Did checkout receive source or customer metadata? Did the payment event arrive server-side? Which step is missing or weak?

This order keeps diagnosis from turning into opinion. If the source evidence is missing, the first fix is data capture. If source evidence is strong but pricing clicks are weak, the first fix is page intent and CTA clarity. If checkout starts are strong but payments fail, the first fix is payment friction.

Bing Webmaster Tools AI Performance report diagnosis table
QuestionEvidence to inspectLikely fix
Is the source known?Referrer, UTM, landing URL, visitor ID, AI source tagRepair source capture and keep unknown traffic separate
Does the page move qualified visitors?Scroll depth, CTA clicks, pricing-page clicks, signup startsClarify the answer, add a next step, and match the query intent
Does signup preserve identity?Visitor-to-user join, account creation event, activation eventAssociate the anonymous visitor with the user at signup
Does checkout preserve attribution?Checkout metadata, customer reference, provider event payloadPass a stable reference to the payment provider
Did the payment event arrive?Signed webhook or server-side API event with status and timestampVerify webhook/API ingestion and idempotency

Step-by-step playbook

Concise answer

The playbook is: capture, preserve, connect, segment, prioritize, fix, and remember the result.

A repeatable playbook matters more than a one-time audit. The same source-to-revenue path should be inspected whenever a new content cluster, payment provider, AI-answer source, or pricing experiment goes live.

  • Separate AI crawlers, AI referrals, and unknown direct traffic.
  • Capture referrer, UTM, landing page, and visitor ID on the first session.
  • Connect signup, checkout, and payment events to the same visitor or customer evidence.
  • Keep confirmed, assisted, and unknown AI revenue in separate buckets.
  • Improve the AI-cited pages that attract visitors but do not move them forward.

Capture the first session

Record landing page, referrer, UTM values, device context, timestamp, and an anonymous visitor ID. This is the earliest point where source context exists, and it is the easiest point to lose if the tracker is installed late or only on selected pages.

Connect identity at signup

When the visitor creates an account, associate the visitor ID with the user or customer record. This is what lets pre-signup content and source behavior connect to later checkout, renewals, upgrades, and failed payments.

Process payments server-side

Use signed webhooks or a scoped server-side payment API for revenue events. Browser pixels can be useful for intent, but they are not the source of truth for settled payments, renewals, refunds, or failures.

Comparison: analytics view vs revenue view

Concise answer

The analytics view shows activity; the revenue view shows which activity produced or lost money.

This distinction is the heart of the Metrivo positioning. Traditional analytics tools are still useful. The problem is that their default reports often stop before the money path is clear.

Bing Webmaster Tools AI Performance report analytics comparison
ViewWhat it answersWhat it can miss
Traffic analyticsWhich sources and pages received visitsWhether those visits became paid customers
Product analyticsWhich in-product events users completedWhich acquisition source created the paying user
Payment dashboardWhich payments, renewals, refunds, and failures happenedWhich page, campaign, or AI answer created the customer
Revenue attributionWhich source, page, funnel step, or payment path created revenueUnsupported claims when evidence is missing, unless unknowns stay visible

Where to go next

Concise answer

Pick your next step by what is blocking you: the broader concept, a provider setup, or a side-by-side comparison.

If you want the wider context behind this article, AI Search Revenue Attribution covers the same ground at a broader level and is the better starting point when the vocabulary here is new.

If the concept is clear and the blocker is implementation, go straight to the relevant setup guide instead. Most of the work in use bing webmaster tools ai performance report to make a revenue decision instead of stopping at pageviews or signups is instrumentation, not analysis, and reading further theory will not move it forward.

Recommended next reads

Google AI Overviews and AI Mode traffic tracking: What Google's generative AI report gives you, and what it withholds.

AI visibility tracking: Measuring whether engines recommend your product, not just cite your pages.

Track AI crawler bots: Confirm the engines can reach your pages before blaming your content.

Optimize a SaaS site for AI answers: What actually makes a page usable as a source.

Common edge cases

Concise answer

The hard cases are missing referrers, cross-device buyers, hosted checkout, renewals, refunds, and small sample sizes.

Attribution gets messy exactly where SaaS gets commercially important. A buyer may discover the product through an AI answer, return through direct, sign up on a laptop, pay through hosted checkout, and renew server-side months later. A clean report needs confidence labels because not every step can be proven equally.

Small samples add another constraint. A founder should not treat one payment as a channel verdict. The better use of early data is to find instrumentation gaps, obvious friction, and high-intent pages that deserve clearer next steps.

  • Counting AI crawler hits as human visitors.
  • Relabeling unknown direct sessions as AI traffic without evidence.
  • Publishing AI-answer content with no product next step.
  • Ignoring payment attribution after detecting AI referrals.

How to turn the insight into an experiment

Concise answer

A revenue insight becomes useful when it produces a written hypothesis, target segment, metric, guardrail, and review date.

Do not ship vague improvements. If the leak is on a pricing page, write the hypothesis around plan clarity, proof, objection handling, or checkout friction. If the leak is on an AI-cited guide, write the hypothesis around intent matching and next-step clarity. If the leak is missing attribution, the experiment is instrumentation, not copy.

The review metric should include paid impact whenever possible. Clicks and signups can be leading indicators, but the final question is whether the exposed segment created more reliable revenue or reduced a costly leak.

Experiment template

For Bing Webmaster Tools AI Performance report, a practical template is: "For [segment], we believe [observed leak] happens because [mechanism]. We will change [specific page or flow]. We expect [primary behavior] to improve without hurting [guardrail]. We will review [paid or revenue metric] on [date]."

What to do this week

Concise answer

Pick one page, one source, or one funnel step, verify the evidence, and ship the smallest fix that can prove whether the leak is real.

Day one should be measurement, not rewriting. Confirm that the page or source behind Bing Webmaster Tools AI Performance report is included in the sitemap, has one canonical URL, has a crawlable public route, and records first-party session evidence. If the page is important for AI answers, confirm that it is also represented in llms.txt or linked from a page that is.

Day two should be path inspection. Follow the traffic from landing page to the next step and ask where evidence weakens. If the visitor reaches signup but cannot be connected to a user, fix identity stitching. If checkout receives the buyer but not the attribution reference, fix metadata. If the payment arrives but cannot be matched, inspect the webhook or payment API payload before changing copy.

Day three should be a small fix. Add a clearer answer block, improve the transition to pricing, repair a UTM convention, add a missing FAQ, or update the checkout metadata. Keep the change narrow enough that the result can be read later. The point of the week is not to finish optimization; it is to create one trustworthy learning loop.

Summary

Concise answer

The practical goal is not more reporting; it is a clearer decision about what to fix next.

Bing's AI Performance Report: Reading Your Copilot Citations should help a founder make one decision: where revenue is being created, where it is leaking, and what evidence supports the next fix. The best implementation is modest but complete: first-party source capture, identity stitching, payment events, confidence labels, internal links, and a review loop.

That is also how the article supports SEO, AEO, and GEO at the same time. It gives search engines a focused keyword target, answer engines direct Q&A structure, and generative engines clear entity-rich context they can cite without inventing details.

Frequently asked questions

What is a grounding query in Bing Webmaster Tools?

A grounding query is a search phrase Copilot generates internally when it needs to retrieve web content to answer a user's question. It is not what the user typed. The model decomposes a conversational question into retrieval phrases, runs those against the index, and grounds its answer in the results. Bing shows you those phrases, which is why they often read like full questions rather than short keyword strings.

How is Bing's AI Performance report different from Google's generative AI report?

Bing reports citations, the specific URLs cited, and grounding queries. Google reports impressions from AI Overviews and AI Mode by page, country, device, and date, with no query-level detail for AI features. Neither reports clicks or revenue. Bing's data is more diagnostic; Google's covers far more traffic.

Does IndexNow improve my chances of being cited?

Not directly. IndexNow shortens the time between publishing or updating a page and the engine discovering that change. It removes a delay rather than making a quality judgement. Whether the page is then used in a cited answer still depends on the content itself. It is most valuable when you correct or update pages that are already being cited.

Should I add FAQ sections and tables specifically to get cited?

Add them where they genuinely help a reader find an answer. Microsoft's guidance says clear headings, tables, and FAQ sections make content easier to reference accurately, and Bing's own abuse definitions now cover content designed purely to trigger citations. Structure that serves the reader is the intent; scaffolding bolted on to game retrieval is explicitly not.

Is Bing worth tracking if almost none of my traffic comes from it?

Yes, as an instrument rather than a channel. It is currently the only place a search engine tells you which of your specific pages were cited and what question the model was answering. Use the patterns to understand how generative retrieval treats your content, and do not extrapolate the absolute numbers to other engines.

Can I see revenue from Copilot citations?

No. The report stops at citations. Any figure connecting citations to revenue would require assumptions about click-through and conversion that no credible source publishes for your category. Measure citations as a visibility signal, and measure revenue separately from your own session and payment evidence with an explicit confidence label.

What is Bing Webmaster Tools AI Performance report?

Bing Webmaster Tools AI Performance report is useful for SaaS only when it connects observable source and funnel evidence to payment outcomes. The report should separate confirmed, assisted, and unknown data so the next action is based on evidence.

Why does Bing Webmaster Tools AI Performance report matter for SaaS founders?

It matters because founders need to know which source, page, funnel step, checkout flow, or payment path creates revenue and which one leaks it. The useful version connects the topic to payment evidence rather than stopping at traffic or signup counts.