AI Visibility

Keyword Research Is Now Prompt Research: Mining AI Queries

Keyword-volume tools cannot see conversational AI queries because those sessions produce no click stream to harvest. Here is how to mine GSC and Bing Webmaster Tools for the demand signal that traditional research misses, and how to turn what you find into posts that get cited.

Schematic ink line illustration on cream paper: layered strata, an angled sieve, and a long meandering query line, with a small solid orange pickaxe head above the strata.

Bottom line

Keyword-volume tools are structurally blind to conversational AI queries because those sessions produce no click stream.

  • Bing Webmaster Tools grounding queries (available since February 2026) show which internal search phrases Copilot uses to retrieve your content and how often it cites each page.
  • Google Search Console question-format long-tails, filtered from the standard Performance report, are the second key data source.
  • Mine both sources, cluster by intent, infer the original user prompt, and publish a standalone post per high-value prompt.
  • The Princeton GEO study found structured, citation-friendly content patterns can lift AI visibility by up to ~40%.
  • The Semrush/Indig study found ChatGPT reasoning mode generates roughly five times more retrieval searches per session than fast mode.

The blind spot keyword tools cannot see

Every keyword research workflow you learned in the last decade starts with a volume number. You type a topic into a tool, it returns a monthly search volume, and you decide whether the prize justifies the effort. The logic is sound for traditional search. It is almost useless for AI-driven search.

When someone types "best CRM for a professional services firm with 12 people and a US and Singapore presence" into ChatGPT, Perplexity, or Google AI Mode, that query never appears in any keyword tool. It has no volume. It exists once, in that session, and it is gone. But it represents a buyer at exactly the moment of active vendor evaluation. Missing it is not a minor gap.

This is the structural problem: keyword tools measure what people typed into a search box with a ten-blue-links interface. They do not measure what people prompt into conversational AI interfaces, because those sessions produce no click stream to harvest. The data simply does not flow into the tools you are accustomed to trusting. HubSpot's State of Marketing 2026 found 41% of marketers had already updated their SEO strategy in response to AI search shifts. The other 59% are running a strategy calibrated for a channel that is shifting underneath them.

The fix is not a new tool. It is a different data source. Keyword research has become prompt research.

Building a prompt set from what you already have

A prompt set is a curated list of the conversational queries your buyers use when they are evaluating, comparing, or trying to understand a topic your product or service addresses. Building one from your own data is not guesswork.

The method has three inputs, each from a source you already control.

Input 1: GSC question-format queries. Export your last 90 days of query data from the GSC Performance report. Filter for queries longer than four words (the conventional keyword length ceiling) or for question starters. Group them by intent cluster: evaluative ("how do I choose"), comparative ("X vs Y"), process-level ("how to"), and outcome-level ("best way to achieve"). These clusters are the topic spine of your prompt set.

Input 2: Bing grounding queries. Pull the grounding query report from Bing Webmaster Tools. These queries reveal how Copilot internally translates user prompts into search phrases to retrieve your content. A grounding query like "AI marketing consultant Singapore B2B" tells you Copilot is fielding user prompts about AI marketing consultants in Singapore and finding your page relevant. A grounding query on a topic your page addresses poorly is a gap to close.

Input 3: Prompt-pattern inference. Work backwards from grounding queries and GSC long-tails to infer the original user prompts that generated them. A grounding query of "first-party data strategy B2B SaaS" most likely came from a user prompt like "how should a B2B SaaS company handle first-party data in 2026." That inferred prompt belongs in your prompt set; it is what your content must answer.

The output of this process is a structured list of 30 to 60 prompt-level questions your buyers are asking AI systems. It is not a keyword list with volumes. It has no monthly search estimates. What it has is intent fidelity: every item on the list represents a real moment of active evaluation, not a historical aggregate of clicks.

Worked prompt-mining template

The steps below are the operational sequence leapbuzz uses when building a prompt set from a client's existing data. Each step has a defined input, a defined output, and a copyable query filter or export instruction.

Prompt-mining workflow: six steps

1

Export GSC query data (90 days)

In Search Console: Performance → Search type: Web → Date: Last 90 days → + New → Query → Contains. Export to CSV. Remove branded queries in the next step.

Filter string: queries with >= 5 words OR starting with [how, what, why, which, who, when, should, can, does, is]


2

Pull Bing grounding queries

Bing Webmaster Tools → AI Performance → Grounding Queries tab. Set date range to 90 days. Export top 200 by grounding event count. Note: these are Copilot's internal reformulations, not the original user prompts.

Priority filter: queries with citations > 0 (already being cited) AND queries with grounding events > 50 but citations = 0 (retrieval without citation, a conversion gap)


3

Cluster by intent type

Group all queries from steps 1 and 2 into four intent buckets: Evaluative (choose / select / compare / best), Process (how to / steps / guide), Outcome (results / benchmarks / ROI), and Context (what is / explain / overview). Each bucket maps to a content format.

Evaluative → decision framework post Process → step-by-step guide post Outcome → benchmark post with honest ranges Context → explainer post with BLUF front


4

Infer the original user prompt

For each grounding query, work backwards to the likely user prompt. A grounding query is a reformulation; the original prompt was conversational and longer. Write the inferred prompt as a full natural-language question. This becomes the target your post must answer in its first 150 words.

Grounding query: "AI marketing consultant Singapore B2B" Inferred prompt: "who is a good AI marketing consultant for B2B companies in Singapore"


5

Prioritise by gap score

Score each inferred prompt on two axes: (a) buyer intent strength (evaluative > outcome > process > context), (b) current content coverage (no post = 0, partial mention in pillar = 1, dedicated post = 2). Prioritise high-intent, low-coverage prompts. These are your production queue, in order.

Gap score = intent_weight x (2 - coverage_score) Evaluative weight: 3 | Outcome: 2 | Process: 2 | Context: 1


6

Write the post, front-load the answer

Open every post with a BLUF block: two sentences that answer the inferred prompt directly. The rest of the post provides evidence, context, and depth. This structure serves both the human reader and the AI retrieval system: the answer is quotable from the first paragraph.

Internal links: connect every new prompt-set post to the topic cluster it belongs to. A post answering a process query should link to the evaluative post that explains why the process matters, and to the outcome post that shows what the process produces.

Turning long-tail prompts into posts

The prompt set tells you what to write. The conversion step is mapping each prompt to a content format that answers it well enough to be retrieved and cited.

Not all prompts produce the same content format. Evaluative prompts ("how do I choose a digital analytics partner") call for decision frameworks: criteria, trade-offs, common mistakes, a structured comparison. Process prompts ("how to set up a first-party data pipeline") call for step-by-step structure with named outputs at each step. Outcome prompts ("what results should I expect from influencer marketing") call for honest benchmarks, context for variation, and the variables that matter most.

The structural principle that applies across all of them: answer the prompt in the first 150 words. AI retrieval systems quote from the opening of a section more often than from its middle. The Princeton GEO study (arXiv:2311.09735) found that citation-adding content patterns, including direct quotable statements and structured fact blocks near the top of a section, lifted AI visibility by up to ~40% on their benchmark. Front-loading the answer is not a stylistic choice. It is a retrieval mechanic.

Each prompt that clears the intent threshold becomes a standalone blog post. Not a section in a longer page. Not a paragraph buried in a pillar. A post with its own URL, its own structured data, and its own BLUF block that answers the prompt in two sentences. This is how leapbuzz approaches content architecture for AI visibility: one prompt, one URL, one authoritative answer. The post you are reading was seeded by exactly this process.

For clients across SG, AU, US, CA, and MY, the prompt-set approach surfaces market-specific demand patterns that aggregate keyword tools miss entirely. A query like "what are the GST implications of influencer marketing payments in Singapore" has essentially no measurable keyword volume, but it represents a live compliance question from a marketing procurement team. A post that answers it precisely gets cited. A pillar page that mentions it in passing does not.

The measurement loop: from grounding query to pipeline

Publishing prompt-set posts without a measurement loop is planting seeds and never checking the soil. The loop is straightforward once the two data sources are configured.

Set a monthly review cadence. In Bing Webmaster Tools, track grounding-query-to-citation conversion rate by page: which pages are retrieved but not cited? That gap is where you improve answer structure, add a BLUF, or sharpen the first paragraph. In the GSC AI performance report, watch which pages gain impressions in AI features month over month. Impression growth without traffic growth is normal right now. AI answers resolve queries without sending clicks. That is brand presence in the answer layer, not a failure state.

For the post-to-pipeline connection, use UTM parameters on CTAs within prompt-set posts. When an AI answer cites your post and a reader clicks through, that session is trackable. The post is the citation vehicle; the CTA within it is the lead channel. Each prompt-set post is a tiny landing page with a specific action and a specific audience.

The Semrush/Indig study published July 1, 2026, found that ChatGPT's reasoning mode generates roughly 1,130 searches per 100 user prompts, against roughly 245 for fast mode. That ratio means reasoning-mode queries drive substantially more retrieval activity per user session. For B2B buyers doing deep research in reasoning mode, the content that gets retrieved and cited is the content that addresses the full prompt with precision, not the content optimised for a two-word keyword. Prompt research is the only method calibrated for that retrieval pattern. A keyword tool cannot help you there.

For deeper coverage of how AI answer engines select and cite sources, the GEO playbook covers the full retrieval-and-citation architecture. The AEO step-by-step guide walks through the entity and schema layer that makes content machine-readable. And ChatGPT reasoning mode citation patterns unpacks the specific mechanics of how reasoning models decide what to quote. If you want to understand the gap between how Bing Webmaster Tools surfaces AI citation data today and what a fuller measurement stack looks like, leapbuzz's visibility optimization practice is where to start.

What GSC and Bing Webmaster Tools now actually show

Two platforms have introduced AI-adjacent query data in 2026. Understanding what each one does and does not show is the starting point for mining it usefully.

Google Search Console. On June 3, 2026, Google began rolling out its Search Generative AI performance report, initially to a subset of UK properties (a Competition and Markets Authority requirement before wider expansion). The report surfaces impressions, pages, countries, and device breakdown for appearances in AI Overviews and AI Mode. What it does not show: query data. There are no keywords, no click data, no CTR. Google has confirmed those metrics will come in future iterations. Right now, the value is directional: you can see which pages are getting surfaced in AI features, which tells you where your content is earning retrieval even before query-level data exists.

The query data you actually have access to today still comes from the standard Performance report. Filter the Queries tab for question-format strings (how, what, why, which, who, when) and sort by impressions. These are not AI prompts, but they are a first-order signal: the conversational surface of your organic search audience. Long-tail question queries that drive low clicks but respectable impressions are frequently the queries that AI systems are answering from your content without sending traffic.

Bing Webmaster Tools. Microsoft moved faster. The AI Performance report launched on February 11, 2026, inside Bing Webmaster Tools, and it shows two metrics that have no equivalent in GSC today. The first is grounding queries: the internal search phrases Copilot generates behind the scenes when it needs to retrieve web content to answer a user's prompt. These are reformulated queries, not the user's original words. The second is citations: how many times Copilot used a specific page from your site in its response.

The gap between grounding queries and citations is instructive. OtterlyAI's own domain generated 647 unique grounding queries and over 30,000 grounding events across 173 pages in three months, yet the vast majority of that activity was invisible to end users. Your content is being consumed by AI systems at a scale that standard analytics cannot see.

For leapbuzz clients operating in Singapore, Malaysia, Australia, the US, and Canada, the Bing AI Performance data is immediately actionable: Copilot has significant market share in the US and Australia, and it is the embedded AI in Microsoft 365 that enterprise buyers use. Grounding query data from Bing is therefore a direct read on B2B buyer prompt patterns in those markets.

Frequently asked questions

What is prompt research and how is it different from keyword research?

Keyword research measures historical search volume for short queries typed into a ten-blue-links search engine. Prompt research identifies the conversational, full-sentence questions buyers type into AI interfaces like ChatGPT, Perplexity, Google AI Mode, and Microsoft Copilot. Those AI sessions produce no click stream, so they never appear in keyword tools. Prompt research mines different data sources, principally GSC long-tail question queries and Bing Webmaster Tools grounding queries, to reconstruct what buyers are actually asking AI systems about your topic.

Where can I find AI query data for my website?

Two platforms provide partial visibility. Bing Webmaster Tools launched its AI Performance report on February 11, 2026, showing grounding queries (the internal search phrases Copilot generates to retrieve your content) and citation counts per page. Google Search Console introduced its Search Generative AI performance report in June 2026, currently showing impressions, pages, countries, and devices for appearances in AI Overviews and AI Mode, but no query-level data yet. For query data today, the GSC standard Performance report filtered for long-tail question-format queries (five or more words, starting with how/what/why/which/who) is the most accessible first-party signal.

What are Bing grounding queries and why do they matter?

Grounding queries are the internal search phrases Microsoft Copilot generates when it needs to retrieve web content to answer a user prompt. They are not the user's original words; they are Copilot's reformulation of the user's intent into a retrieval query. Grounding queries matter because they reveal which topics Copilot associates with your content and how it translates buyer intent into retrieval actions. A page with many grounding events but zero citations is being read but not quoted, which tells you the answer structure needs improving. A grounding query on a topic your site does not cover well is a gap worth filling with a dedicated post.

Does the Google Search Console AI performance report show keyword data?

Not yet. As of its June 2026 launch, the GSC Search Generative AI performance report shows impressions, pages, countries, and devices for appearances in AI Overviews and AI Mode, but no query data and no click data. Google has confirmed additional metrics will come. The practical implication: today the report tells you which pages are being surfaced in AI features, not what prompts triggered those appearances. For query-level insight, use the standard Performance report filtered for question-format long-tail queries, and use Bing Webmaster Tools for grounding query data.

How many prompts should be in a prompt set?

A working prompt set for a single topic cluster typically contains 30 to 60 inferred user prompts. Fewer than 30 and you are likely missing meaningful buyer intent variants. More than 60 per cluster and the prompts start overlapping enough that consolidation is more efficient than separate posts. Build the prompt set from your own GSC long-tail data and Bing grounding queries; do not fabricate prompts by guessing. The gap-score prioritisation step (intent weight times coverage gap) then reduces the full list to a production queue of the 10 to 15 highest-value posts to write first.

What content format works best for AI retrieval?

The format that earns the most AI retrieval is whichever one answers the inferred prompt directly in the first 150 words. The Princeton GEO study (arXiv:2311.09735) found that citation-adding patterns, including direct quotable statements and structured fact blocks near the top of a section, lifted AI visibility by up to ~40% on their benchmark. Beyond that opening structure, the format maps to intent: evaluative prompts get decision frameworks, process prompts get step-by-step guides, outcome prompts get honest benchmarks with context for variation. Standalone posts with their own URL and structured data outperform topic mentions buried inside longer pillar pages.

How do I connect AI citations to business pipeline?

Use UTM parameters on CTAs within each prompt-set post. When an AI answer cites your post and a user clicks through, that session carries the UTM tag and is attributable in your analytics. The post functions as the citation vehicle; the CTA within it is the lead channel. Track this by post and by intent cluster so you can identify which inferred prompt categories produce the most qualified inbound sessions. This is not a perfect attribution model, but it is tractable and improves as your prompt-set post library grows and Bing citation data accumulates.

Does ChatGPT reasoning mode change what content gets cited?

Yes, meaningfully. The Semrush/Indig study published July 1, 2026, found that ChatGPT reasoning mode generates roughly 1,130 searches per 100 user prompts, against roughly 245 for fast mode. That means reasoning-mode sessions drive substantially more retrieval activity per user session, pulling from more sources before generating a response. The implication: content optimised for a short keyword will struggle against reasoning-mode retrieval, which is looking for precise, expert-level answers to the full, contextual user prompt. Prompt-set posts structured around the inferred full prompt are calibrated for reasoning-mode retrieval in a way keyword-optimised pages are not.

Should every long-tail prompt become its own post?

Not automatically. Apply the gap-score filter first: high buyer intent plus low current coverage equals publish. Low intent or already well-covered by an existing post equals internal link from that post instead. A prompt that represents a genuine buyer decision moment, one where a buyer is actively evaluating or comparing, almost always earns its own post. A prompt that is definitional or contextual and already answered well elsewhere earns an anchor link within the relevant existing post. The goal is one authoritative answer per distinct buyer intent, not one post per query string.

Related

Work with leapbuzz

Ready to replace keyword guesswork with a prompt set built from your own GSC and Bing data?

leapbuzz builds AI-visibility content strategies for businesses across Singapore, Malaysia, Australia, the US, and Canada. The prompt-research workflow, the post architecture, and the citation measurement loop are all included.

Talk to us