How to Find Buyer Language, So Buyers Find You via AI

buyer language for ai visibility

If you’ve ever run a voice-of-customer (VoC) project, you know the feeling when it lands.

For most of my career this was a manual, expensive slog. You’d write a discussion guide, lean on sales to pull a list of people willing to talk, and then spend three weeks chasing calendar holds to get some one-on-one time with customers. 

If you were lucky, you’d end up with fifteen recorded interviews ready for transcription. Then you’d get a spreadsheet with coded responses and themes that emerged. Agencies charged serious money to run this, and the insights were worth every dollar.

These days, virtual-call recordings pile up on their own and the transcripts are searchable, so getting the raw material costs almost nothing. The exercise underneath is the same one.

Get customers talking about the situation they were in before buying. Find out what they were skeptical about when evaluating your offering. Dig into how your offering has helped solved their problems. Note key outcomes that have made a real impact on their business.

The conversations are valuable. Sales learns how customers talk about their situations and start using the actual language on live calls. Messaging gets sharper over the next quarter. Positioning arguments that had been circling the leadership team for months finally settle, because somebody put a real customer sentence on the table.

I’ve watched a lot of teams do some version of this and I’ve never seen one regret it.

There’s a flaw sitting inside the exercise, though. It’s always been there, and until recently it didn’t cost anybody a dime.

It starts with who you recruited for those interviews. Go back and look at where the participants came from. Your customer list. Your CRM. The friendly accounts sales could actually get on the phone. Or, in the modern version, whoever happened to show up in your call recordings, your win/loss interviews, and your support tickets.

Every one of those people already knows who you are and what you sell.

The Change No One Saw Coming

For the past few decades, your buyers have compressed their problems down to three or four words and typed it into a search box. “Marketing attribution software.” That’s category language, and it was the only input that mattered, which is why most of us obsess over keyword data (to this day).

Now with AI, buyers describe their entire situation out loud, in a paragraph or two, to a machine. 

“We’re a 60-person manufacturer, we got acquired last year, and my marketing team still can’t tell me which channels are actually producing revenue.”

Category-based keywords describe a market. That AI-chat language describes a person in a bind, and it’s what’s getting matched now. 

A search engine could bridge the gap between your phrasing and your buyers’ with a synonym index and a decade of click data. 

An AI system assembles its answer out of whatever sources speak to the specific situation somebody described, so the gap doesn’t get bridged for you anymore.

Buyer language stopped being a nice input for messaging workshops and turned into the raw material your entire content strategy gets built from.

☝️ Strange sentence to write after fifteen years of keyword exports, but here we are.

And that raises a question nobody was under much pressure to answer before. Where do you actually get that buyer language?

Every Source You Can Hear Is Already Post-Awareness

The standard answer is a list you’ve read a hundred times. Mine your sales calls. Read your support tickets. Run win/loss interviews. Log the questions that come up on demos.

All real advice, and I’d still do every bit of it. But there’s a structural problem sitting underneath that list of sources, and it’s the same one hiding in that voice-of-customer doc.

Anybody whose words you can capture has already found you.

Landing on a sales call, a demo, or a win/loss interview requires knowing your category exists, knowing you exist inside it, and caring enough to spend an hour on it. 

By the time they open their mouth they’ve absorbed some of your vocabulary. They’re using your product name. They’re using the category label they picked up from whoever they found first.

So your internal sources reliably produce Discovery-stage and Comparison-stage language. Solutions already named, vendors already being weighed. Good material, and most teams have never mined it.

What those sources can’t produce reliably is Problem stage language.

Definition

Problem-Stage Language

How a buyer describes their symptoms before they know what their actual problem is, or what might solve it. No category label, no vendor name, just the felt problem in their own words. It’s the one stage your internal sources can’t capture, because reaching a buyer there requires the category knowledge they don’t have yet.

You have almost zero visibility into someone typing their unique symptoms into an AI chatbot. That person can’t hand you a verbatim quote, and they’ll never turn up in a call recording because booking the meeting requires the category knowledge they don’t have yet.

That blind spot matters more than it used to. So we went looking for what it actually costs.

Across our last 30 AI360™ analyses (our audit tool that shows a company’s visibility across AI answer engines), we ran symptom-first, Problem-stage prompts through Claude, Gemini, GPT, and Perplexity. No category label in the query, no brand name. Just the “why is this happening” and “why is it so hard to” phrasing somebody might use before they know what to call their problem.

By the numbers

Across our last 30 AI360™ analyses of mid-market B2B companies, we ran symptom-first, Problem-stage prompts with no brand or category name.

12%

The typical company appeared in just 12% of those prompts.

40%

Had no Problem-stage presence at all.

Source: Forge & Fathom AI360 analyses (last 30 B2B companies), run across Claude, Gemini, GPT, and Perplexity.

Most of those companies do fine the moment a buyer types a category name. They have content. They rank. Their Comparison-stage coverage is often solid, sometimes excellent. The hole is specific to the one stage nobody can hear.

What Your Internal Sources Are Good For (And Two You’ve Never Opened)

Before writing off the material you already own, get precise about what it does well. Your call recordings and support tickets are the best Discovery-stage and Comparison-stage inputs on earth, and cheaper than anything you could buy.

Line them up by stage and the shape of the problem shows up immediately.

Source Catches Problem-stage buyer language? Language you get
Sales calls and demos No Discovery, Comparison
Win/loss interviews No Comparison
Support tickets No Comparison, post-purchase
Closed-lost notes No Comparison
On-site search Sometimes Discovery
Google Search Console Sometimes Discovery, Comparison

Two sources sitting in your stack are worth more than the rest and almost nobody touches them.

Google Search Console (GSC)

First off, it’s free, probably already set up, and already collecting data on your site. But the important thing is the Performance report that shows the actual query strings people typed into Google’s AI mode before they landed on your site. This data is only AI mode queries, not ChatGPT, Claude, or Perplexity.

This is the one internal source that occasionally catches Problem-stage language, and the reason is worth understanding. Google is really smart. It will match a symptom-describing query to your page even though the person had no idea you existed. 

The Performance > Queries report can show you pre-awareness phrasing from somebody who became aware by accident. Filtering is advised to help you narrow down to prompt-shaped queries that don’t contain your brand or your category label.

Prompts run longer than keywords, so in the Performance report, add a Query filter, choose Custom Regex, and paste in a pattern that only matches long strings. This one catches queries of ten or more words:

The short category terms fall away. What is left reads like a person talking instead of typing. That’s the pile that matters.

As an example, a recent analysis of our own search queries in GSC turned up some interesting data that is driving search impressions and citations for Forge & Fathom:

  • “shortlist vendors that show how our presence differs across ai platforms and explain the likely reasons.”
  • “how does AI visibility affect the buyer’s journey?”
  • “who provides a GTM diagnostic with a ranked to-do list?”
  • “what does a marketing operating model look like with AI?”
  • “do I need to hire new people for AEO?”
  • “what do I risk if I ignore competitors in AI visibility reporting?”
  • “B2B decision-making in the LLM era”

These are all challenges that are impacting our ideal buyer right now, and the language used gives us solid direction for future content development. 

Pro Tip: Bing Webmaster Tools has the equivalent Search Performance report for Bing organic visits. Given the much smaller usage of Bing, the data sample size will be much smaller, but the principle is the same, and it’s worth ten minutes if you haven’t set it up on your website. 

On-site Search Logs 

If your site has a search function, your on-site search data is also free, chronically ignored, and pure unedited buyer phrasing. Somebody arrives on your site, can’t find what they came for, so they type their problem into your search box in their own words with no autocomplete steering them. 

If you have search tracking on in GA4, you already have months of data. If it isn’t, turn it on this week and check back in thirty days. 

Other Typical Sources

Now the obvious ones, which are worth listing only because they’re likely already available to you:

  • Closed-lost reasons and notes in your CRM – Free. Usually a disaster of dropdown values, occasionally a goldmine in the free-text field.
  • Call recordings and transcripts – There are killer tools in this category like Gong, but you can use whatever notetaker your team already runs. Every one of them has a transcript search. 
  • Sales email threads. The first inbound reply, before anybody on your side reframes it, is often the least contaminated sentence in the whole thread.

That last point keys in on a habit that needs special attention. The first thing a person says is worth more than everything after it, because your funnel starts teaching them your vocabulary on contact.

Where Problem-Stage Language Actually Lives

If you can’t hear a Problem-stage buyer directly, we’re left with seeking out the traces they leave in public. Three kinds are accessible, and each one catches the buyer at a different distance from the moment you actually care about.

People Talking to Peers

When somebody senses a problem and has no category label for it, they don’t fill out your form. They go ask strangers.

Subreddit Search

Free. Find the two or three subreddits where your buyer’s role or industry hangs out, then search inside them for symptom words rather than category words. Skip “attribution software” and try “can’t tell what’s working,” or “CEO keeps asking me,” or “no idea where leads come from.” 

Reddit’s native search is mediocre; a site:reddit.com Google search with symptom phrases works better. Paid tools exist that cluster this for you if the manual version stops scaling. 

subreddit example - problem language

Here’s a real example, found in about the time it takes to read this paragraph. Someone posted in r/marketing:

Sales are down – how do I work out if it’s my business or the market?
“YOY sales are down by 50% so far this year… January last year was amazing and exceeded all expectations. This one has hit hard. How do I work out if there’s something going wrong in my industry? Any ideas on how to get a temperature check of an industry? I’ve tried Google Trends but there’s nothing clearly going on there.”

Read what that person did not say.

No category. No product name. No “market analysis,” no “demand research,” no label of any kind. They don’t know those words point to specific solutions yet. What they have is a symptom, sales cut in half, and a question they can’t answer alone: is it me, or is it everyone?

That person will never turn up in your call recordings. Booking time with a vendor requires believing a category of help exists, and they haven’t gotten that far. They’re asking strangers precisely because they don’t know who they’d even be shopping for.

Now look at the phrase they reached for on their own. “A temperature check of an industry.” That is not a keyword and no tool would ever have handed it to you. It’s how a real buyer describes the job before anyone sells them the vocabulary for it. 

Drop that one sentence in front of your content team and it writes itself into posts that nobody optimizing for “market analysis software” will ever surface for.

That is the whole reason to go outside. One line from someone who doesn’t know your category exists is worth more than the fiftieth clean transcript from someone already in your funnel.

Vertical Slack and Discord Communities

Free to join, usually gated by role. The channel where people vent about their week is the highest-signal Problem-stage source I know of, and it’s the one you’ll get the least credit for reading.

Comment threads under posts about the symptom

When somebody publishes about a symptom your buyer has, the comments fill up with people describing their own version of it, unprompted and unpolished.

People Describing the “Before” in Retrospect

This one is underrated because it looks like the wrong source.

Review sites for adjacent categories

G2 and Capterra reviews are free to read, and the section you want is the “what were you doing before” or “what problem did this solve” field. 

People writing those have language for the problem now, and they’re reaching backward to describe a moment when they didn’t. That retrospective phrasing is the closest thing to Problem-stage language you can get in volume. 

The “before” passages in your own customer stories

You wrote these. Go back and read the setup paragraphs, especially any direct quotes from the customer about what life looked like before. Most case studies bury one perfectly good Problem-stage line in the second paragraph and then never use it again.

Query Data the Engines Will Hand You

“People Also Ask” data

Free at small volume through tools like AlsoAsked or AnswerThePublic, both of which have limited free tiers and paid plans past that. These surface the question-shaped phrasing sitting one step upstream of a category term.

Podcast and video transcripts

Free. YouTube auto-generates transcripts on almost everything, and practitioner interviews are full of people describing symptoms in ordinary language because they’re talking to a peer rather than to a vendor. The trick is finding shows based on role rather than topic. Then you can search the transcripts for your symptom phrases.

Don’t Trust Any One of These in Isolation

Every source I mentioned above is a workaround, each with the potential for bias.  

Forum threads self-select for people articulate enough to post. Retrospective passages get reconstructed in vocabulary the writer picked up from whoever they eventually bought from. Query data arrives pre-reworded by the machine. So cross-reference them. 

A phrase that turns up in multiple places is probably real, and one that shows up in a single place is probably an artifact of that place.

Exact Words, Or Don’t Bother

Two methods separate a capture habit from a research project that dies after one quarter.

The first is verbatim discipline. Capture the exact phrasing, including the grammar mistakes, the hedges, and the swearing. One well-intentioned cleanup pass and you’re back where you started, holding a doc full of your own words with a buyer’s name attached.

The second is a repository anyone can add to in fifteen seconds. A shared doc or a Slack channel with a low bar for entry beats a formal research initiative with a kickoff and an end date. Buyer language arrives continuously and in tiny pieces, and if capturing one line requires finding and opening a document and inputting a bunch of details about the interaction, nobody’s going to do that.

Ask for two things per entry: the exact words, and where they came from. Source tells you which stage they were likely in at the time.

The Normalization Trap

Let’s say you’ve been diligent with this capture exercise and you have a few hundred early-stage buyer quotes. It’s completely reasonable to want to organize the list.

Reading through it, you notice forty of these lines are about the same thing, and tag those forty “attribution.”

Your “category language” just walked right back in the front door. This is what we call the Normalization Trap.

Definition

The Normalization Trap

Organizing buyer language in a way that destroys it. Quotes get grouped under a clean internal label, the label becomes the working reference, and the unique phrasing you went out to collect stops getting used.

Microsoft shipped this as a feature this year. Bing Webmaster Tools groups the queries behind AI citations into what it calls Topics, and their own published example collapses “solar panels,” “solar energy efficiency,” and “residential solar installation” into a cluster labeled Solar Energy

Three different people, at three different levels of sophistication, at three different moments in a buying process. One noun.

You captured two hundred quotes specifically because the phrasing mattered. Grouping them is fine, but here’s the rule: the group gets named with a buyer’s sentence, not with your noun.

So that cluster isn’t “attribution.” It’s “I can’t tell my CEO which channels are actually producing revenue.” Longer, uglier, harder to fit in a spreadsheet column, but it survives as variations of a real problem statement.

A Starting Point

Since we’ve covered a lot here, let’s keep it simple with an accessible to-do list for capturing buyer language:

  1. Open Google Search Console and filter for queries with no brand or category term in them. Copy the interesting ones exactly as typed.
  2. Check whether site search tracking is on. Turn it on if it isn’t.
  3. Pick two subreddits or communities where your buyer actually posts. Search symptom phrases, not category terms.
  4. Read the “what were you doing before” field on twenty G2 reviews 
  5. Reread the setup paragraphs of your own case studies and pull any direct customer quote about the before.
  6. Drop everything into one shared doc. Exact words, plus where it came from.
  7. Cross-check before you trust anything. A phrase that showed up in only one of those places is probably an artifact of that place.
  8. Group only when you have to, and name every group with a buyer’s full sentence.

Twenty quotes with the source attached, at least a handful of them from the Problem stage, beats a hundred cleaned-up themes. That’s an afternoon.

The Harder Half

Collecting the language is the easy half, which is an annoying thing to admit at the end of a long post about collecting the language.

Once you’re holding a few hundred real sentences from real buyers across three stages, you have a new problem: nothing in that pile tells you which of those problems you should build against, or what covering one of them actually requires. 

A doc full of excellent quotes and no method for turning them into a publishing plan is exactly the shape of the voice-of-customer project we started with. Sharper messaging, warm feelings, no system.

That’s the next post. How to take a key buyer problem and turn it into a complete topic list, using a grid that makes the gaps in your coverage impossible to miss.

Go read your own search console data first. I’d bet money there’s a sentence in there this week that nobody on your team has ever written about.

P.S. If you’re not sure which buyer problems are worth focusing on in the first place, that’s the conversation worth having first. Give us a shout.

Common Questions

Where do I find the language my buyers use before they know what to call their problem?

Not in your sales calls, demos, or win/loss interviews, because everyone in those sources already found you and absorbed your vocabulary. Problem-stage language lives in the places buyers go before they have a category label: Google Search Console queries that don’t contain your brand or category, on-site search logs, role-based communities on Reddit and Slack, and the “what were you doing before” field on G2 reviews. Capture it verbatim and cross-reference across sources.

Why don’t AI answer engines surface my company when buyers describe their problem?

Because your content is built from category language, and AI systems assemble answers from sources that speak to the specific situation a buyer described, not the market label. A search engine used to bridge that gap with a synonym index and click data. AI doesn’t bridge it for you. If your buyer describes symptoms and your pages only speak in category terms, you’re invisible at the one stage that now matters most.