H

Comparison

HVAC AI Visibility Report: 66 AI Answers, 12 US Metros

Published 2026-08-17Updated 2026-08-1721 min read7 verified sources

HVAC AI visibility, measured: in 66 recorded runs on 2026-08-17, ChatGPT (gpt-5.5) and Perplexity (sonar-pro) named at least one specific HVAC company in every answer — 194 distinct brands across 12 US metros. In matched pairs of answers, the two engines shared a mean 0.1383 of all brands either one named; 8 of 24 pairs shared none.

This report publishes a first-party prompt panel, not a survey and not a ranking study. HVAC AI visibility here means one thing only: whether a web-connected AI engine names your company in its answer to a buyer's question. GEO means generative engine optimization — trying to earn those mentions — never geographic targeting. Metro means the city written into the prompt string, not the location of a real asker, because no geolocation was applied.

Two things this page deliberately does not do. It does not estimate how many homeowners ask AI in the first place; that is survey territory, covered in the survey evidence on how many homeowners actually ask AI. And it does not explain why any company was named. The panel recorded 66 answers and measured nothing about the companies inside them. Every observation below is a property of the answers, not of the businesses.

01

How we ran the study (methodology)

The panel sent two buyer-intent prompts to two web-connected commercial LLM endpoints across 12 US metros on 2026-08-17, capturing every response verbatim inside a 5.5-minute window (14:24:35 to 14:30:05 UTC). All 66 requests returned status code 20000; none failed. Total API cost was $3.044076, summed from the per-run cost field rather than estimated.

Design elementWhat we actually ran
Runs recorded66 total — 48 main grid, 18 repeat runs
Main grid12 metros × 2 prompts × 2 engines, n=1 per cell
Repeat subset3 metros (Phoenix, Atlanta, Chicago) × prompt P1 × 2 engines × 3 extra repeats — n=4 in 6 cells
EnginesChatGPT gpt-5.5-2026-04-23 (web search on, reasoning on); Perplexity sonar-pro (web search on)
MetrosLos Angeles, Houston, Phoenix, Tampa, Atlanta, Chicago, Dallas, Philadelphia, San Antonio, Las Vegas, Charlotte, Orlando
TransportDataForSEO MCP, ai_optimization_llm_response, API version string 0.1.20260806
Endpoint failures0 of 66
Cost$3.044076

The two prompts were sent exactly as written, with {metro} substituted and nothing else added:

  1. P1 — "Who is the best HVAC company in {metro}?"
  2. P2 — "My AC stopped working in {metro}. Which company should I call?"

Each run is a fresh single-turn request: no system prompt, no chat history, no supplied search results. Fifteen fields were recorded per run, including the DataForSEO task id, the full verbatim response text, the ordered list of every business named, every citation URL returned, whether the answer pointed the reader to a directory or registry, and any company the answer cautioned against.

Business names were read off the verbatim response text by hand, not by a classifier, and brand variants were merged only through an explicit alias list — never by fuzzy matching. That is why this dataset counts "Legacy Air Conditioning and Heating" and "Legacy Heating, Cooling, Plumbing & Electrical" as two brands: same-entity identity could not be verified in session, so we did not assume it.

What went wrong, and where the design is weak

One design gap belongs on the record before anyone cites this. The panel's build note recorded location as "not settable on this endpoint." That is wrong. DataForSEO's documentation for the ChatGPT llm_responses endpoint lists both web_search_country_iso_code and web_search_city, and the Perplexity endpoint exposes web_search_country_iso_code for Sonar models. Those fields existed and we left them unset, so the honest statement is "geolocation was available and unused," not "geolocation was impossible." A homeowner physically in Phoenix may see different answers than an unlocated API call that merely says "Phoenix." Version 2 sets those fields.

Sampling parameters were also left at endpoint defaults, which are not zero: DataForSEO documents a default temperature of 0.94 and top_p of 0.92 on the ChatGPT endpoint, and 0.77 / 0.9 on Perplexity's. One qualifier belongs with that ChatGPT figure — the same documentation adds that temperature is "not supported in reasoning models," and our ChatGPT runs used a reasoning model, so the 0.94 default cannot be assumed to have applied to them. Non-zero sampling is one obvious contributor to the run-to-run movement reported below, and this design cannot separate that contribution from index churn.

The full limitation set:

  • Single day. All 66 runs sit inside one ~5.5-minute window. Nothing here describes a trend or a stable state.
  • n=1 in 42 of 48 cells. Each main-grid number rests on one observation. The repeat subset shows single observations are unstable, so read every main-grid cell as one draw, not a measurement.
  • n=4 in 6 cells. Four observations demonstrate that variance exists; they cannot estimate it.
  • Not a ranking-factor study. No attribute of any recommended company was measured, so no causal or correlational claim about what earns an AI recommendation is supported.
  • Two engines only. Google AI Overviews, Gemini, Claude and Microsoft Copilot were not tested, though the endpoint exposes claude and gemini families. Scope stayed at two engines so the repeat subset could run properly.
  • APIs, not apps. gpt-5.5 via API is not the ChatGPT app; sonar-pro via API is not perplexity.ai in a browser. App-level personalization, memory, ad units and UI ranking are absent.
  • No dated Perplexity build. The endpoint exposes no dated build for Sonar models, so the version behind "sonar-pro" on 2026-08-17 cannot be pinned.
  • Analyst-assigned labels. Business-type and citation-class buckets are judgment calls from in-session evidence. No filings, franchise disclosure documents or ownership registries were checked; another analyst would bucket some domains differently and move the percentages.
  • Citation counts are not volume-comparable across engines. ChatGPT repeats a URL across claims, so its per-run lists were de-duplicated; Perplexity's count is the source list it returned, and that list came back at 20 URLs in 30 of its 33 runs and 19 in the other three — a ceiling on what the engine reports, not a count of what it consulted.
  • Nothing the models asserted was verified. Star ratings, review counts and awards inside the answers were recorded as model output only. Several are likely wrong.
  • Absence proves nothing. A company unnamed in these 66 runs may still be recommended on another day, phrasing, or engine.

The prompts, model ids, parameters and task ids are all recorded, so the calls are re-issuable. The answers are not reproducible: these endpoints are non-deterministic and the underlying web index changes daily.

02

Which companies AI recommends most, by metro

Every one of the 66 runs named at least one specific company. None refused, and none substituted a directory for a name. The two engines produced 334 brand mentions resolving to 194 distinct brands, at a mean of 5.06 businesses per run — 5.39 for ChatGPT, 4.73 for Perplexity.

Read the run column first in the table below, which is counted from the panel's raw run ledger. Phoenix, Atlanta and Chicago carry 10 runs each because they were the repeat-subset metros, so their brands had more chances to be named. Cross-metro comparison of these counts is not valid.

The right-hand column follows one fixed rule and breaks no ties: it lists every brand named in 3 or more of that metro's runs, and where no brand cleared 3 it lists every brand named twice instead. Brands named exactly twice are otherwise omitted — between one and six of them per metro — because picking a subset of tied names would be an editorial choice this dataset cannot justify.

MetroRunsDistinct brands namedBrands named in 3+ runs (runs naming them)
Los Angeles412Brody Pennell (3)
Houston415none above 2 — Cool Care Heating and Air Conditioning (2), Air Tech of Houston (2), Abacus (2)
Phoenix1028Island Breeze (5), Goettl (5), Benefit Air (5), Day and Night (4), Howard Air (3)
Tampa418Air Masters of Tampa Bay (4)
Atlanta1018Moncrief (9), TE Certified (9), Estes Services (8), PV Heating (6), Hope Heating (4), Reliable Heating and Air (4)
Chicago1017All Temp (9), King Heating (5), Air-Rite (4), Deljo (4), Hero Air (4), Four Seasons (3)
Dallas416Texas Airzone (3)
Philadelphia411Airmaster (3), Gen3 (3)
San Antonio412Cowboys AC (4)
Las Vegas416none above 2 — Super Service (2), Nevada Residential Services (2), Aloha Air Conditioning (2), Angel's Air (2)
Charlotte414Horne Heating and Air (4), Morris-Jenkins (3), McClintock (3)
Orlando421Downtown Air and Heat (3)

These are the panel's canonical merged labels, not the companies' legal names, and we verified nothing about any of them. Moncrief and TE Certified were each named in 9 of Atlanta's 10 runs, which tells you the engines were consistent about those two names inside one five-minute window on one day. It does not tell you either company did anything to earn it.

Concentration matters more than any single name: 125 of the 194 distinct brands were named in exactly one run. Roughly two-thirds of the companies these engines produced appeared once and never again. Position needs the same care — ordered lists reflect the order names appear in the answer text, and the models rarely state a strict ranking, so position 1 means "named first," not "ranked first."

Two of the 66 runs did the opposite of recommending: the answer text cautioned the reader against a specific named company. We are not reprinting those claims, because the panel records them as unverified model output and no rating, review count or award in any answer was checked against Yelp, Angi, BBB or the company itself. The behaviour is the finding — these engines will occasionally warn a homeowner away from a named contractor, on evidence nobody has audited.

03

Nothing this panel can prove. That is the honest headline, and it is the sentence most AI-visibility marketing gets wrong. The panel recorded what two engines said; it inspected no website, review profile, schema block or backlink belonging to any of the 194 brands. No tactic can be credited here — not by us, and not by anyone quoting this dataset.

A separate measurement pass on 2026-08-17 did inspect 94 of those brands — Google listing, homepage reachability, robots.txt and homepage JSON-LD — and its findings are published as an observational profile of what the AI-named HVAC companies actually look like. That file measures companies, not answers, and it establishes no cause either.

What the panel can describe is the shape of the answers and the evidence surfaces the engines reached for.

The named companies were overwhelmingly small and local-looking. Of 194 distinct brands, 190 were labelled apparent single-market contractors, 3 were multi-market brands (Cool Today, Goettl, One Hour Heating and Air Conditioning) and 1 was a retailer. By mentions: 323 single-market, 10 multi-market, 1 retailer. A brand earned the multi-market label only where a cited page on its own domain used a per-metro location path, proving multi-metro operation. "Apparent" means exactly that — brands labelled single-market may be roll-up or franchise owned, since no ownership registry was checked.

Third-party review platforms dominated the citations. The 66 runs returned 767 citation URLs, 465 of them unique, spanning 200 unique domains.

Citation mix by class for each engine: ChatGPT drew 80 percent of its citations from third-party directories and 2 percent from the contractor's own site, against 56 and 32 percent for Perplexity.

Citation classAll citationsShareChatGPTPerplexity
Third-party directory or review platform45759.58%88369
Vendor-owned page21027.38%2208
Community forum668.60%1848
Vendor-owned "best of" listicle303.91%030
Government / regulator20.26%20
Retailer20.26%02

The two engines behaved nothing alike as citers. ChatGPT returned 110 citation URLs across its 33 runs — 84 unique, from just 22 domains, a mean of 3.33 per run, with a per-run range of 2 to 6. Perplexity returned 657 across its 33 runs — 381 unique, from 191 domains — but its per-run count barely moved: 20 URLs in 30 of the 33 runs and 19 in the other three, a floor of 19 and a ceiling of 20. A list that comes back at exactly 20 URLs in 30 of 33 runs is a capped list, not a measure of how widely the engine searched, so the 110-against-657 gap is not evidence of citation breadth and should not be quoted as if it were. The contrast that survives the cap is domain diversity: 22 domains for ChatGPT against 191 for Perplexity.

Most-cited domainCitation URLsRuns it appeared in
reddit.com6644 of 66
consumeraffairs.com4848 of 66
expertise.com3929 of 66
serviceagent.ai3729 of 66
angi.com3331 of 66
bbb.org2925 of 66
hvacservice.io2320 of 66
forbes.com2323 of 66
bestpickreports.com2121 of 66
downtobid.com1818 of 66

Reddit as the most-cited domain is the finding most worth dating, because it inverts the closest prior benchmark. BrightLocal's November 2024 study of 800 manual ChatGPT local searches reported that "Forums like Reddit and Quora did not appear as sources throughout this study," alongside a mix of 58% business websites, 27% business mentions and 15% directories. Twenty-one months later, on a different engine pair and a different unit of measurement, reddit.com contributed 66 citation URLs and appeared in 44 of our 66 runs. BrightLocal classified source types across manual searches; we counted API-returned citation URLs — so treat these as two dated snapshots pointing opposite ways, not a trend line. Expertise.com appears in both: 18% of BrightLocal's directory sources then, third most-cited domain here now.

Contractor-owned "best HVAC company in {city}" pages did get cited — rarely. Thirty citations across 18 distinct URLs on contractors' own domains matched the best/top/compare/ranking path rule, including four pages from one San Antonio contractor and three from one Orlando contractor. That class is 3.91% of all citations. Whether publishing such a page changes whether an engine names you is not something this panel can test, and the small share should temper confidence either way.

Directory pointers came only from ChatGPT, and they supplemented named contractors rather than replacing them. Seven of the 66 runs (10.61%) told the reader to go to a directory, review platform or state registrar as an action — Yelp, Angi, BBB, Consumers' Checkbook. All seven were ChatGPT runs: 7 of 33 ChatGPT runs (21.21%) against 0 of 33 Perplexity runs, a split the panel-wide 10.61% hides. In zero of those runs did the directory stand in for a named company. The pattern is a shortlist plus a verification step, as in run r005 (ChatGPT, Phoenix, P1): "There isn't one objectively 'best' HVAC company in Phoenix, but my short answer would be: start with North Valley Mechanical HVAC & Plumbing or Larson Air Conditioning, then get at least 2–3 quotes before authorizing major work."

04

How often engines disagree with each other

The two engines disagreed far more than they agreed. Across 24 matched cells — same metro, same prompt, same day, one ChatGPT run against one Perplexity run — mean overlap was 0.1383. Eight of the 24 cells shared no brand at all, and only four cleared 0.25.

Overlap here is the Jaccard index: brands both engines named, divided by all distinct brands either named.

Cross-engine agreement on which HVAC companies to name, 24 metro-prompt cells, ChatGPT gpt-5.5 against Perplexity sonar-pro: 8 cells shared no company at all and mean overlap was 0.138. A cell at 0.0 means the two engines produced entirely separate lists of HVAC companies for the identical question, minutes apart.

MetroPromptChatGPT namedPerplexity namedSharedOverlap
Los AngelesP14300.0
Los AngelesP24520.2857
HoustonP17410.1
HoustonP22510.1667
PhoenixP14600.0
PhoenixP24300.0
TampaP111610.0625
TampaP24510.125
AtlantaP15620.2222
AtlantaP24620.25
ChicagoP15410.125
ChicagoP24510.125
DallasP16410.1111
DallasP25600.0
PhiladelphiaP15420.2857
PhiladelphiaP23400.0
San AntonioP15410.125
San AntonioP25540.6667
Las VegasP18400.0
Las VegasP24400.0
CharlotteP17430.375
CharlotteP26520.2222
OrlandoP110510.0714
OrlandoP26600.0

San Antonio's P2 cell is the high-water mark at 0.6667 — four shared brands out of six distinct. Las Vegas produced zero overlap on both prompts. With one run per engine per cell, no individual number here is a measurement; the distribution is the point.

The same engine, asked four times

The repeat subset re-ran prompt P1 four times per engine in three metros, all inside roughly six minutes. Only one of the six cells kept the same first-named brand across all four runs.

MetroEngineFirst-named brand, runs 1→4Brands named ≥1Named in all 4Mean pairwise overlap
PhoenixChatGPTNorth Valley Mechanical → Day and Night → Day and Night → Day and Night1710.1339
PhoenixPerplexityBenefit Air → Island Breeze → Island Breeze → AC and Plumbing Doctors830.5772
AtlantaChatGPTPV Heating → PV Heating → PV Heating → TE Certified1320.3586
AtlantaPerplexityHope Heating → TE Certified → Hope Heating → Hope Heating650.9167
ChicagoChatGPTAll Temp → All Temp → All Temp → All Temp1210.2632
ChicagoPerplexityHero Air → Deljo → Hero Air → Deljo540.9000

The engines differ in kind, not degree. ChatGPT's mean pairwise overlap across the repeat cells was 0.2519; Perplexity's was 0.7980. Phoenix on ChatGPT is the extreme: 17 distinct brands surfaced across four identical requests, 15 of them in only one of the four, and one brand present in all four. Perplexity behaved like a narrow, stable shortlist that reshuffled its order; ChatGPT behaved like a wide, shifting pool.

Run-to-run agreement for four repeats of an identical HVAC prompt in three metros: ChatGPT averaged 0.25 agreement with itself while Perplexity averaged 0.80.

Rewording the question changes the answer too

Prompt framing moved the output as much as engine choice did. Across the 48 main-grid runs, P1 ("Who is the best HVAC company in {metro}?") produced 131 brand mentions and 117 distinct brands, a mean of 5.46 per run, pointing to a directory in 4 runs. P2 ("My AC stopped working in {metro}. Which company should I call?") produced 110 mentions and 95 distinct brands, a mean of 4.58 per run, pointing to a directory in 2 runs. Comparing P1 against P2 within the same engine and metro, mean overlap was 0.1733, with 8 of 24 pairs sharing no brand.

Multi-market brands drew more mentions from the emergency-framed prompt: 2 mentions under P1 against 5 under P2. That is five mentions in a 334-mention dataset — an observation to test properly, not a finding to act on.

None of this is unique to HVAC. SISTRIX's Johannes Beus, whose team tracked 82,619 prompts across 17 weeks and six countries, measured weekly source replacement of 5% in AI Overviews, 56% in Google AI Mode and up to 74% in ChatGPT Search, and reached the conclusion this panel independently walked into: "a single citation placement is not a reproducible result but merely a snapshot." His companion line belongs on every AI-visibility proposal: "GEO is a continuous process, not a one-off optimisation with permanent results."

One quirk shapes how these answers read. Counting every Perplexity response whose recorded text contains the exact phrase "you provided", 8 of the 33 Perplexity runs credited search results or sources the caller never supplied: r002, r006, r008, r020, r024, r050, r056 and r062. Run r006 (Phoenix, P1) contains the phrase "Based on the results you provided, the strongest contenders are". Every run was a fresh single-turn request; nothing was provided. That is the model's own framing of its retrieval, recorded verbatim and uncorrected — and it is why the run ids are published, so the count can be re-run against the ledger rather than taken on trust.

05

What HVAC owners should do with this data

Treat this panel as a measurement pattern to copy, not a scoreboard to react to. One run tells you nothing, one day tells you nothing, and one engine tells you about one engine. The defensible reading, in order:

  1. Stop chasing a single answer. With 8 of 24 matched cells at zero overlap and only 1 of 6 repeat cells holding its first-named brand, any screenshot of "ChatGPT recommends us" — or of a competitor — is one draw from a distribution.
  2. Measure with a panel, not a query. Fix the prompts, fix the engines, log every run, repeat on a schedule. Our 66 runs cost $3.044076; the scarce resource is the discipline of recording verbatim output, not the API budget. And set the geolocation fields your endpoint exposes — the gap we published above is the one to avoid.
  3. Fix eligibility before anything clever. Google's own documentation on AI features is unambiguous about the floor: "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. There are no additional technical requirements." The same page notes AI Overviews "are only shown when our systems determine that it is additive to classic Search." Start by confirming crawlers reach you — test whether these engines' crawlers can even fetch your site in about a minute.
  4. Take third-party review surfaces seriously, on evidence. Third-party directory and review platforms carried 59.58% of the 767 citations here, and consumeraffairs.com appeared in 48 of 66 runs. That describes where these engines were looking on one day, not proof that a profile edit changes an answer.
  5. Expect the answer to move, and budget for it. Between SISTRIX's up-to-74% weekly churn in ChatGPT Search and the within-six-minute movement recorded here, AI visibility is a monitoring commitment, not a project with an end date.
  6. Refuse guarantees, including ours. Nobody can promise you a slot in an AI answer, and this dataset is why: one engine, one question, four times in six minutes, 17 different companies in a single metro.

That last point is the basis of how we work — measuring one contractor's AI answers month over month against a fixed prompt panel, reporting what moved and what did not, with the raw runs attached. That panel is not sold separately: it sits inside the $2,500/month retainer on our pricing page, because a measurement you can switch off is a measurement nobody runs. For the engine-specific mechanics of being findable at all, the companion piece is the tactics guide for ChatGPT specifically.

FAQ

Frequently asked questions

What does AI visibility mean for an HVAC company?

AI visibility means a web-connected AI engine names your company in its answer when a homeowner asks who to call. It is distinct from ranking: an engine can name a company whose website never reaches the top 10 blue links, and can omit one that ranks first. This panel measured only the naming behaviour of two engines on one day.

How many metros and engines did this study actually cover?

Twelve metros, two engines, two prompt phrasings, one day, 66 runs. The metros are Los Angeles, Houston, Phoenix, Tampa, Atlanta, Chicago, Dallas, Philadelphia, San Antonio, Las Vegas, Charlotte and Orlando. Results do not generalize beyond that set, and Google AI Overviews, Gemini, Claude and Copilot were not tested.

How does this compare to the other HVAC AI visibility index?

The HVAC & Plumbing AI Visibility Index 2026 from 5W Public Relations covers a wider surface: 65+ prompts through ChatGPT, Claude, Perplexity and Google AI Overviews over January–March 2026, spanning two trades. Its headline stat, quoted in full — "87% — estimated share of independent HVAC and plumbing contractors with effectively zero AI citation share in their own metro and category" — is labelled an estimate and scoped to a contractor's own metro and category, and its unit is brand citation share rather than brands named per answer, so the two datasets are not comparable figure to figure. This panel is narrower by design and publishes its per-run ledger, prompt text and task ids so the method can be audited.

Why do ChatGPT and Perplexity name different companies?

Different indexes, different retrieval, different sampling. Perplexity's citations spanned 191 domains across its 33 runs against ChatGPT's 22 domains, which is the comparison that holds; the raw volumes (657 against 110) do not, because Perplexity's returned list was capped at 20 URLs per run. Sampling was not deterministic either: DataForSEO documents default temperatures of 0.94 for ChatGPT and 0.77 for Perplexity, though it also notes temperature is "not supported in reasoning models," and our ChatGPT runs used one. This panel records the disagreement; it does not diagnose it.

Can I reproduce these results?

The calls, yes — prompts, model ids, parameters and task ids are all recorded. The answers, no. These endpoints are non-deterministic and the underlying web index changes daily, which is what the repeat-run section quantifies: four identical Phoenix requests to ChatGPT surfaced 17 distinct companies, 15 of them only once.

Is AI search big enough to matter for HVAC yet?

Adoption should be read from surveys, not from this panel. BrightLocal's 2026 Local Consumer Review Survey found 45% of US consumers used AI tools for local business recommendations, up from 6% a year earlier, with ChatGPT at 31% and Google AI Mode at 23%, from a panel of 1,002 US adults. Google still ranks first as a recommendation source in that survey; AI is third.

Talk to us before you decide.

A 30-minute strategy call with your market's real data. No pressure, no 12-month contracts.