Generative engines —ChatGPT, Perplexity, Claude, Gemini, Copilot— ignore your firm because its content is not ingestible: no structured Schema.org JSON-LD, a robots.txt with no explicit permission for AI crawlers and no llms.txt as a knowledge map. This is not an SEO problem. It is semantic infrastructure.
1 · The symptom: how to verify in 3 minutes that your firm is invisible
Before theorising, measure. Open five tabs with ChatGPT, Claude, Perplexity, Gemini and Microsoft Copilot. In each one, run this query literally, replacing the brackets:
best [your practice area] firms in [your city] 2026
Count how many of the five engines cite your firm within the answer. Not in the source links — in the body text of the answer. That number is your Citation Presence baseline. In firms with no GEO infrastructure, the most common result in 2026 is 0/5 or 1/5. If your firm bills €1-10M and your Citation Presence is zero, you are technically invisible to the 67 % of B2B decision-makers who use generative engines as their first research source before picking up the phone (KnewSearch, The State of B2B Decision-Making in the AI Era, 2026).
Repeat the exercise with a variant using your own brand:
[Exact name of your firm]
If this second test also returns vague answers, or confuses your firm with a namesake, the problem is not one of relative visibility. It is one of structured identity. The generative engine has no node in its graph identifying your firm as an entity. You are asking the AI to cite someone who, as far as it is concerned, does not exist.
This three-minute check is what we at AuditScale call the Bar Test T+0: the empirical measure of the current state before any intervention. It is the baseline that every serious GEO project documents with timestamped screenshots across the five engines.
2 · Differential diagnosis: why this is not an SEO problem
The immediate reaction of most firms on seeing their Bar Test T+0 result is to assume the problem is SEO and to call their digital marketing agency. That is a category error.
Classic SEO optimises so that a human user clicks your link on the Google results page. The metric is position 1-10 in the SERP. A generative engine does not work that way. ChatGPT does not return ten links. It returns one answer, synthesised from sources its model judged authoritative at the moment of inference. Either your firm appears inside that answer —cited by name, with your URL as the source— or you do not exist in that exchange.
The difference has measurable technical consequences. In February 2026, the overlap between the pages Google ranks in its top ten and the pages ChatGPT, Perplexity and Claude cite in their answers fell from 76 % to 20 % in the B2B professional services segment (aggregated data from Profound, SimilarWeb and SemRush AI Search Index, Q1 2026). In other words: optimising for page one of Google no longer guarantees visibility in the AI answer. These are two diverging discovery infrastructures that share less and less content.
To this we can add one relevant sector figure: 52 % of the law firms measured in the first GEO benchmark of British firms (GTM Signal Studio, UK Law Firms GEO Readiness Index, Q1 2026) scored 2/25 or less on AI citation presence, despite having active classic SEO budgets. Spend on traditional SEO is not converting into generative visibility.
Your firm may have impeccable PageRank, quality backlinks, extensive content and Core Web Vitals in the green — and still be invisible to ChatGPT. Because the problem is not ranking. It is the engine's ability to ingest your content as a citable structured entity.
3 · Truth table · Traditional SEO vs GEO
| Dimension | Traditional SEO | GEO (Generative Engine Optimization) |
|---|---|---|
| Unit of success | Click through to the website from the SERP | Citation inside the AI engine's answer |
| Metric | Position 1-10 in Google | Citation Presence + AI Share of Voice |
| Primary optimisation | Keywords + backlinks + Core Web Vitals | Schema.org JSON-LD + llms.txt + explicit welcome for AI crawlers |
| Relevant crawlers | Googlebot, Bingbot | GPTBot, ClaudeBot, PerplexityBot, CCBot, Google-Extended |
| Technical validator | Google PageSpeed, Search Console | geo_scanner.py + manual Bar Test across 5 engines |
| Validation cycle | 4-12 weeks (Google algorithm) | 14-21 days (re-indexing after IndexNow + GSC) |
| Winning content | Long-form keyword-optimised, 2,000+ words | Direct answer + structured data + triangulated authority |
| Typical silent failure | Ranked on page two (indexed, invisible) | Indexed but not cited (high PageRank, no valid schema) |
The table does not imply that SEO is obsolete. It implies that SEO and GEO are parallel disciplines with different technical stacks. The newsrooms of leading media outlets have been building GEO teams separate from their SEO teams since Q4 2025. Professional services firms have not.
4 · The root cause: the four infrastructure failures that produce algorithmic blindness
When a firm scores 0/5 or 1/5 on the Bar Test, the reason almost always comes down to four infrastructure failures, in this order of severity.
4.1 · Schema.org JSON-LD missing or malformed
Schema.org is the standardised semantic vocabulary that Google introduced in 2011 together with Bing and Yahoo. Injected into the <head> of your HTML as a <script type="application/ld+json"> block, it tells the engine exactly what kind of entity you are (@LegalService, @ProfessionalService, @AccountingService), your legal name, structured postal address, E.164 telephone number, geographic service area and the URLs of verifiable external profiles (sameAs).
Without this block, the generative engine processes your home page as plain text and has to infer your identity. That inference fails more than 80 % of the time in firms whose name contains common words (surnames, "associates", "solicitors", "consulting"). The node "your firm" is never created in the engine's knowledge graph — or it is created badly, merged with a namesake.
The AuditScale scanner (geo_scanner.py) audits six critical Schema.org fields and returns a weighted 0-100 score. More than 70 % of the Spanish professional firms audited so far in 2026 score 0/100 because the block is not injected at all, not because it is badly written.
4.2 · robots.txt with no explicit welcome for AI crawlers
robots.txt is the file at the root of the domain that tells bots what they may crawl. Most firms have a robots.txt inherited from their CMS that allows Googlebot but does not explicitly mention the crawlers of the AI engines: GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot, CCBot (Common Crawl, the basis of many LLMs), Google-Extended (which separates classic indexing from indexing for Gemini training).
OpenAI published in August 2023 that GPTBot respects robots.txt and only crawls sites that explicitly allow it. Anthropic followed in April 2024 with ClaudeBot. If your robots.txt does not contain lines such as:
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
...those crawlers assume ambiguous permission and, in many cases, do not come in. The result: your content never reaches the engine's training or retrieval set. You are blocking the entrance and then asking why nobody arrives.
4.3 · No llms.txt as a knowledge map
llms.txt is an emerging standard proposed by Jeremy Howard in September 2024 (llmstxt.org) that performs for AI engines the function sitemap.xml performs for Google: a manifest at the root of the domain describing in natural language who you are, what you do, which key documents the engines should read in order to understand you, and what citation policy you apply.
Unlike sitemap.xml, which is structural XML enumerating URLs, llms.txt is readable Markdown that explains the site in semantic terms. A good llms.txt has six to eight sections: identity, services, guarantees, documentation, contact, citation policy. The engine processes it once and obtains dense context that would otherwise take thousands of fetches to build.
As of today (May 2026), fewer than 3 % of Spanish B2B corporate websites publish llms.txt. Those that do gain a disproportionate advantage in contextual retrieval.
4.4 · sameAs missing: no authority triangulation
Within the Schema.org JSON-LD block, the sameAs field declares the external URLs that confirm your identity: your company LinkedIn page, the partners' personal LinkedIn profiles, GitHub if you use it, profiles in verifiable professional directories, cross-market domains of the same firm.
Without sameAs, the generative engine has only your word —that of your own domain— as proof that you exist. With sameAs triangulating against LinkedIn (which the engine trusts very highly), the Google Knowledge Graph (which the engine consults) and other verifiable domains, your identity solidifies in the AI's graph. A sameAs field with three to five verifiable external URLs raises the probability of being cited by name by between 20 % and 40 %, depending on the query category.
5 · The AuditScale method · three actionable steps, not recommendations
The four failures in the previous section are symptoms. The AuditScale method turns them into measurements. Three steps, in this order, each of which produces a number or a table — not a PDF of good intentions.
Step 1 · Semantic infrastructure audit
Automated technical scan via geo_scanner.py across the six critical Schema.org fields: name, address (with its four sub-fields: streetAddress, postalCode, addressLocality, addressCountry), telephone (E.164 validation), openingHours, areaServed (single vs plural), sameAs (≥2 external URLs). Each field is graded in three states: green (1.0), amber (0.5), red (0.0). Final score: simple arithmetic over six. Output: 0-100, quantified and reproducible.
To this we add the HTTP status of robots.txt, llms.txt and sitemap.xml, and the presence or absence of an explicit welcome for the main AI crawlers. Total time: 90 seconds. Marginal cost: zero.
Step 2 · Empirical Bar Test T+0
Selection of 5-6 queries representative of your firm's specific sector and location (examples from the Spanish market: M&A Madrid, tax Barcelona, employment Valencia) plus one query with your exact brand name as an identity anchor. Manual execution across the five generative engines: ChatGPT, Claude, Perplexity, Gemini and Microsoft Copilot. Timestamped screenshots per query × engine. Final count: X / N, where N is the total of queries × engines and X the actual citations by name.
We coined the term Bar Test in May 2026, inspired by the documented case of an independent consultant in Austin who generated $49,200 in six months auditing law firms with an equivalent method. The idea is the same: pass the bar test — if your firm is not cited in AI answers when a decision-maker asks in natural language, you do not exist for them.
Step 3 · Competitive triangulation
Identification, via search and verification (WebFetch or equivalent), of 3-5 real competitors in the same vertical and location. For each one: parsing of their Schema.org JSON-LD, a fetch of their robots.txt, a check for a welcome for AI crawlers, and a record of their public pricing or explicit lack of it.
The end product is a competitive table in which your firm is compared across five or six technical dimensions with real, verifiable competitors. Along the way it detects what we at AuditScale call technical hypocrisy: competitors who sell digital transformation or AI services on their website but do not welcome GPTBot in their own robots.txt. It is the most easily usable commercial insight and it appears in 30-40 % of audits.
The three steps together produce a documented GEO Dashboard with a technical score, a Bar Test T+0, a competitive map and a remediation roadmap. It is the baseline against which any later intervention is measured.
6 · Evidence: the AuditScale case study (dogfooding)
Before auditing third parties, we applied the method to AuditScale itself. On 9 May 2026, the initial audit of auditscale.es returned a technical score of 33/100 and a Bar Test of 0/5. Three technical dependencies explained the whole picture: partial schema with no triangulated sameAs, a robots.txt with no explicit welcome for AI crawlers, and no llms.txt at all.
We documented the full remediation process, with commits, deploys and measurements per iteration, in this article: A GEO audit applied to ourselves. Worth reading before starting a similar project at your own firm.
The result as of 27 May 2026 (18 days after the initial audit): technical score 100/100 verified by an independent scanner; complete robots.txt + llms.txt + sitemap.xml infrastructure plus a Schema.org JSON-LD @graph with dual @type; sameAs triangulating against five verifiable external URLs. The effective Bar Test across the five engines is measured at T+14d from the Bing IndexNow submission and the manual indexing request in Google Search Console — checkpoint scheduled for 10 June 2026.
We publish the full numbers of the process, including the failed attempts. We do not publish success stories; we publish reproducible audits. The distinction matters in an industry —digital marketing— where the success story with no verifiable method is the norm.
7 · Frequently asked questions
Does GEO replace SEO or sit alongside it?
It sits alongside it. They are two parallel disciplines with different technical stacks. SEO is still needed to capture the traffic that arrives via classic Google (still the majority in low-ticket transactions). GEO captures the traffic —and, more importantly, the decisions— of the B2B decision-makers who already consult generative engines as their first research source. Cancelling SEO in order to invest in GEO is a mistake. Ignoring GEO because "we already do SEO" is another.
How long does it take to see the effect of a GEO audit once it has been implemented?
The technical score changes immediately: as soon as valid Schema.org is injected and robots.txt/llms.txt are published, the scan returns the new number. The effective Bar Test —that is, real citations in answers from ChatGPT, Claude, Perplexity, Gemini and Copilot— requires a re-indexing period of 14-21 days from the Bing IndexNow submission and the manual request in Google Search Console. There is no shortcut. The engines process in batches and the models' training data is updated in cycles.
Do I have to rebuild my website to implement GEO?
No. More than 90 % of GEO interventions are surgical injections into specific files: a <script type="application/ld+json"> block in the <head> of the HTML, a rewrite of robots.txt, the creation of llms.txt at the root, an adjustment to sitemap.xml. None of them require touching the design, the layout or the CMS. In firms with a well-built static site, the full implementation fits into a single development session.
How do I objectively measure whether my firm appears in ChatGPT?
Three levels, in increasing order of rigour: (a) manually run 5-6 queries across the five generative engines and count citations by name (Bar Test); (b) monitor monthly with the same set of queries to detect variation (what we at AuditScale call the AI Authority Monitor); (c) run a quarterly technical audit of the Schema.org and the semantic infrastructure to anticipate regressions (firms on a CMS accidentally modify the <head> fairly often and break the JSON-LD block).
Does GEO work for small firms or only for large ones?
It works in inverse proportion to size. Large firms have a pre-existing brand advantage: the AI engine already knows them. Boutique firms of 5-50 employees, with no established brand in the AI's graph, are the ones that gain the greatest percentage improvement from implementing GEO correctly. Moving from 0/5 to 3-4/5 in Citation Presence is the difference between invisible and citable. That is where the early competitive advantage is consolidated — before the rest of the sector understands the shift.
8 · Next steps · your own diagnosis in 60 seconds
Before hiring anyone, measure your own position. Three terminal commands give you a verifiable initial diagnosis:
# 1 · Does your robots.txt welcome the AI crawlers?
curl -sL https://yourfirm.com/robots.txt | grep -E "GPTBot|ClaudeBot|PerplexityBot"
# 2 · Do you have Schema.org JSON-LD on your home page?
curl -sL https://yourfirm.com/ | grep -i "application/ld+json"
# 3 · Do you publish llms.txt?
curl -sI https://yourfirm.com/llms.txt | head -1
If the first command returns nothing, the three main AI crawlers have no explicit permission from your site. If the second returns no lines, your site has no semantic infrastructure for the engines to identify you as an entity. If the third returns 404, you are not publishing a knowledge map for the LLMs.
Three consecutive 0s = complete algorithmic blindness. Three 1s do not guarantee citation — they guarantee that the door is not shut for trivial technical reasons.
Update (July 2026): the wall came down, and it is measured
This article was written when ChatGPT answered "I cannot find any company by that name" when asked about AuditScale. Two months later, without changing the strategy —only executing it— ChatGPT describes the entity correctly: Spanish company, GEO, services. In parallel, the site went from zero citations in AI answers to more than three hundred. The wall this article describes is not permanent: it is infrastructure and authority, and both are corrected with method and time. We crossed it in our own house first so that we could state it with data rather than promises.
Do you want to go from three zeros to a verified 100/100?
The initial AuditScale GEO audit includes a quantified 0-100 technical scan, a Bar Test T+0 documented across the five engines, a verifiable competitive map and a remediation roadmap. Delivered in 7-10 days with a PDF dossier and an asynchronous video. From €1,500.
Request an audit