7 min read
How AI-ready are mid-market B2B websites? We checked 10 in New Jersey
We read the homepage and robots.txt of 10 owner-led companies in northern New Jersey. None carried FAQ structured data, five had no meta description, and one had made any decision about AI crawlers.
- Sample
- 10 owner-led B2B companies, $1M to $249M in revenue, 20 to 499 employees, northern New Jersey and the New York metro area. Not a random sample: a roster that reached us through a relationship.
- Method
- The homepage HTML and the /robots.txt of each domain, read on 27 August 2026. Nothing beyond those two public files: no login, no internal crawl, no client data. Full criteria and aggregate results are published alongside the piece.
In August 2026 we read two public files from each of ten owner-led B2B companies in northern New Jersey and the New York metro area: the homepage and the robots.txt. Every one of the ten is run by its owner or its president, none of them sells software, and their revenue runs from $1M to $249M. Not one of the ten carried FAQ structured data. Five had no meta description at all. Two declared, in terms a machine can read, what kind of business they are. One had made any decision whatsoever about AI crawlers.
That last number is the one worth sitting with. Every week a buyer asks an assistant some version of the same question: who should I hire for this. ChatGPT, Claude, Gemini, Perplexity and Google’s AI Overviews all answer it, and each one answers with a handful of names. The handful changes with every question, but the names keep coming out of the same short group. If a company is not in that group, the buyer never learns it existed. There is no page two to lose on any more.
So we went looking for how ready a normal mid-market company actually is for that. Not a tech company, not a startup: a company with trucks, a warehouse, a plant or a clinic, run by the person whose name is on the building.
What we looked at, and what we deliberately did not
Ten companies, between 20 and 499 employees, most of them decades old. All of them grew the same way: somebody told somebody else.
We read the HTML of the homepage and the contents of /robots.txt. That is the entire surface of this study. No crawling past those two files, no login, no client data, and no tools an ordinary visitor could not use. We are not naming any of the ten, and everything below is the aggregate.
We should also say plainly where the list came from, because it changes what the numbers mean. This is not a random sample of the American mid-market. It is a roster that reached us through a relationship, in one region, read on one day. It tells you what the surface looks like in one corner of one market. It does not tell you the pattern holds nationally.
What we found
| What we checked | Result |
|---|---|
| Carries FAQ structured data | 0 of 10 |
| Carries no JSON-LD at all on the homepage | 2 of 10 |
| Says nothing machine-readable about the business itself | 1 of 10 |
| Has no meta description | 5 of 10 |
| Declares what kind of business it is, in machine-readable terms | 2 of 10 |
| Names an AI crawler in robots.txt | 1 of 10 |
Four things stand out.
Nobody had FAQ schema. Zero out of ten. This is the single format that maps most directly onto how an assistant works: a question, and under it an answer that stands on its own. A model reading a page with FAQ markup gets the answer already separated from the layout. A model reading a page without it has to guess which sentence on the page is the answer, and it often guesses on a competitor’s page instead.
Two sites carry no JSON-LD, and one of them says nothing about itself at all. JSON-LD is the format search engines and assistants actually read. One of those two does declare itself an organization the old way, through inline microdata, which is worth something. The other one’s only machine-readable statement describes its navigation menu. For a company with decades of trading behind it, that is decades of reputation stored in a format a machine has to reverse engineer from marketing prose.
Half of them have no meta description. Not an empty one: the tag is absent from the HTML. In SEO terms this is old news and it usually gets treated as cosmetic. It is not cosmetic any more. That description is often the shortest, cleanest, human-written sentence about what the company does, and it is the first thing a summarizer reaches for.
Two of the ten declare what they are. Two sites state their business type in machine-readable terms. The other eight leave it to be inferred from the copy. An assistant asked for a specific kind of supplier in a specific town has to work out, from marketing prose, whether each of these companies even qualifies for the list.
The one file that decides whether any of the rest matters
Nine of the ten robots.txt files never mention an AI crawler.
That file deserves more attention than it gets, because it is the gate standing in front of everything else on this list. The AI companies now run two different kinds of crawler. One collects training data. The other fetches a page in the moment, to answer the question somebody just typed. The second kind is the one that produces a citation, and a page it cannot fetch cannot be quoted, no matter how good the page is.
The tenth file is the interesting one, and not because it gets it right. It names five AI agents and allows each of them. But the file already opens the whole site to everything with a blanket rule at the top, so those five lines change no crawler’s behavior at all. They are a statement of intent, not a configuration. And the agents it names are mostly the collectors. The ones that go fetch a page at the moment of answering are not on the list. The only company of the ten that made a decision about this made half of one.
Then there is the finding that is smaller and more uncomfortable. In one of these ten sites, both Sitemap: lines in the robots.txt point at a domain with no DNS record at all. So does the url field inside that site’s own structured data, which is the field that tells a machine which address is authoritative for this business. The real site lives at a different domain and resolves normally. Somebody did the work, carefully, and typed the wrong address into the two places a machine looks first.
Why this lands hardest on companies that grew by referral
A company that grew on referrals has usually never needed to be findable. The pipeline arrived through people. The website got built to confirm what a referral had already said, not to introduce the company to a stranger.
That worked for as long as the introduction always came from a person. It stops working the moment the introduction comes from a model, because a model cannot be told about you at a golf outing. It can only read what is on the page, and what is on the page, for most of these ten companies, is prose written for someone who already knows who they are.
The uncomfortable part is that these are good companies. Decades of delivery, real customers, real reputation. None of it is in a form an assistant can pick up and repeat.
Three checks you can run on your own site today
- Open your homepage source and search for
application/ld+json. If there is nothing, an engine is reading your site the way a stranger reads a brochure in a language they half know. - Fetch
yourdomain.com/robots.txtand read every line out loud. Check that each URL in it is spelled the way your site is actually spelled, and that the sitemap line resolves. Then check whether it says anything about the crawlers that fetch pages at answer time. - Ask ChatGPT or Perplexity the question your best customer would ask before they knew your name. Not your company name: the need. See who gets named. That list is your real competitive set now, and it is often not the one you think you are in.
What this measurement does not show
Ten companies is a small sample, all in one region, all read on one day, two files deep per company, and drawn from a roster rather than at random. It describes a surface.
It does not measure whether any of these companies actually gets cited by an assistant, which is a separate test with a separate method. It does not establish that the pattern holds outside this segment. And a site can carry perfect markup and still go uncited, because markup is a precondition, not a cause.
We are publishing it because the direction is not subtle. Zero out of ten is not a close call.
- AEO
- structured data
- New Jersey
- mid-market
So what does AI say about your business?
The free audit runs real buying queries against the five platforms and shows you who gets named in your place. It lands in your inbox within 48 business hours.
Request the free audit →