Formidable research
The State of Inbound 2026
What 15,237 B2B conversion paths actually ask a buyer. Our own crawl, our own numbers, the method and the limits written out in full - including what this data cannot tell you.
The whole report is on this page. No form, no email address.
What 15,237 B2B conversion paths actually ask a buyer
Second edition, 7 September 2026. Supersedes the 31 August edition, which measured 10,258 companies — two-thirds of the book — because a DNS fault in our own crawler silently skipped 5,678 domains. Those domains have since been crawled. Every number here is re-derived from scratch on the complete corpus.
Formidable rendered and parsed the inbound conversion path of 15,237 companies — the demo form, the contact form, the booking link, the free-trial button — and read every field label, every dropdown option, every required flag and every submit button off the page. This report is what that corpus says. It measures what companies ask. It does not measure what happens next: we never submitted a form, so there is no conversion rate, no response time and no lead-quality data anywhere in this document, and no sentence here should be read as evidence that one design outperforms another.
0. What changed since the first edition
The first edition carried a correction inside it: the crawl log appeared to say a third of the universe was dead, and it was really our own crawler failing to resolve api.firecrawl.dev on 29 August. 5,678 domains were never contacted. They have now been crawled, and this edition is written on the finished book.
| 31 Aug edition | This edition | |
|---|---|---|
| Companies with a classified conversion path | 10,258 | 15,237 |
| Share of the 17,286-domain universe | 59.3% | 88.1% |
| Domains never attempted | 5,678 | 0 |
| Apparent dead-domain rate in the raw log | 35.8% | 4.2% |
| Forms parsed | 7,534 | 11,282 |
| Field observations | 48,667 | 73,254 |
| Apollo leadership census coverage | 10,083 (98.3%) | 15,237 (100%) |
The headline distribution barely moved, which is the strongest thing that could have happened to the first edition. Primary conversion kind shifted by at most 1.3 points on any bucket; median form length, median required fields and the qualification share are unchanged. The first edition's conclusions survive a 48.5% increase in sample.
But the missing third was not a random third, and it is worth saying exactly how it differed. Measuring the 4,979 recovered companies as their own cohort against the 10,258 originally seen:
| Originally seen (10,258) | Recovered (4,979) | |
|---|---|---|
| No founder, marketing or sales lead on Apollo | 36.0% | 41.7% |
| Founder + marketing + sales all present | 20.9% | 15.8% |
Conversion form is native / unrecognised | 52.3% | 57.8% |
| Conversion form is HubSpot | 15.2% | 11.0% |
| Any scheduler fingerprinted | 26.9% | 22.3% |
| Median form fields | 6 | 6 |
| Forms requiring ≤3 fields | 45.8% | 47.2% |
The recovered companies are less vendored and less visible on Apollo — the outage skewed the first edition toward the better-resourced end of the market. It did not skew what those companies ask: form length and required-field discipline are indistinguishable between the two cohorts. So the vendor-penetration and team-shape numbers in the first edition were mildly optimistic; the form-anatomy numbers were not.
Two things also changed in the database between editions, independently of the crawl:
- ICP labelling improved. 71.8% of the classified set carried no ICP judgement in August; that is now 48.1%. The dataset still does not support ICP-scoped market claims, but it is closer.
- The census is now complete. Every classified company has an Apollo leadership row, so §7 no longer rests on a 98.3% subset.
A note on method continuity. The first edition's analysis scripts were not kept. This edition rebuilds them, and the rebuild is validated by re-running it on the original 28–29 August cohort: every purely measured figure reproduces the published number exactly — 10,258 classified, 7,534 forms, 48,667 field observations, the entire field-count distribution, 3,447 forms requiring ≤3 fields, 5,062 sites with no named vendor, 1,796 with both a form and a scheduler, 334 non-rendering HubSpot embeds. The one component that could not be recovered byte-for-byte is the semantic field classifier, which is reconstructed here from documented rules (§4) and lands within about a point of the original on every category. Where this edition's classifier disagrees with the first edition's, both cohorts are re-scored with the new one, so every comparison in this report is like-for-like.
1. Method and sampling frame
What the universe is
The universe is 17,286 domains assembled from five named acquisition lists. It is not the internet, not a random sample of B2B, and not a census of anything. Every company in it was selected because a technology detector saw a CRM or marketing-automation product on its domain, or because a VC had it in a portfolio, or because it tripped a visitor-resolution pixel.
| Source list | What it is | Domains | % of universe |
|---|---|---|---|
multi-crm-2026-07 | Technographic pull: Salesforce / Pipedrive / Zoho installs | 11,746 | 68.0% |
portco-legacy-2026-07 | VC portfolio companies | 3,470 | 20.1% |
hubspot-builtwith-2026-07 | BuiltWith list of confirmed HubSpot installs | 1,603 | 9.3% |
diginius-intent-2026-08 | Purchased intent / visitor-resolution feed | 329 | 1.9% |
v1-workbench-qualified | Hand-triaged by Formidable | 122 | 0.7% |
v1-legacy, scout, unassigned | Residue | 16 | 0.1% |
| Total | 17,286 | 100.0% |
Any claim of the form "B2B companies do X" is really "companies that had a CRM tag, a HubSpot tag, a VC on the cap table or an intent hit in July–August 2026 do X." The frame is pre-filtered for companies that already bought sales software. We state it once here and it applies to every number in the report.
One list is a different population altogether. diginius-intent-2026-08 resolves whoever tripped a pixel, and its contents are radio stations, regional news sites, corporate newsrooms and local service businesses — not B2B software. Its conversion profile is visibly a different distribution. It is reported separately wherever it matters and never pooled silently.
How each page was read
Stage one is pure code. A headless browser (Firecrawl) renders the homepage with JavaScript executed, so forms injected at runtime actually appear. A small language model reads the page's link list and picks up to three further pages worth visiting; a deterministic pass overrides one pick if the most prominent call-to-action points elsewhere. Cap: four pages per company. Those pages go through a real HTML parser — every <form>, every input, every label resolved through label[for], a wrapping label, or aria-label; every scheduler and chat widget matched against a fingerprint list. This stage cannot invent anything.
Stage two is a model choosing, never writing. The extracted facts go to alibaba/qwen3.7-flash at temperature 0, which decides which of the extracted forms is the real conversion form and which extracted fields are signal rather than a newsletter box. Its answer passes a validation gate: every label it keeps must appear verbatim in the deterministic extraction, the scheduler must be one actually fingerprinted, the conversion URL must be a page actually fetched. Field objects are copied from the extractor by construction — the model physically cannot rename, retype or fabricate a field. Of 15,237 classified companies, 15,100 (99.1%) were structured within the gate and 137 (0.9%) fell back to deterministic selection.
Measured versus inferred — the line this report holds
| Measured (code read it off a rendered DOM) | Inferred (a model or a classifier decided) |
|---|---|
Field labels, types, required flags, <select> options | Which extracted form is "the" conversion form |
| Field counts | Which extracted fields are noise |
| Form and scheduler provider fingerprints | conversionKind (form / scheduler / both / …) |
| Submit-button text, page headings, CTA text | The semantic category assigned to each label |
| Chat widget fingerprints, page URLs and titles | hasSelfSignup |
Label quality is itself measurable. Across 73,254 field observations, 56,915 (77.7%) came from a real accessible label — label[for] 49,549, wrapping label 5,248, aria-label 2,118. The other 22.3% came from a placeholder (10,738) or a raw name attribute (5,601), which are weaker evidence of what a company means to ask.
The inference is fallible in a known direction. 249 captured "conversion forms" (2.2% of 11,257) contain a password field — those are signup screens, not sales forms. That is a small error rate but a real one, and it inflates the short-form end of the distribution.
The semantic classifier, stated in full
Every measured label is assigned to exactly one category by an ordered rule list, first match wins. The ordering carries the judgement — "Company email" is an email field, not a company field; "How did you hear about us" is attribution, not a message box — so it is published rather than appealed to:
anti-bot → consent → account (password) → attribution → budget → timeline → company size → industry → job title → website → phone → email → company name → name → location → free-text message → need / use case → unclassifiable.
The full regular expressions are in scripts/analyze-inbound.ts. Two consequences worth knowing: need / use case is not counted as qualification in the strict measure (it routes a topic, it does not size an opportunity), and the residual unclassifiable bucket is 15.4% of fields, dominated by multi-select interest arrays, placeholder artefacts ("john", "doe", "example text") and non-English identity fields.
The crawl outcome
Every domain in the universe has now been attempted exactly once, so the universe and the denominator are the same number.
inbound_crawl_status | n | % of 17,286 |
|---|---|---|
ok | 15,111 | 87.4% |
no_cta — reached, no conversion path our renderer could see | 1,137 | 6.6% |
unreachable | 729 | 4.2% |
timeout at 35s | 138 | 0.8% |
partial | 126 | 0.7% |
blocked (401/403/bot wall) | 45 | 0.3% |
Classified (ok + partial) | 15,237 | 88.1% |
The real dead-domain rate in this universe is 4.2%. The 35.8% figure that appeared in the first edition's crawl table was our own DNS failure and should not be quoted from any copy of that document.
Per list, against every domain attempted:
| Source list | Attempted | Classified | Dead | No CTA | Timeout / blocked |
|---|---|---|---|---|---|
hubspot-builtwith-2026-07 | 1,603 | 1,541 (96.1%) | 37 (2.3%) | 14 (0.9%) | 11 (0.7%) |
v1-workbench-qualified | 122 | 115 (94.3%) | 1 (0.8%) | 5 (4.1%) | 1 (0.8%) |
multi-crm-2026-07 | 11,746 | 10,351 (88.1%) | 593 (5.0%) | 676 (5.8%) | 126 (1.1%) |
portco-legacy-2026-07 | 3,470 | 2,981 (85.9%) | 55 (1.6%) | 401 (11.6%) | 33 (1.0%) |
diginius-intent-2026-08 | 329 | 235 (71.4%) | 42 (12.8%) | 40 (12.2%) | 12 (3.6%) |
List decay is real but modest, and it is not what the raw first-edition log suggested: the spread on classification is 71.4%–96.1%, and the intent feed is the outlier at both ends — highest dead rate (12.8%) and highest no-CTA rate (12.2%), consistent with a list that is not B2B software.
The survivorship caveat, stated precisely
Findings here describe 15,237 domains that (1) appeared on a July–August 2026 CRM-install, technographic, VC-portfolio or intent list, (2) resolved and served a page inside 35 seconds without a bot wall, and (3) exposed at least one conversion path our renderer could see. Each filter selects toward companies that still exist, still pay for sales tooling, and still ship a conversion path a headless browser can reach.
The first edition carried a fourth filter — surviving a DNS outage — which is now gone. That is the single biggest methodological improvement in this edition, and §0 quantifies what that filter had been doing.
Two instrument biases that push in our favour, disclosed
Crawl depth predicts the answer. Mean form length rises from 4.36 fields on companies where only the homepage was read (n=155) to 6.50 where all four pages were read (n=6,230). Our field counts are a floor. Separately, the pages-crawled = 2 cell shows 59.6% scheduler presence against a 25.4% baseline — that is the crawler stopping early because it hit a booking page, not a real concentration.
Half the classified set hit the four-page cap (7,853 of 15,237, 51.5%), which means for half the set there may be a better conversion page we never opened.
What this dataset cannot support
- Anything about conversion rate, lead quality, routing or response.
postSubmitisunknownfor 100% of the 11,282 form-bearing companies. We never submitted a form. - Any trend. The crawl ran on three days inside one week. There is no time axis. No sentence in this report says "rising," "falling," or "increasingly."
- Any geographic finding. The
countryfield is 36.5% blank and its encoding is collinear with source list. Geography here is a source-list cut wearing a hat, and we publish none of it. - Any ICP-scoped claim beyond n=257 verified. 7,326 of the classified set (48.1%) still carry no ICP judgement. This dataset supports claims about how companies convert. It does not support claims about how our ICP converts.
- A single pooled headline for anything that varies by list — scheduler adoption in particular, which ranges from 3.4% to 59.6% across the five lists.
Where this report references the well-known relationship between form length and conversion rate, that is external context from published marketing research, not a measurement from this crawl, and it is labelled as such every time.
2. The headline findings
1. The form has not been replaced. 11,347 of 15,237 classified companies (74.5%) put a form on the primary path between a buyer and a conversation. Self-serve signup as the only front door is 648 companies (4.3%) — and that 4.3% is an upper bound.
2. But the forms are not qualifying anybody. Of 11,257 parsed forms, 196 (1.7%) ask about budget or timeline on a strict label match; the semantic classifier finds budget on 237 (2.1%) and timeline on 108 (1.0%). 533 (4.7%) ask company size on the strict match, 722 (6.4%) on the classifier; industry is 551 (4.9%) strict, 656 (5.8%) loose. The median form asks zero questions that would tell you whether the buyer is worth a call. The industry's story about long forms — that sales insists on qualifying — is not what the pages say.
3. The friction number everyone quotes is roughly 50% too big. The median form has 6 fields but only 4 required, and 5,212 of 11,257 forms (46.3%) require three fields or fewer. Quote the total and a prospect who counts the asterisks on their own form will conclude you cannot read a page.
4. The best-resourced marketing teams went the other way. Among the 1,007 companies running Marketo or Pardot — platforms nobody deploys without a demand-gen team — 549 (54.5%) ask eight or more fields, against 2,810 of 10,249 (27.4%) on HubSpot or hand-rolled forms. Only 29 of 1,007 (2.9%) ask three fields or fewer. This replicates in all three major source lists.
5. The scheduler brand, not the scheduler, tells you whether there is a sales team. Of 443 companies running ChiliPiper, 299 (67.5%) have a sales leader listed on Apollo. Of 2,812 running Calendly, 476 (16.9%) do — half the 33.1% base rate. Eleven Calendly companies out of 2,812 have a double-digit sales bench. Treating "has a scheduler" as one feature destroys the strongest signal in the dataset.
6. Product-led is a layer, not a replacement. 3,832 companies (25.1%) have a self-serve signup path, and 3,184 of them (83.1%) also run a form, a scheduler or a contact page. Companies that add PLG do not lighten the sales side: their demo forms are statistically the same length, still demand a phone number 61.2% of the time, and are more likely to ask company size.
7. Half the market runs no named vendor at all on the money page. 7,822 sites (51.3%) expose no recognised marketing-automation form, no recognised booker and no recognised chat widget. The modal inbound funnel in this universe is a hand-assembled form on a page, with no routing layer and no way to book.
3. The shape of inbound
The primary conversion mechanism
The conversionKind label is inferred — a model chose which mechanism is primary. The underlying artefacts (form fields, scheduler scripts, signup links) are measured.
| Conversion kind | n | % of 15,237 | (31 Aug edition) |
|---|---|---|---|
form — a form, no scheduler | 8,797 | 57.7% | 56.4% |
both — form and scheduler | 2,550 | 16.7% | 17.4% |
scheduler — calendar, no form | 1,317 | 8.6% | 9.4% |
contact_only — contact page, no form | 1,121 | 7.4% | 7.4% |
none — no mechanism found | 804 | 5.3% | 5.3% |
self_signup — product signup is the route | 648 | 4.3% | 4.1% |
| Total | 15,237 | 100.0% |
Adding half the book again moved no bucket by more than 1.3 points. That is the single best evidence available that this distribution describes the frame rather than the sample.
The shape is substantially an artefact of sourcing
This is the table that most needs saying out loud, because the one above looks like a fact about B2B until you cut it by list.
| Source list | n | form | both | scheduler | contact_only | none | self_signup |
|---|---|---|---|---|---|---|---|
multi-crm-2026-07 | 10,351 | 60.6% (6,275) | 16.3% (1,683) | 7.2% (748) | 7.0% (726) | 5.2% (538) | 3.7% (381) |
portco-legacy-2026-07 | 2,981 | 60.7% (1,810) | 8.7% (258) | 7.5% (223) | 8.9% (266) | 7.1% (212) | 7.1% (212) |
hubspot-builtwith-2026-07 | 1,541 | 32.1% (495) | 37.6% (580) | 22.0% (339) | 3.8% (59) | 1.8% (27) | 2.7% (41) |
diginius-intent-2026-08 | 235 | 58.3% (137) | 2.6% (6) | 0.9% (2) | 26.0% (61) | 10.6% (25) | 1.7% (4) |
v1-workbench-qualified | 115 | 61.7% (71) | 18.3% (21) | 4.3% (5) | 7.0% (8) | 0.9% (1) | 7.8% (9) |
Scheduler presence — any booking tool detected anywhere on the path — ranges from 3.4% to 59.6% purely by list:
| Source list | n | any scheduler | % |
|---|---|---|---|
hubspot-builtwith-2026-07 | 1,541 | 919 | 59.6% |
multi-crm-2026-07 | 10,351 | 2,431 | 23.5% |
v1-workbench-qualified | 115 | 26 | 22.6% |
portco-legacy-2026-07 | 2,981 | 482 | 16.2% |
diginius-intent-2026-08 | 235 | 8 | 3.4% |
A "% of B2B now offers instant booking" headline is a number we could set anywhere between those poles by choosing the list. We publish no pooled scheduler-adoption figure. The high end is partly a bundling artefact — HubSpot ships Meetings in the box — though only partly: within the HubSpot cohort, Calendly still outnumbers HubSpot Meetings by roughly seven to one. What the HubSpot list mostly selects for is marketing-stack maturity, and stack maturity is what determines the shape of the front door.
A note on our own labels: icp_status remains collinear with source list, and conversion shapes cut by it describe our labelling rather than the market. We do not report that cut as a finding.
How many routes in?
Three independent, non-exclusive routes derived from the evidence rather than the single primary label: a form (provider detected or ≥1 field read), a scheduler (booking script or embed), a self-serve signup.
| Routes offered | n | % of 15,237 |
|---|---|---|
| 0 | 1,914 | 12.6% |
| 1 | 8,300 | 54.5% |
| 2 | 4,311 | 28.3% |
| 3 | 712 | 4.7% |
Only 712 companies (4.7%) offer all three — down from 5.3% on the smaller sample, because the recovered tranche is less vendored. Nearly a third of them come from the 1,541-company HubSpot cohort. The modern multi-path funnel is real, but it is a small minority and it is over-represented in exactly the list you would build if you were sourcing on marketing stack.
Adding live chat as a fourth route (detected on 2,345 companies, 15.4%) moves the picture only slightly: 1,696 companies (11.1%) present none of form, scheduler, signup or chat. Chat is the only route for 218 companies (1.4%).
contact_only and none — 1,925 companies with nowhere to go
These two buckets are 12.6% of the classified set, and they are the most misread rows in the table. They are not "companies with no CTA" — that is a separate crawl outcome. These are companies the crawler mostly read three or more pages of (1,600 of 1,925, 83.1%), with CTAs harvested, and still found nothing to convert into.
Demo-intent CTAs by bucket, among companies where any CTA text was captured:
| Primary kind | n with CTAs read | demo / sales / call CTA | % |
|---|---|---|---|
scheduler | 882 | 668 | 75.7% |
none | 282 | 187 | 66.3% |
both | 1,951 | 1,175 | 60.2% |
form | 6,326 | 3,450 | 54.5% |
self_signup | 427 | 205 | 48.0% |
contact_only | 929 | 331 | 35.6% |
Of the 804 companies where we found no conversion mechanism, 282 had CTAs captured and 187 of those (66.3%) carry a "Book a Demo," "Request a Demo" or "Contact Sales" button. The button is measurably there; what it opens was not reachable within four pages. (This rate fell from 74.0% on the smaller sample — the recovered tranche's none companies advertise a demo less often, which is what you would expect if more of them are simply not selling to a B2B buyer.)
Three readings are live in the same rows and we hold all of them:
- A share are genuinely broken funnels — a company loudly promising a demo whose button leads to no form, no calendar and no signup. That is the most actionable shape in the dataset.
- A share are JavaScript-gated forms our renderer could not see. Heavily client-rendered React homepages defeat a headless first pass more often than anyone would like.
- A share were never trying to convert a B2B buyer at all. The
nonebucket is 23.1% country-code-TLD domains (186/804) and 4.1% edu/gov/org, against 9.5% and 1.1% forform. (ccTLD here excludes.io,.co,.ai,.meand.tv, which are sold generically.) Sampled business summaries include news sites, job boards, school districts and rural utilities — organisations that got into the frame because they had a CRM installed.
Treat "66.3% of none companies advertise a demo they do not deliver" as an upper bound on a real phenomenon, not a clean count.
4. Form anatomy
This is the material nobody else has: 73,254 field observations read off 11,282 rendered forms. The working set below is the 11,271 forms on companies whose primary conversion kind is form or both, less fourteen reporting 41+ fields (page-level parse failures — the worst read an entire multi-page intake wizard as one form), leaving 11,257 forms and 71,965 measured fields. Where a table below is computed on a wider set it says so.
The distribution is tighter than the folklore
Median 6 fields. Interquartile range 4–8. Mean 6.39. Percentiles: p10 = 2, p25 = 4, p50 = 6, p75 = 8, p90 = 10, p95 = 12, p99 = 20. Every one of those is identical to the first edition except the mean, which moved by 0.05.
| Fields | Forms | % | Cumulative |
|---|---|---|---|
| 1 | 890 | 7.9% | 7.9% |
| 2 | 365 | 3.2% | 11.1% |
| 3 | 663 | 5.9% | 17.0% |
| 4 | 1,273 | 11.3% | 28.3% |
| 5 | 1,700 | 15.1% | 43.4% |
| 6 | 1,622 | 14.4% | 57.9% |
| 7 | 1,385 | 12.3% | 70.2% |
| 8 | 1,060 | 9.4% | 79.6% |
| 9 | 768 | 6.8% | 86.4% |
| 10 | 476 | 4.2% | 90.6% |
| 11–12 | 498 | 4.4% | 95.1% |
| 13–15 | 279 | 2.5% | 97.5% |
| 16–20 | 189 | 1.7% | 99.2% |
| 21–40 | 89 | 0.8% | 100.0% |
The fourteen-field monster demo form is real but rare: 751 of 11,257 forms (6.7%) have twelve or more fields. The mass sits at 4–8 fields, where 7,040 forms (62.5%) live.
The 1-field bucket is the noisiest part of the distribution. 890 forms (7.9%) have exactly one field, and 528 of those (59.3%) are a single <input type="email"> — 566 (63.6%) once text inputs labelled "email" are counted too. Some are genuine progressive-capture demo requests; 27 have the submit label "Subscribe," which means our inference picked a newsletter box. Read the short end with that in mind.
Total fields is the wrong friction number
| Metric | Value |
|---|---|
| Forms with parsed fields | 11,257 |
| Median total fields | 6 |
| Median required fields | 4 |
| Mean required fields | 3.84 |
| Forms requiring ≤3 fields | 5,212 (46.3%) |
| Forms where every field is required | 2,935 (26.1%) |
| Forms with zero required fields | 2,275 (20.2%) |
Nearly half of all forms in this universe require three fields or fewer. If a report leads with "the median form asks six questions," it is quoting a number the reader can disprove on their own site in ten seconds.
Phone is the field to watch. 6,483 forms (57.6%) ask for a phone number and 3,430 of them (52.9% of askers, 30.5% of all forms) mark it required. Nearly a third of forms in this set will not accept a submission without a phone number — while the requirement rate on phone fields (51.8%) sits well below email's (79.6%), which suggests companies are themselves ambivalent about it.
The free-text box runs the other way: of 5,406 forms with a message field, only 2,071 (38.3%) require it. The one place a prospect could say something genuinely qualifying is optional in three-fifths of forms.
The label corpus: what the questions actually say
Normalised labels (lowercased, punctuation and asterisks stripped), ranked by how many distinct companies ask them. % required is measured off the required attribute.
| Rank | Label | Forms | % of 11,257 | % required |
|---|---|---|---|---|
| 1 | first name | 5,579 | 49.6% | 80% |
| 2 | last name | 5,543 | 49.2% | 79% |
| 3 | 4,397 | 39.1% | 80% | |
| 4 | phone number | 2,515 | 22.3% | 60% |
| 5 | company name | 2,462 | 21.9% | 81% |
| 6 | company | 2,286 | 20.3% | 62% |
| 7 | phone | 2,030 | 18.0% | 50% |
| 8 | message | 1,940 | 17.2% | 40% |
| 9 | name | 1,571 | 14.0% | 74% |
| 10 | work email | 1,537 | 13.7% | 86% |
| 11 | country | 1,536 | 13.6% | 76% |
| 12 | job title | 1,453 | 12.9% | 73% |
| 13 | email address | 1,110 | 9.9% | 82% |
| 14 | business email | 1,008 | 9.0% | 86% |
| 15 | full name | 651 | 5.8% | 76% |
| 16 | how did you hear about us? | 506 | 4.5% | 48% |
| 17 | subject | 435 | 3.9% | 44% |
| 18 | industry | 405 | 3.6% | 63% |
| 19 | how can we help? | 330 | 2.9% | 49% |
| 20 | your name | 329 | 2.9% | 79% |
Fourteen of the top fifteen labels are name, email, phone, company and country. The only exception, and the first thing in the corpus that is not a piece of contact information, is "message," at rank 8 — and the first thing that is a question is "How did you hear about us?", at rank 16.
These are identity forms, not qualification forms
Applying the classifier from §1 to all 71,965 measured labels. The labels are measured; the mapping from "Number of seats" to "company size" is our judgement.
| Category | Forms asking | % of 11,257 | Field occurrences | % of all fields | % required |
|---|---|---|---|---|---|
| 10,511 | 93.4% | 10,832 | 15.1% | 79.6% | |
| Name | 9,381 | 83.3% | 15,963 | 22.2% | 77.8% |
| Phone | 6,483 | 57.6% | 6,727 | 9.3% | 51.8% |
| Company name | 6,240 | 55.4% | 6,819 | 9.5% | 69.4% |
| Free-text message | 5,406 | 48.0% | 5,924 | 8.2% | 37.0% |
| (unclassifiable) | 4,260 | 37.8% | 11,106 | 15.4% | 34.4% |
| Location (country/state/city/zip) | 2,941 | 26.1% | 4,358 | 6.1% | 60.5% |
| Job title / role | 2,434 | 21.6% | 2,571 | 3.6% | 65.0% |
| Need / use case / topic | 2,213 | 19.7% | 2,536 | 3.5% | 46.7% |
| Marketing consent | 995 | 8.8% | 1,080 | 1.5% | 40.2% |
| Attribution ("how did you hear") | 917 | 8.1% | 928 | 1.3% | 44.3% |
| Website | 835 | 7.4% | 882 | 1.2% | 35.7% |
| Company size / headcount | 722 | 6.4% | 738 | 1.0% | 70.7% |
| Industry | 656 | 5.8% | 673 | 0.9% | 62.4% |
| Account creation (password) | 249 | 2.2% | 342 | 0.5% | 41.8% |
| Budget / revenue | 237 | 2.1% | 258 | 0.4% | 54.7% |
| Anti-bot / captcha | 117 | 1.0% | 117 | 0.2% | 6.8% |
| Timeline | 108 | 1.0% | 111 | 0.2% | 44.1% |
Rolled up:
| Bucket | Share of 71,965 fields |
|---|---|
| Contact identity (name, email, phone, company, title, location, website) | 66.9% |
| Open free text | 8.2% |
| Qualification (size, budget, timeline, industry, need) | 6.0% |
| Consent / attribution / captcha / password | 3.4% |
| Unclassifiable | 15.4% |
At the form level:
- 196 of 11,257 forms (1.7%) ask about budget or timeline on a strict label match.
- 533 (4.7%) ask company size or headcount; 551 (4.9%) ask industry.
- Median strict qualification fields per form: 0. Mean 0.16. Only 1,494 forms (13.3%) ask a single strict qualifier of any kind.
- 2,047 forms (18.2%) are nothing but identity fields — no message box, no dropdown, no qualifying question at all.
- A further 2,141 (19.0%) are identity fields plus a free-text box and nothing else. Together, 37.2% of every form in this corpus is contact details and possibly a comment box.
We sampled the unclassifiable bucket to check it was not hiding qualification. It is not: after the classifier's non-English rules, what remains is overwhelmingly multi-select interest arrays (interest[], services[], area-of-interest[]), placeholder artefacts and address fragments — so the 66.9% identity share is understated and the conclusion is stronger than the table shows.
The prevailing story about B2B demo forms is that they are long because sales insists on qualifying. The data does not support it. The median form in this universe collects who you are and where to reach you, and asks nothing about whether you are a fit.
The modal form is the 1998 contact form
Collapsing each form to its set of semantic categories, 11,257 forms produce 1,166 distinct signatures.
| Signature | Forms | % |
|---|---|---|
| email only | 682 | 6.1% |
| name + email + phone + company + message | 557 | 4.9% |
| name + email + phone + message | 523 | 4.6% |
| name + email + phone + company | 427 | 3.8% |
| name + email + message | 401 | 3.6% |
| name + email | 301 | 2.7% |
| name + email + phone | 300 | 2.7% |
| name + email + company + message | 285 | 2.5% |
| (nothing classifiable) | 237 | 2.1% |
| name + email + message + need | 223 | 2.0% |
| name + email + company | 215 | 1.9% |
| name + email + phone + company + location + message | 197 | 1.8% |
Ten of the top twelve shapes are permutations of the same five ingredients.
Two measured details inside "name" and "email": of the 9,381 forms asking a name, 5,740 (61.2%) split it into separate First and Last fields, spending two fields where one would do. And of the 10,511 asking an email, 3,359 (32.0%) demand a "work," "business," "corporate" or "professional" address in the label. The most common act of qualification in the entire corpus is not a dropdown — it is a gate written into an email field's label.
How forms get long, and it is not qualification
Presence of each category by form length:
| Fields | Forms | name | phone | company | message | title | location | need | industry | size | attrib. | consent | budget | timeline | Median req'd | All req'd | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–2 | 1,255 | 67% | 8% | 2% | 3% | 5% | 1% | 5% | 3% | 0% | 1% | 0% | 2% | 0% | 0% | 1 | 52% |
| 3–4 | 1,936 | 92% | 84% | 29% | 21% | 45% | 2% | 4% | 13% | 0% | 1% | 1% | 3% | 0% | 0% | 3 | 33% |
| 5–6 | 3,322 | 97% | 94% | 61% | 59% | 54% | 12% | 11% | 17% | 2% | 4% | 5% | 6% | 1% | 0% | 4 | 26% |
| 7–8 | 2,445 | 98% | 96% | 79% | 81% | 52% | 35% | 39% | 21% | 7% | 11% | 12% | 10% | 3% | 1% | 6 | 22% |
| 9–10 | 1,244 | 99% | 97% | 85% | 84% | 60% | 51% | 63% | 35% | 17% | 13% | 19% | 19% | 5% | 2% | 7 | 15% |
| 11–14 | 706 | 98% | 95% | 86% | 80% | 61% | 48% | 65% | 39% | 19% | 12% | 18% | 20% | 6% | 4% | 7 | 7% |
| 15–40 | 349 | 97% | 91% | 79% | 69% | 62% | 46% | 68% | 42% | 19% | 11% | 14% | 23% | 10% | 7% | 6 | 3% |
The accretion order is legible: email → name → free text → company → phone → job title → location → consent and industry. Company size, budget and timeline never take off. Even in the 15–40 field bucket — forms four times the median length — only 11% ask company size, 10% ask budget and 7% ask timeline.
So what is in a twelve-plus field form? Of the 4,677 unclassifiable fields on long forms, 2,583 (55.2%) are checkboxes — measured examples: interest[], services[], area-of-interest[], automationtype[] — and 3,675 of the 4,677 (78.6%) are optional. Long forms are long because of multi-select interest arrays and postal-address blocks.
Note the collapse in the right-hand columns: 52% of 1–2 field forms make everything mandatory, against 3% of 15–40 field forms. Teams that build long forms know they are long and hedge by marking most of it optional — which means those fields deliver neither data nor a clean experience.
Consistent with that, even among the longest forms:
| Total fields | Forms | Ask ≥1 strict qualifier | % |
|---|---|---|---|
| 0–3 | 1,918 | 20 | 1.0% |
| 4–6 | 4,595 | 275 | 6.0% |
| 7–10 | 3,689 | 867 | 23.5% |
| 11–15 | 777 | 250 | 32.2% |
| 16+ | 278 | 82 | 29.5% |
Among companies asking sixteen or more questions, 70.5% ask nothing that would qualify the lead. The curve stops climbing after eleven fields and then bends back down: past that point, extra length is not extra qualification.
Dropdowns: what companies try to learn when they bother
9,830 select and radio fields across 5,041 forms carried option lists. Our extractor truncates option lists at 12, and 3,468 of 9,830 fields report exactly 12 — roughly a third of dropdowns are truncated and their true option counts are unknown. Median observed options is 7; read anything at 12 as "12 or more."
| Dropdown purpose | Fields | % of 9,830 |
|---|---|---|
| Unclassifiable (mostly product/interest arrays) | 3,175 | 32.3% |
| Location (country / state / dial code) | 2,302 | 23.4% |
| Need / interest / inquiry type | 1,207 | 12.3% |
| Industry | 550 | 5.6% |
| Company size | 495 | 5.0% |
| Job title / role | 463 | 4.7% |
| Attribution | 434 | 4.4% |
| Budget / revenue | 181 | 1.8% |
There is no standard company-size taxonomy. The most common bands are 501-1000, 1-10, 11-50, 201-500, 51-200, 1000+, 51-100, 1001-5000, 5000+, 251-500 — and no single band accounts for more than 2.0% of size options. Companies inventing their own segmentation means the field is not comparable even between two companies that both ask it.
When companies build a picklist, the modal option is "none of the above." In industry dropdowns (544 forms) the most common option is "Other" at 192 occurrences — ahead of education (128), construction (104), healthcare (101), automotive (96) and manufacturing (90). Same in job-title dropdowns (420 forms): "Other" (213) leads Director (67), Manager (64) and Marketing (52). Same in attribution picklists (431 forms): Other (280), then LinkedIn (170), Social media (138), Word of mouth (85), Referral (78).
Timeline questions are so rare (108 forms) that the corpus is almost entirely bespoke.
The submit button says nothing
10,539 of 11,257 forms exposed a submit-button label.
| Submit label | Forms | % of 10,539 |
|---|---|---|
| submit | 3,919 | 37.2% |
| send | 656 | 6.2% |
| send message | 599 | 5.7% |
| request a demo | 244 | 2.3% |
| book a demo | 241 | 2.3% |
| request demo | 182 | 1.7% |
| get started | 158 | 1.5% |
| contact us | 134 | 1.3% |
| continue | 125 | 1.2% |
| get in touch | 116 | 1.1% |
The single most common thing a B2B conversion button says is "Submit" — 3,919 sites, 37.2% of every form we could read. Widen the class to include Send, Send message, Continue, Next and OK and the share of buttons that describe a database operation rather than the thing the buyer is about to receive runs between 51.0% (5,372 exact labels) and 59.0% (6,221, any label containing one of those words). Median submit-label length is one word.
The most expensive real estate on the page — the last thing a hesitating buyer reads before committing — carries the HTML default on more than a third of these sites.
The most-written bespoke question is about marketing, not the buyer
The measured long-form questions (real <label> elements, 25–95 characters, containing a "?") that appear on more than one site:
How did you hear about us? (409, plus 35 "where did you hear", 29 asterisked and 9 marked optional) · What are you interested in? (34) · What can we help you with? (30) · Anything else we should know? (11) · What do you need help with? (11) · What are you looking for? (11) · What would you like to discuss? (11) · Which best describes you? (10) · What product are you interested in? (8) · What industry are you in? (7) · Which product are you interested in? (7) · What solution are you interested in? (7)
The most common question a B2B company writes onto its own form, by a factor of twelve, is "How did you hear about us?" — an internal attribution question, asked of the prospect, at the moment of highest intent, marked required on 44% of the forms that ask it. At the form level it is more common than company size, industry, budget and timeline (917 forms against 722 / 656 / 237 / 108).
5. The vendor landscape
Market share on the conversion path
native is a residual bucket, not a vendor. The provider enum is closed at five values, so "native" means not a HubSpot, Marketo, Pardot or Typeform embed — it silently contains Webflow, Gravity Forms, Formstack, Unbounce, Salesforce Web-to-Lead, Zoho and hand-rolled React. Read it as "no recognised marketing-automation embed," never as "built it themselves."
| Form provider | n | % of 15,237 | % of form-havers |
|---|---|---|---|
| native / unrecognised | 8,245 | 54.1% | 72.6% |
| HubSpot | 2,104 | 13.8% | 18.5% |
| Marketo | 589 | 3.9% | 5.2% |
| Pardot | 418 | 2.7% | 3.7% |
| Typeform embed | 2 | 0.0% | 0.0% |
| (no form on the path) | 3,879 | 25.5% | — |
Excluding the BuiltWith list entirely, HubSpot is 1,682 of 13,696 (12.3%) — the list inflates its share by about 1.5 points, not by an order of magnitude. Typeform's n=2 is a measurement limit, not a finding: we only recognise a Typeform embed, so a Typeform reached by outbound link lands in native. Do not cite it.
| Scheduler | n | % of 15,237 | % of scheduler-havers |
|---|---|---|---|
| Calendly | 2,812 | 18.5% | 72.7% |
| ChiliPiper | 443 | 2.9% | 11.5% |
| HubSpot Meetings | 392 | 2.6% | 10.1% |
| Cal.com | 187 | 1.2% | 4.8% |
| TidyCal / YouCanBookMe / Acuity / SavvyCal | 34 | 0.2% | 0.9% |
Calendly is not the leader; it is the category. It outnumbers every other booker combined by 2.7 to 1.
| Chat widget | n | % of 15,237 |
|---|---|---|
| HubSpot Chat | 883 | 5.8% |
| Intercom | 574 | 3.8% |
| Qualified | 415 | 2.7% |
| Drift | 229 | 1.5% |
| Tawk.to | 153 | 1.0% |
| Crisp | 89 | 0.6% |
| LiveChat | 84 | 0.6% |
All prevalences are lower bounds — a widget that loads lazily, sits behind a consent gate, or uses a CDN we do not fingerprint is invisible to us.
Half the market runs no named vendor at all
| Named vendors on the path | n | % of 15,237 |
|---|---|---|
| 0 | 7,822 | 51.3% |
| 1 | 5,637 | 37.0% |
| 2 | 1,570 | 10.3% |
| 3 | 203 | 1.3% |
| 4 | 5 | 0.0% |
7,822 sites expose no recognised marketing-automation form, no recognised booker and no recognised chat widget — two points higher than the first edition, because the recovered tranche is less vendored. Some fraction genuinely runs a vendor we do not fingerprint, so treat 51.3% as an upper bound on "truly unvendored." Even discounting heavily, the modal inbound funnel here is a hand-assembled form on a page, with no routing layer and no way to book. Only 2,561 sites (16.8%) have both a form and a scheduler.
Scheduler choice is determined by form stack, and it inverts
| Form provider | Sites | Has scheduler | Calendly | ChiliPiper | HubSpot Meetings | Cal.com |
|---|---|---|---|---|---|---|
| native | 8,245 | 1,985 (24.1%) | 1,632 (82.2%) | 126 (6.3%) | 123 (6.2%) | 85 (4.3%) |
| HubSpot | 2,104 | 445 (21.2%) | 150 (33.7%) | 182 (40.9%) | 110 (24.7%) | 3 (0.7%) |
| Marketo | 589 | 60 (10.2%) | 23 (38.3%) | 37 (61.7%) | 0 | 0 |
| Pardot | 418 | 70 (16.7%) | 37 (52.9%) | 30 (42.9%) | 1 (1.4%) | 1 (1.4%) |
| (no form) | 3,879 | 1,307 (33.7%) | 969 (74.1%) | 68 (5.2%) | 158 (12.1%) | 98 (7.5%) |
Native-form shops pick Calendly over ChiliPiper 13.0 : 1. HubSpot-form shops pick ChiliPiper over Calendly 1.21 : 1; Marketo shops 1.61 : 1. That reversal is not a BuiltWith artefact — it holds with the HubSpot-sourced list excluded.
Read from the vendor's side: 56.2% of all ChiliPiper installs (249/443) sit on a marketing-automation form stack, against a 20.4% base rate — a 2.8× over-index. Cal.com is at 2.1% (4/187), a 10× under-index. ChiliPiper is not competing with Calendly for the same buyer; it is a routing layer sold to the team that already owns the form. The behaviour confirms it: 375 of 443 ChiliPiper sites (84.7%) put a form in front of the booker, against 1,843 of 2,812 (65.5%) for Calendly and 89 of 187 (47.6%) for Cal.com.
How the booker is installed says the same thing. Native shops link out to a Calendly page (32.8% of their installs are a plain href); Marketo and Pardot shops embed the booker inside a page they control (76.7% and 61.4% script embeds, 8.3% and 14.3% links) so they can pass hidden fields to it.
A technographic tag is a claim about a tag, not about a page
1,541 companies were selected because BuiltWith saw HubSpot on the domain. On the actual conversion path:
| HubSpot signal in the inbound path | n | % of 1,541 |
|---|---|---|
| HubSpot is the conversion form provider | 422 | 27.4% |
| Any HubSpot artefact (form, Meetings, Chat, or a detected-but-unrendered embed) | 732 | 47.5% |
| No HubSpot fingerprint on the conversion path at all | 809 | 52.5% |
Conversion form is native | 617 | 40.0% |
More than half of a list of confirmed HubSpot customers has no HubSpot on the page where the money enters. The tag is on the site; the money page is a native form, a Marketo instance or a Calendly link. Buying a list on "installs X" and pitching against X is a coin flip. (Caveat both ways: BuiltWith records may be stale, the tag may live on a subdomain we did not crawl, and our own embed sentinel proves we sometimes see a HubSpot form we cannot read — so 47.5% is a floor on HubSpot presence.)
HubSpot does not own its own customers' funnels either. Of 2,104 sites running a HubSpot form, only 110 (5.2%) also run HubSpot Meetings and 36 (1.7%) run form + Meetings + Chat. Among HubSpot-form shops that book meetings at all, 335 of 445 (75.3%) book on somebody else's tool. Running it the other way, 158 of the 392 HubSpot Meetings installs (40.3%) sit on sites with no form at all — it is being adopted as a free Calendly substitute more than as part of a HubSpot funnel.
Your form platform predicts your form length
All counts are visible fields on the rendered primary conversion form, 1–40 fields.
| Provider | Forms | Mean | Median | p25 | p75 | p90 | ≤3 fields | ≥8 fields |
|---|---|---|---|---|---|---|---|---|
| native | 8,221 | 6.11 | 5 | 4 | 8 | 10 | 1,641 (20.0%) | 2,086 (25.4%) |
| HubSpot | 2,028 | 6.73 | 7 | 5 | 8 | 10 | 248 (12.2%) | 724 (35.7%) |
| Pardot | 418 | 7.86 | 7 | 6 | 9 | 11 | 16 (3.8%) | 207 (49.5%) |
| Marketo | 589 | 8.11 | 8 | 7 | 9 | 11 | 13 (2.2%) | 342 (58.1%) |
The ladder is monotonic and it replicates in all three major lists.
What the extra fields are matters more than how many there are:
| Provider | Forms | Asks phone | Has a <select> | Has free text | Asks location | Asks company size |
|---|---|---|---|---|---|---|
| native | 8,221 | 55.4% | 39.8% | 50.9% | 20.6% | 4.7% |
| HubSpot | 2,028 | 58.3% | 48.9% | 36.8% | 31.7% | 11.3% |
| Pardot | 418 | 74.2% | 71.1% | 51.9% | 46.7% | 7.9% |
| Marketo | 589 | 74.2% | 85.7% | 44.0% | 69.8% | 13.1% |
Native forms are conversational; vendor forms are structured. A native form is the least likely to include a dropdown. Marketo inverts that — 85.7% carry a <select> and 69.8% ask for a location. The extra Marketo fields are not a longer conversation; they are picklists feeding routing, territory assignment and lead scoring.
Label vocabulary confirms it: the top twenty labels account for 69.9% of Marketo fields and only 47.9% of native fields, which use 12,024 distinct labels across 50,254 fields including a long non-English tail. Marketing-automation forms converge on a shared machine-readable vocabulary; native forms are idiosyncratic.
One instrument artefact, do not cite: only 13.2% of Pardot form fields carry a required attribute in the DOM, against 55.6% native, 78.4% HubSpot and 86.2% Marketo. Pardot validates server-side and does not mark the markup. Pardot's "asks phone" figure is sound; its "requires phone" figure is a measurement floor.
Chat widgets are three different products wearing the same UI
| Widget | n | Also has scheduler | Also has self-signup | Mean form fields | On a MAP form stack |
|---|---|---|---|---|---|
| Qualified | 415 | 26 (6.3%) | 91 (21.9%) | 7.23 | 237 (57.1%) |
| HubSpot Chat | 883 | 318 (36.0%) | 250 (28.3%) | 6.13 | 419 (47.5%) |
| LiveChat | 84 | 24 (28.6%) | 24 (28.6%) | 7.73 | 9 (10.7%) |
| Drift | 229 | 27 (11.8%) | 34 (14.8%) | 6.59 | 40 (17.5%) |
| Tawk.to | 153 | 65 (42.5%) | 44 (28.8%) | 5.99 | 3 (2.0%) |
| Intercom | 574 | 241 (42.0%) | 295 (51.4%) | 5.44 | 97 (16.9%) |
| Crisp | 89 | 39 (43.8%) | 44 (49.4%) | 5.00 | 5 (5.6%) |
| no widget (baseline) | 12,892 | 3,163 (24.5%) | 3,080 (23.9%) | 6.41 | 2,319 (18.0%) |
Three clusters:
- Qualified is the enterprise marketing-ops widget. 57.1% of Qualified sites run a recognised MAP form, nearly 3× the 20.4% base rate. Qualified sites also carry the longest forms of any widget cohort with meaningful n (7.23 mean). The AI chat agent is layered on top of the long form, not instead of it.
- Intercom and Crisp are the product-led widgets. 51.4% and 49.4% have self-signup against a 25.1% base — roughly 2× — and they carry the shortest forms (5.44 and 5.00).
- HubSpot Chat is a bundling artefact. 419 of its 883 installs are on a MAP stack, nearly all of it HubSpot's own.
Qualified's 6.3% scheduler rate is not evidence that its customers cannot book. We cannot see booking that happens inside a chat widget, and routing a live visitor to a rep in-conversation never produces a Calendly or ChiliPiper fingerprint in the markup. The honest statement is: on Qualified sites the visible page offers a long form and nothing else. Drift's 11.8% is directionally the same story.
A measured defect: 460 embeds that never painted
The extractor records hubspotFormDetectedNotRendered when it sees a HubSpot embed script but finds no rendered <form>. 460 sites (3.0%) hit that condition — 384 where another form rendered alongside it (two forms, two destinations, one page) and 76 where there was nothing else, of which 71 have no scheduler fallback either.
What is measured: on 460 domains, an automated visit to the company's own stated demo page produced a HubSpot script tag and no form. What is inferred: that any human ever saw the same thing. Script timing, consent gating, geo-blocking and bot detection all plausibly explain it. The distribution across lists (242 multi-CRM, 104 portco, 101 HubSpot-BuiltWith) is consistent with a real phenomenon rather than one list's quirk — note that the BuiltWith list, 9% of the universe, contributes 22% of the defect.
6. Product-led versus sales-led
Self-signup is common; self-signup instead of a sales path is rare
| Segment | n | % of 15,237 | % of self-signup cohort |
|---|---|---|---|
| Self-signup coexisting with a form, scheduler or contact page | 3,184 | 20.9% | 83.1% |
| Self-signup only, no human path found | 648 | 4.3% | 16.9% |
| Sales-led only, no self-signup detected | 10,601 | 69.6% | — |
| No conversion mechanism found at all | 804 | 5.3% | — |
13,785 of 15,237 (90.5%) offer at least one route to a human. In this universe the sales-led motion is the near-universal substrate and product-led is a layer added on top of it.
Self-signup is also badly under-counted by the primary label. Read as a secondary route it appears on 24.0% of form companies, 27.7% of both and 27.0% of scheduler — six times the 4.3% you get from the primary column. Roughly a quarter of every form-fronted company in this set is simultaneously running a product-led path; the form is competing with a "start free" button on the same page.
Evidence quality: of the 3,832, 2,191 (57.2%) had a /signup-style URL captured and 559 (14.6%) a free-trial URL — 71.8% anchored to an unambiguous pattern. The remainder rest on the model's read of a less obvious URL and are softer.
The 4.3% "pure PLG" figure is an upper bound, and the error runs one way
Of the 648 companies classified self_signup, 427 had CTAs captured, and 205 of those (48.0%) advertise a demo, a sales conversation or a call on their own homepage. What happened is that the crawler followed the conversion URL to a /sign-up page, found no form fields to parse, and the model labelled the whole company self-signup.
The measurement error converts sales-led companies into apparent PLG ones, never the reverse. Any "sales-led is disappearing" claim built from this distribution is inflated by it. (The mirror risk exists — the crawler can miss signup routes too — but a /signup link in a nav bar is a far easier target than a JavaScript-rendered multi-step form, so the errors are not symmetric in magnitude.)
Adding PLG does not lighten the sales side
Holding conversion kind constant and excluding forms under three fields:
| Conversion kind | Self-signup? | n | Mean fields | Median | Mean required |
|---|---|---|---|---|---|
| Form only | No | 6,058 | 7.20 | 7 | 4.35 |
| Form only | Yes | 1,786 | 7.16 | 6 | 4.25 |
| Form + scheduler | No | 1,617 | 6.59 | 6 | 3.92 |
| Form + scheduler | Yes | 541 | 6.06 | 5 | 3.80 |
For form-only companies the difference is 0.04 fields — nothing. The content is likewise near-identical:
| Field asked (forms with 3+ fields) | Sales-led (n=7,675) | Self-signup present (n=2,327) |
|---|---|---|
| Phone number | 65.5% (5,029) | 61.2% (1,425) |
| Company / organisation | 61.6% (4,725) | 64.3% (1,472) |
| Free-text message | 55.9% (4,292) | 45.2% (1,051) |
| Location | 28.6% (2,192) | 29.4% (683) |
| Job title | 24.3% (1,864) | 24.1% (561) |
| Company size / employees | 6.3% (485) | 9.8% (227) |
Two things stand out. PLG companies are more likely to ask company size — the one field that exists purely to route a self-serve user toward or away from sales. And 61.2% still demand a phone number on a site where the visitor could have created an account thirty seconds earlier without one.
And they use schedulers more, not less
| No self-signup (n=11,405) | Self-signup (n=3,832) | |
|---|---|---|
| Has a scheduler | 2,806 (24.6%) | 1,062 (27.7%) |
| Has a form | 8,534 (74.8%) | 2,824 (73.7%) |
Form ownership is flat while scheduler ownership is 3.1 points higher. The two-funnel company does not remove the sales path; it adds the lower-friction sales path alongside the form it already had.
Two-thirds of scheduler adopters keep the form
Of the 3,867 companies whose primary conversion kind involves a booking tool, 2,550 (65.9%) also run a form and 1,317 (34.1%) are booking-only. The modal behaviour when a company adds a scheduler is addition, not substitution.
The one signal that does point toward shorter
Companies offering a scheduler alongside a form run a shorter form, and this holds in every provider and every list:
| Provider | Form only (n, mean) | Form + scheduler (n, mean) | Gap |
|---|---|---|---|
| native | 6,247 — 6.32 | 1,974 — 5.45 | −0.87 |
| HubSpot | 1,588 — 6.88 | 440 — 6.17 | −0.71 |
| Marketo | 529 — 8.19 | 60 — 7.42 | −0.77 |
| Pardot | 348 — 7.95 | 70 — 7.41 | −0.54 |
| Source list | Form only (n, mean) | Form + scheduler (n, mean) |
|---|---|---|
multi-crm-2026-07 | 6,218 — 6.93 | 1,679 — 6.00 |
portco-legacy-2026-07 | 1,789 — 5.59 | 258 — 4.98 |
hubspot-builtwith-2026-07 | 491 — 6.14 | 578 — 5.08 |
Four of four providers and three of three lists move the same direction. This is correlational and confounded. Apollo finds a marketing leader far more often at form companies than at scheduler-only ones — smaller companies run shorter forms and are likelier to bolt on a calendar. It is not evidence that shortening a form caused anything, and we do not present it as such.
7. GTM team shape versus stack
Joining apollo_leadership_census, which now covers 15,237 domains — 100% of the classified set (it was 98.3% in the first edition). It is still strictly downstream of a successful crawl and inherits every bias above.
Apollo absence is not organisational absence. has_sales = 0 means Apollo lists no sales leader, and Apollo coverage rises with company size and US-market profile. Read the ordering as the finding, never the absolute levels. Where it matters, figures are repeated under a coverage control — restricted to companies where Apollo lists at least one founder (n=8,218), which at least proves Apollo has the company indexed.
The second survivorship layer: a readable funnel is not a reachable human
| Leadership personas on Apollo | n | % of 15,237 | (31 Aug edition) |
|---|---|---|---|
| None of the three | 5,773 | 37.9% | 36.1% |
| Founder only | 3,232 | 21.2% | 21.6% |
| Founder + Marketing + Sales | 2,931 | 19.2% | 20.9% |
| Founder + Sales | 1,149 | 7.5% | 7.3% |
| Founder + Marketing | 906 | 5.9% | 6.0% |
| Marketing + Sales | 549 | 3.6% | 4.0% |
| Sales only | 416 | 2.7% | 2.4% |
| Marketing only | 281 | 1.8% | 1.8% |
37.9% of companies whose conversion form we can read in full have no founder, no marketing lead and no sales lead listed on Apollo. Reading the funnel and reaching the owner of the funnel are two different survival filters, and only 62.1% clear both — slightly worse than the first edition reported, because the recovered tranche is less well covered by Apollo.
That gap tracks conversion sophistication, which is itself a warning: any list built as crawled ∧ enriched is doubly skewed toward the well-resourced end — the companies least likely to need a diagnosis.
Stack predicts team, at 2–3× the base rate
| Signal | n | P(marketing leader) | P(sales leader) | P(nobody on Apollo) |
|---|---|---|---|---|
| Base rate — all | 15,237 | 30.6% | 33.1% | 37.9% |
| Marketo form | 589 | 72.7% | 75.9% | 16.1% |
| ChiliPiper scheduler | 443 | 61.6% | 67.5% | 16.5% |
| HubSpot form | 2,104 | 55.5% | 62.6% | 15.2% |
| Pardot form | 418 | 54.3% | 62.4% | 23.4% |
| Form ≥8 fields | 3,361 | 42.0% | 49.0% | 31.7% |
| Self-signup present | 3,832 | 32.4% | 33.9% | 32.9% |
| Form 1–3 fields | 1,925 | 27.2% | 27.9% | 38.5% |
| HubSpot Meetings | 392 | 25.3% | 33.9% | 19.6% |
| Native form | 8,245 | 22.9% | 25.3% | 42.9% |
| Calendly | 2,812 | 15.9% | 16.9% | 46.0% |
| cal.com | 187 | 11.2% | 11.8% | 35.3% |
Under the coverage control (Apollo lists ≥1 founder, n=8,218) the spread widens rather than collapsing: Marketo 88.1% marketing / 90.4% sales, ChiliPiper 73.2% / 80.4%, Calendly 25.3% / 27.9%, against a controlled base of 46.7% / 49.6%.
And it is not the BuiltWith list leaking. P(marketing leader) computed inside each source:
| Source list | n | base | Marketo | HubSpot form | Native form | Calendly |
|---|---|---|---|---|---|---|
multi-crm-2026-07 | 10,351 | 30.3% | 72.3% (n=498) | 57.5% (n=1,213) | 22.3% (n=5,872) | 15.1% (n=1,904) |
portco-legacy-2026-07 | 2,981 | 30.2% | 80.4% (n=56) | 48.1% (n=420) | 23.9% (n=1,571) | 19.9% (n=221) |
hubspot-builtwith-2026-07 | 1,541 | 32.8% | 64.3% (n=28) | 56.9% (n=422) | 22.5% (n=617) | 16.9% (n=674) |
Marketo lands at 64–80% against a 30–33% base in all three, including the VC-portfolio slice where companies are younger — though the Marketo cells outside the CRM pull are thin (n=56 and n=28) and the n=28 cell should not be quoted on its own. Calendly sits below base in all three.
Two schedulers, opposite signals
This table counts sales leaders Apollo lists with an email, a stricter measure than the persona-present flag used above.
| Scheduler | n | 0 sales leaders (with email) | ≥10 sales leaders |
|---|---|---|---|
| (none) | 11,369 | 7,930 (69.8%) | 556 (4.9%) |
| ChiliPiper | 443 | 193 (43.6%) | 24 (5.4%) |
| Calendly | 2,812 | 2,428 (86.3%) | 11 (0.4%) |
| HubSpot Meetings | 392 | 290 (74.0%) | 0 (0.0%) |
| cal.com | 187 | 174 (93.0%) | 0 (0.0%) |
Eleven Calendly companies out of 2,812 have a double-digit sales bench. ChiliPiper implies a sales team to route to; Calendly implies there is nobody to route to and the founder takes the meeting. This is mechanically sensible — ChiliPiper is lead routing, only worth buying when there are multiple reps — and it means the scheduler brand on a page is a cheap, purely measured proxy for whether there is a sales team behind the booking.
Consistent with that, scheduler-primary companies are the least likely to have an identifiable sales leader (16.4%, n=1,317) of any conversion kind. The calendar there is a substitute for a sales team, not a tool of one.
What the form asks predicts team shape better than how long it is
| Form asks… | n | P(marketing) | P(sales) | P(nobody) |
|---|---|---|---|---|
| Base — forms parsed | 11,268 | 32.6% | 36.1% | 35.7% |
| Attribution ("how did you hear") | 917 | 48.9% | 53.3% | 19.4% |
| Job role / title | 2,436 | 48.8% | 54.0% | 24.9% |
| Geo (country / region) | 2,942 | 47.7% | 53.9% | 29.8% |
| Company size / employees | 724 | 46.0% | 51.1% | 25.4% |
| Phone | 6,486 | 35.5% | 40.5% | 34.8% |
| None of the above | 3,395 | 22.8% | 23.5% | 42.5% |
Counting how many of those five signal classes a form carries:
| Classes asked | n | P(marketing) | P(sales) | P(nobody) |
|---|---|---|---|---|
| 0 | 3,395 | 22.8% | 23.5% | 42.5% |
| 1 | 3,988 | 27.7% | 31.1% | 37.2% |
| 2 | 2,412 | 40.8% | 46.5% | 30.4% |
| 3 | 1,213 | 54.7% | 61.0% | 24.8% |
| 4+ | 260 | 58.1% | 65.0% | 23.5% |
Monotonic all the way. "How did you hear about us" is the cleanest pure-marketing artefact in the dataset — it exists only if someone is accountable for attribution, and it lifts P(marketing leader) by 16.3 points over the base.
Raw field count reverses at the top end (8–11 fields: 43.8% marketing, n=2,610; 12+: 35.4%, n=751) because the 12+ bucket is 75.5% native forms — job applications, registration flows, or an inference failure. That reversal is an artefact, not a finding.
Team shape and stack are correlated, not locked — and the mismatches are the interesting cases
| Leadership shape | n | Light stack (Calendly/cal.com + native or no form) | No form and no scheduler detected |
|---|---|---|---|
| all three | 2,931 | 144 (4.9%) | 433 (14.8%) |
| marketing + sales | 549 | 34 (6.2%) | 85 (15.5%) |
| founder + marketing | 906 | 137 (15.1%) | 193 (21.3%) |
| founder + sales | 1,149 | 175 (15.2%) | 141 (12.3%) |
| founder only | 3,232 | 882 (27.3%) | 454 (14.0%) |
| none | 5,773 | 1,316 (22.8%) | 1,155 (20.0%) |
433 companies that Apollo says have a founder, a marketing leader and a sales leader have no detectable form and no detectable scheduler on their conversion path. Another 144 route a full GTM org through a personal-calendar link and a hand-rolled form. That is 577 companies — 19.7% of the fully-staffed cohort — whose org chart has outgrown their inbound plumbing.
Caveat honestly: part of that 19.7% is our renderer failing to see a JavaScript-gated form, not a company failing to build one. The two are indistinguishable from the outside, which is itself the point — if our headless browser cannot find the form, some share of real buyers on slow connections and locked-down browsers cannot either.
8. What this means
For a company reading this about its own funnel:
-
Count your required fields, not your fields. The median form here shows six and demands four. 46.3% demand three or fewer. If you are benchmarking your form against an industry number, make sure you are comparing the same thing.
-
If your form is long, it is probably long for the wrong reason. In this corpus, long forms are long because of address blocks and multi-select interest checkboxes — 78.6% of which are optional — not because anyone is qualifying. If you are asking sixteen questions, there is a 70.5% chance none of them tells you whether the buyer is worth a call.
-
You are probably not qualifying at all. 1.7% of these forms ask budget or timeline; 4.7% ask company size. Only 13.3% ask a single strict qualifier of any kind. The most common bespoke question anyone writes is "How did you hear about us?" — a question that serves your attribution reporting, not the buyer's next step, asked at the moment of highest intent and marked required by 44% of the companies that ask it.
-
Read your own submit button. 3,919 of 10,539 say "Submit." The last thing a hesitating buyer reads before committing describes a database operation.
-
Check whether your CTA actually goes anywhere. 187 of the 282
none-classified companies whose buttons we could read advertise a demo that leads to no form, no calendar and no signup within four pages. A further 460 have a form embed that did not render for an automated visit, 76 of them with nothing else on the page. Some of these are our renderer's failure; some are real. The only way to know which is to test your own page in a cold browser. -
If you are buying lists on installed technology, discount them. 52.5% of a list of confirmed HubSpot customers had no HubSpot anywhere on the path a buyer walks. The tag describes what a company bought, not what a prospect touches.
-
The scheduler brand on a page tells you more than the form does. ChiliPiper means there is a bench to route to (67.5% have a sales leader). Calendly usually means there is not (16.9%). If you are building a target list or a routing rule, do not collapse them into "has a scheduler."
What this report deliberately does not say: that shorter forms convert better. We measured what forms ask, never what fraction of visitors completed them. The best-resourced demand-gen teams in this sample — the 1,007 running Marketo or Pardot, who have the testing budget to know — went the other way: 54.5% of them ask eight or more fields and 2.9% ask three or fewer. A "shorter is better" pitch would send a Marketo customer to check their peer set and find the opposite. The defensible statement is different and stronger: your form collects six contact fields, requires four of them, and asks nothing that would tell you whether this buyer is worth a call — and neither does most of your peer set.
9. Limitations, and what we would need next
Structural limits of this dataset:
| Limitation | Consequence |
|---|---|
| Three-day snapshot (28, 29, 31 Aug 2026) | No trend of any kind is computable. Forms behind A/B tests, geo-gates or logged-in states were seen in one variant only. |
postSubmit unknown for 100% of form companies | Nothing about routing, response time, instant-book or lead handling. |
| Sampling frame is five named lists | No "% of B2B" claim. Provider mix is partly manufactured by list selection. |
| Four-page crawl cap, hit by 51.5% of the set | For half the set there may be a better conversion page we never opened. |
| Field counts rise with crawl depth (4.36 → 6.50) | Our field counts are a floor. |
| Which form is "the" conversion form is inferred | 249 captured forms contain a password field — known-direction error inflating the short end. |
| Dropdown options truncated at 12 | 3,468 of 9,830 option lists hit the cap; option counts are not complete. |
| Semantic classifier is a reconstruction | Category boundaries are ours, published in §1; measured counts are not affected. |
country 36.5% blank and collinear with source list | No geographic findings published. |
| 48.1% of the classified set has no ICP judgement | ICP-scoped claims are limited to n=257 verified. |
| Apollo census covers only crawl survivors | Every leadership finding is conditional on the site being live and crawlable; Apollo absence ≠ organisational absence. |
| Universe contains non-B2B organisations | The none bucket is 23.1% country-code TLD and includes school districts, news sites and utilities. |
| Long-tail cell counts include duplicate operators | Some operators contribute a dozen or more near-identical domains with near-identical forms. |
Resolved since the first edition: the DNS outage that dropped 5,678 domains (now crawled), the 5,678-row survivorship filter, the 98.3% census coverage, and the 71.8% ICP-unlabelled share (now 48.1%).
What we would need to answer the next question. The question this dataset raises and cannot answer is whether any of this costs anybody anything. To get there we would need:
- A second crawl. One more pass at a meaningful interval turns a cross-section into a panel and makes every "is this changing" question answerable. Nothing else in this list matters as much — and the corpus is now complete enough that a re-crawl would be a true like-for-like panel.
- Submission telemetry. Even a small consented panel of companies willing to share form-completion data would convert every anatomy finding into a performance finding. Absent that, we are describing questions, not outcomes.
- A rendering fallback for JavaScript-gated forms. The 460 non-rendering embeds, the 804
nonerows and part of the 1,137no_ctarows are one problem. A second-pass renderer with a longer wait and click-through on modal CTAs would tell us how much of the "broken funnel" segment is real. - A size or revenue field. Almost every stack finding here is confounded with company size. Firmographic size would separate "Marketo predicts a marketing team" from "Marketo predicts money."
- A non-technographic control list. Every finding is scoped to companies that already bought sales software. A list assembled without that filter — even a small one — would tell us how much of the shape is the market and how much is the frame.
- ICP labels on the remaining 7,326. Until they exist, this dataset describes how companies convert, not how our buyers convert.
Appendix: raw distributions
A1. Universe and crawl outcome
| Metric | n |
|---|---|
| Domains in universe | 17,286 |
| Attempted | 17,286 |
Classified (ok + partial) | 15,237 |
| Unreachable | 729 |
| Reached, no CTA found | 1,137 |
| Timeout | 138 |
| Blocked | 45 |
A2. Per-list outcome
| Source list | Attempted | Classified | % | Dead | % | No CTA | % |
|---|---|---|---|---|---|---|---|
hubspot-builtwith-2026-07 | 1,603 | 1,541 | 96.1% | 37 | 2.3% | 14 | 0.9% |
v1-workbench-qualified | 122 | 115 | 94.3% | 1 | 0.8% | 5 | 4.1% |
multi-crm-2026-07 | 11,746 | 10,351 | 88.1% | 593 | 5.0% | 676 | 5.8% |
portco-legacy-2026-07 | 3,470 | 2,981 | 85.9% | 55 | 1.6% | 401 | 11.6% |
diginius-intent-2026-08 | 329 | 235 | 71.4% | 42 | 12.8% | 40 | 12.2% |
A3. Conversion kind (classified set)
| Kind | n | % |
|---|---|---|
| form | 8,797 | 57.7% |
| both | 2,550 | 16.7% |
| scheduler | 1,317 | 8.6% |
| contact_only | 1,121 | 7.4% |
| none | 804 | 5.3% |
| self_signup | 648 | 4.3% |
A4. Route combinations (form / scheduler / self-signup)
| Form | Scheduler | Self-signup | n | % |
|---|---|---|---|---|
| ✓ | — | — | 6,685 | 43.9% |
| ✓ | — | ✓ | 2,112 | 13.9% |
| — | — | — | 1,914 | 12.6% |
| ✓ | ✓ | — | 1,849 | 12.1% |
| — | ✓ | — | 957 | 6.3% |
| ✓ | ✓ | ✓ | 712 | 4.7% |
| — | — | ✓ | 658 | 4.3% |
| — | ✓ | ✓ | 350 | 2.3% |
A5. Form field count (11,257 forms, 1–40 fields)
| Fields | n | % | Fields | n | % |
|---|---|---|---|---|---|
| 1 | 890 | 7.9% | 11 | 304 | 2.7% |
| 2 | 365 | 3.2% | 12 | 194 | 1.7% |
| 3 | 663 | 5.9% | 13 | 122 | 1.1% |
| 4 | 1,273 | 11.3% | 14 | 86 | 0.8% |
| 5 | 1,700 | 15.1% | 15 | 71 | 0.6% |
| 6 | 1,622 | 14.4% | 16 | 56 | 0.5% |
| 7 | 1,385 | 12.3% | 17 | 29 | 0.3% |
| 8 | 1,060 | 9.4% | 18 | 38 | 0.3% |
| 9 | 768 | 6.8% | 19 | 42 | 0.4% |
| 10 | 476 | 4.2% | 20 | 24 | 0.2% |
| 21–40 | 89 | 0.8% |
A6. Required field count
| Required | n | % |
|---|---|---|
| 0 | 2,275 | 20.2% |
| 1 | 890 | 7.9% |
| 2 | 801 | 7.1% |
| 3 | 1,246 | 11.1% |
| 4 | 1,438 | 12.8% |
| 5 | 1,366 | 12.1% |
| 6 | 1,143 | 10.2% |
| 7 | 843 | 7.5% |
| 8 | 565 | 5.0% |
| 9+ | 690 | 6.1% |
A7. Field type (all 73,254 measured fields)
| Type | % |
|---|---|
| text | 48.2% |
| select | 12.8% |
| 12.6% | |
| textarea | 8.9% |
| checkbox | 8.1% |
| tel | 6.4% |
| radio | 1.1% |
| number | 0.7% |
| password | 0.5% |
A8. Label provenance (73,254 observations)
| Source | n | % |
|---|---|---|
label[for] | 49,549 | 67.6% |
placeholder | 10,738 | 14.7% |
name attribute | 5,601 | 7.6% |
wrapping <label> | 5,248 | 7.2% |
aria-label | 2,118 | 2.9% |
| Grounded in an accessible label | 56,915 | 77.7% |
A9. Providers
| Form provider | n | % of 15,237 |
|---|---|---|
| native / unrecognised | 8,245 | 54.1% |
| (no form) | 3,879 | 25.5% |
| HubSpot | 2,104 | 13.8% |
| Marketo | 589 | 3.9% |
| Pardot | 418 | 2.7% |
| Typeform embed | 2 | 0.0% |
| Scheduler | n | Chat widget | n |
|---|---|---|---|
| (none) | 11,369 | (none) | 12,892 |
| Calendly | 2,812 | HubSpot Chat | 883 |
| ChiliPiper | 443 | Intercom | 574 |
| HubSpot Meetings | 392 | Qualified | 415 |
| Cal.com | 187 | Drift | 229 |
| TidyCal | 15 | Tawk.to | 153 |
| YouCanBookMe | 10 | Crisp | 89 |
| Acuity | 5 | LiveChat | 84 |
| SavvyCal | 4 |
A10. Pages crawled
| Pages | n | % |
|---|---|---|
| 4 (cap) | 7,853 | 51.5% |
| 3 | 6,188 | 40.6% |
| 2 | 804 | 5.3% |
| 1 | 392 | 2.6% |
A11. Apollo leadership shape (15,237 joined)
| Shape | n | % |
|---|---|---|
| none of the three | 5,773 | 37.9% |
| founder only | 3,232 | 21.2% |
| founder + marketing + sales | 2,931 | 19.2% |
| founder + sales | 1,149 | 7.5% |
| founder + marketing | 906 | 5.9% |
| marketing + sales | 549 | 3.6% |
| sales only | 416 | 2.7% |
| marketing only | 281 | 1.8% |
A12. ICP labelling (classified set)
icp_status | n | % |
|---|---|---|
| predicted | 7,635 | 50.1% |
| unreviewed | 7,326 | 48.1% |
| verified | 257 | 1.7% |
| legacy / rejected / manual | 19 | 0.1% |
A13. Cohort comparison — the replication evidence
Both columns are computed by the same code in this edition. The "originally seen" column reproduces the 31 August edition's published figures exactly on every purely measured quantity.
| Measure | Originally seen (10,258) | Recovered (4,979) | Full book (15,237) |
|---|---|---|---|
form primary | 56.4% | 60.5% | 57.7% |
both primary | 17.4% | 15.3% | 16.7% |
scheduler primary | 9.4% | 7.0% | 8.6% |
self_signup primary | 4.1% | 4.6% | 4.3% |
| Any scheduler | 26.9% | 22.3% | 25.4% |
| Native form | 52.3% | 57.8% | 54.1% |
| HubSpot form | 15.2% | 11.0% | 13.8% |
| Median form fields | 6 | 6 | 6 |
| Mean form fields | 6.34 | 6.49 | 6.39 |
| Forms requiring ≤3 | 45.8% | 47.2% | 46.3% |
| No leadership on Apollo | 36.0% | 41.7% | 37.9% |
| Founder + marketing + sales | 20.9% | 15.8% | 19.2% |
Source: data/leads.db, tables v2_companies (17,286 rows) and apollo_leadership_census (15,237 rows). Crawl 28–29 and 31 August 2026. Every figure is re-derived by scripts/analyze-inbound.ts, which is read-only and takes a --cohort flag so any number here can be recomputed on the original, recovered or full book. Percentages rounded to one decimal place; exact n given throughout.
This report measures what forms ask. Formidable is what happens when you stop asking and start having the conversation instead. Try the live demo.
Crawl run 28-29 and 31 August 2026. Second edition, 7 September 2026.