LCHWM LLC · US business lead research  |  ← back to the universe map
internal draft · v1 · source map

Where to find US business & business-person leads

Every source class ranked by new contactable entities per hour of work — not by row count. Status is against what this corpus already holds: 186,461,039 rows / 8.1 GB of Parquet across 1,356 shards.

The honest funnel

186.5 M
rows in corpus
1,356 sources · 8.1 GB Parquet
unknown
distinct entities
never measured — no normalization pass yet
unknown
with a contact channel
the number that actually sells
~34–36 M
real US businesses
incl. nonemployers; 330 M ≈ 97% of the population

Rows ≠ entities ≠ leads. A business appears in a registry, a licence file, a tax permit and three permit sets — six rows, one lead. Everything below is judged on new entities and contact density, never on volume.

What we already hold — volume by source family

unclassified (wave 6) 88,900,158
occupational licensing 46,895,732
arcgis / socrata sweep 30,520,090
federal 15,580,374
permits 2,435,553
health practitioners 1,206,782
derived (dentists) 922,350
QUEUED — not yet pulled ~23,280,927

Nearly half the corpus is still unclassified by family — that bucket is wave-6 files whose slugs don't say what they are, and classifying it is a prerequisite for knowing what we can already sell. Of the 779 queued sources / ~23.3 M rows, 765 are ArcGIS and the biggest are permits — buying signals attached to businesses we already hold, not new entities.

The source map

Source classAccessEntitiesContact density Volume availableStatusWhy / next move
Occupational & professional licensing
state boards
Socrata / bulkboth 46.9 M held
~32 M more
have gap 148 of 178 inventoried sources don't match the corpus. Big missing: TX insurance agents 4.42 M, CT licences 2.66 M, WA providers 2.45 M, CO 1.61 M. Licence numbers fill the empty identifiers table.
Sales / use tax permit holders
state dept of revenue
bulk / APIbusiness 10–30 M est. untapped Best underused class. Only businesses with actual taxable revenue must register — filters out shell entities. CA CDTFA, TX Comptroller, WA DOR publish lists.
OpenStreetMap POIs
shop / amenity / office / craft
Geofabrik bulkbusiness 1–15 M US untapped Highest contact density of any free source — phone, website, street address, hours as first-class tags. Nothing of this class is in the corpus. One 11 GB download.
SAM.gov entity registrations API + keybusiness ~1 M+ needs key POC name + postal address only. Email/phone are FOUO → federal-only. Check the bulk extract before paging a 10 k/day API.
Federal award / vendor data
USAspending · FPDS · Grants.gov
API / bulkbusiness millions of awards untapped Entities proven to win paid work, with dollar amounts — a qualification signal nothing else gives.
IRS Business Master File
tax-exempt orgs
bulk CSVbusiness ~1.9 M partial Nonprofits and foundations with addresses; partial coverage already held via wave 6.
Commercial property & assessor rolls
county parcel + deed
ArcGIS / portalboth county × 3,000 untapped Owner-of-record names and mailing addresses — the landlord/property-holding segment, largely business entities.
UCC filings
secured lending
state SOS portalsbusiness state × 50 untapped Filing means a lender underwrote them — an active, creditworthy filter. Names + addresses.
Health & food inspection data ArcGIS / Socratabusiness city/county × 1,000s partial Restaurants and food service with premises addresses; continuous publishing, so it stays fresh.
Alcohol / cannabis / gaming licences state ABC bulkbusiness ~500 k+ partial TTB permittees partly held (1.01 M); state ABC lists add on-premise venues.
Childcare / private school licensing state + NCESbusiness ~500 k untapped Licensed facilities with addresses and often director names; NCES PSS covers private K-12.
Job postings
hiring = spending
Adzuna API · USAJobsbusiness continuous untapped Strongest timing signal available — a company hiring this week is buying this month.
Trade association & chamber directories scrapebusiness thousands of lists untapped Member lists publish contact details by design. Small volumes, excellent density, high-effort to collect.
Trade-show exhibitor lists scrapebusiness event × 1,000s untapped Exhibitors are pre-qualified buyers with budget. Usually the densest contacts per row anywhere.
Court · liens · bankruptcy · tax-delinquent court portals / stateboth continuous untapped Distress signals; also a rare source for sole-proprietor and officer names.
Franchise disclosure (FTC FDD) portal / scrapebusiness franchise systems untapped Franchisors plus their unit counts; franchisees are high-spend, well-capitalised operators.
Insurance producer / agency licences state DOIboth TX alone 4.42 M gap Already in the licensing inventory and unmatched — one of the largest single missing rosters.
Website enrichment of rows we already hold crawlbusiness = existing rows untapped Cheapest path to the metric that matters: turn rows we already own into rows with an email or phone. Respect robots.txt; expect IP throttling.
Fresh registration feeds Socrata / RSSbusiness ~39 k/month OR alone untapped A business registered this month needs everything. Highest intent, smallest volume — find the equivalents in each state.
FCC ULS licensees + tower registrations browser onlyboth ~1.5 M 403 to scripts Ten files lost to bot-blocking. Needs a real browser session — not a retry.
ATF federal firearms licensees browser onlybusiness ~140 k 403 to scripts Same gate; small volume, complete coverage of the segment.
California CSLB + DRE browser onlyboth ~280 k + ~600 k postback-gated The largest single blocked roster — ~880 k contractors and agents in one state.
FL Sunbiz · NJ · OH · VA · GA · HI · KY browser / accountbusiness millions gated Login walls, WAFs and JS-only apps. Browser session, one source at a time.
Google Places · SafeGraph · ZoomInfo · Data Axle paidboth paid Listed for completeness only. Not recommended without explicit authorisation to spend.

Contact density = how often a row carries a phone, email or website: high · medium · low.

Next five moves, in order

  1. Build the scoreboard first. Normalize the shards into a unified view and measure distinct entities and contact coverage. Without it, nothing below can be shown to have added anything — and this is also the honest answer to "how many businesses do we actually have?".
  2. Pull the missing licensing rosters (~32 M rows). Biggest verified inventory, rosters not permits, and the licence numbers repair the empty identifiers table — which is what makes cross-source dedupe possible at all.
  3. OpenStreetMap POIs. Highest contact density available for free and entirely absent today. One bulk download, then extract tags — the fastest route to "entities with a phone".
  4. Sales/use tax permit holders + fresh-registration feeds. Businesses with revenue and businesses that just formed — the two strongest filters on commercial intent that exist.
  5. Enrich the rows we already own. Crawl the website fields for emails and phones. Cheaper than any new source and it moves the only metric that matters.

Decisions needed

QuestionWhy it blocks
Is the goal 330 M rows, or distinct contactable entities? 330 M rows is reachable; 330 M US businesses is arithmetically impossible (~34–36 M exist). The plan optimises entities and cannot honestly chase the row number.
Public corpus or private bucket? Decides whether we publish a business-only cut with contacts stripped, or pay ~$1–2/month for a private bucket. Publishing as-is would expose consumer complaint records and individual practitioners.
API key: personal SAM.gov account, or register LCHWM LLC as a federal entity? A personal non-Federal account unlocks all Public data. Registering the LLC is a much heavier process and only pays off for contract bidding, not for data.
How much browser-babysitting is worth it? FCC, ATF and California are hours of hands-on work each. Breadth over depth says skip them and take the Socrata/ArcGIS and licensing rosters first.
Authorise paid sources? Currently treating every gated source as blocked. Nothing is purchased without an explicit yes.

Text version (survives a broken render)

WHAT WE HOLD — 186,461,039 rows / 8.1 GB / 1,356 shards
  unclassified (wave 6)      88,900,158   47.7%
  occupational licensing     46,895,732   25.2%
  arcgis / socrata sweep     30,520,090   16.4%
  federal                    15,580,374    8.4%
  permits                     2,435,553    1.3%
  health practitioners        1,206,782    0.6%
  derived (dentists)            922,350    0.5%
  QUEUED (779 sources)      ~23,280,927     mostly permits

FUNNEL        186.5M rows  ->  distinct entities UNKNOWN  ->  contactable UNKNOWN
              real US businesses ~34-36M   (330M would be ~97% of the population)

PRIORITY
  1 scoreboard: normalize + measure distinct and contactable
  2 missing licensing rosters (~32M rows)
  3 OpenStreetMap POIs (highest free contact density, absent today)
  4 sales/use tax permit holders + fresh registration feeds
  5 website enrichment of rows already owned
sources checked live this session: open.gsa.gov/api/entity-api (SAM.gov sensitivity tiers, API-key location) · sam.gov/data-services/Entity Management/Extract Data (empty for anonymous visitors) · api.data.gov/signup (no form served) · corpus_manifest.csv (186,461,039 rows / 8.1 GB / 1,356 shards) · wave6_jobs.csv vs .done markers (779 queued / ~23.3M rows) · licensing_sources.csv (178 sources / 37.9M verified rows; 148 unmatched) · DuckDB read_parquet union (17 s)
unverified / estimated: sales-tax and OSM volume ranges, all "untapped" volume figures, and the distinct-entity and contact-coverage numbers — which remain unknown until the scoreboard exists.