← The Observatory

Issue 07 / February 2026

The Small-Business Data Blind Spot

When the category conceals the business

“Small business” is one of the economy's most quoted categories and one of its least coherent. A better evidence system begins by asking what the average leaves out.

Evidence cutoff / February 27, 2026

The Small-Business Data Blind Spot

When the category conceals the business

A self-employed interpreter, a 70-person machine shop, a venture-backed software company, a family restaurant, and a 450-employee distributor can all appear under the heading “small business.” The category is administratively useful. Analytically, it can be close to absurd.

Each firm faces markets, financing, management, and risk. Beyond that, their operating realities diverge. One has no payroll; another has supervisors and a compliance team. One sells time; another holds inventory. One is designed to scale rapidly; another is designed to support a household and remain local.

When public debate says small businesses are optimistic, credit-constrained, adopting AI, hiring, struggling, or thriving, which of these firms is speaking?

The February problem is not a shortage of data. It is a shortage of fit between the categories we collect and the decisions we need to make.

Categories are instruments, not portraits

The small-business data problem begins with a category asked to do too many jobs. A threshold useful for administration becomes a portrait of economic life, even though firms cross legal forms, employment states, and operating models in ways the category cannot explain.

Definitions are necessary. Agencies need rules for eligibility, statistics, procurement, and enforcement. The U.S. Small Business Administration uses industry-specific size standards; many public discussions use a broad threshold of fewer than 500 employees. The Federal Reserve Small Business Credit Survey studies both employer and nonemployer firms through separate reports.

Each choice reveals something and hides something. A single employee threshold treats headcount as the decisive feature even when capital intensity, revenue, business model, and ownership structure differ. Revenue thresholds improve some comparisons and distort others. Administrative datasets offer coverage but may lag. Surveys provide experience but depend on sampling and response.

The mistake is not using a definition. It is forgetting that the definition is an instrument.

Suppose a program is designed to improve technology adoption among “SMEs.” Evidence that large-small firms adopt less than large enterprises may be true in aggregate. It says little about whether a four-person construction firm needs a scheduling tool, whether a manufacturer needs robotics integration, or whether a consultant needs secure knowledge management. The average barrier may not exist for any actual participant.

Nonemployer firms are particularly easy to misread. They can be treated as nascent employer firms, informal side businesses, or lifestyle enterprises. Some are all three. Many are none.

The Federal Reserve's 2025 nonemployer report described firms with and without hiring plans, varied financing behavior, and distinct operating challenges. The important analytical move is separation: a firm that intends to hire is facing a transition problem; a firm designed around one professional may be optimizing a different form of scale.

Consider two illustrative businesses with the same annual revenue. One is a solo consultant with low fixed costs and several clients. The other is a product seller with inventory, storage, shipping, and seasonal cash needs. A revenue statistic makes them peers. A financing model based on their operating cycles would not.

The nonemployer category also obscures labor. A business with no employees may coordinate contractors, family help, platforms, and outsourced services. It may create significant work without creating payroll. Conversely, a registered business may be dormant. Counting legal entities is not the same as understanding economic systems.

The most revealing small-business events often happen between datasets. A sole proprietor incorporates. A nonemployer hires one person. A shop closes its legal entity and reopens under a new owner. A consultant keeps the same tax identity while gradually becoming an agency. A founder takes a wage job but continues selling on weekends. Whether the business is recorded as surviving, dying, starting, or transforming depends partly on the administrative lens.

The Census Bureau's longitudinal work is valuable precisely because it tries to connect records across time. Its research on the transition from nonemployer to employer status found that the crossing is rare: among roughly 8.75 million firms first observed as nonemployers in 2011 or 2012, only about 1.8 percent became employers during the observation window, and most of those transitions occurred in the first four years. That statistic overturns an easy story. The vast nonemployer population is not simply a queue waiting to hire. Employer formation is a distinct event with predictors, costs, and timing.

The same research found associations with legal form, capital investment, time commitment, prior business experience, a website, and the kind of customer served. None is destiny. Together they show why a headcount category cannot explain the decision. A firm serving other businesses or government may encounter contracts that support payroll. A firm that supplies the owner's primary income may justify a different commitment than a side activity. Capital equipment can signal both ambition and the need for complementary labor. The transition emerges from a configuration.

Now consider the policy consequence. If a program judges success by jobs created, it may push stable solo businesses toward a boundary they do not need to cross. If it treats every nonemployer as a hobby, it may miss precisely the few approaching the transition. A better question is not “How do we make nonemployers hire?” It is “Which firms are approaching a recurring workload and revenue condition in which employment is the appropriate institution, and what prevents them from crossing safely?”

Data becomes useful when it helps identify a decision without pretending to decide it.

Categories allocate more than attention. They allocate eligibility, enforcement, finance, and public sympathy. A size threshold can exempt a business from one burden while excluding it from another benefit. A disadvantaged-business designation can create access and also require owners to translate identity into administrative proof. Industry codes determine which comparisons and programs appear relevant, even when a company spans several activities.

Classification can also harden a theory of what a legitimate business looks like. Firms with regular payroll, formal accounts, conventional premises, and stable ownership are easiest to observe. Household enterprises, seasonal businesses, multilingual firms, cooperatives, and companies mixing digital and physical work can fit badly. The more an institution automates decisions, the more consequential that fit becomes.

There is no classification without a boundary case. The answer is not infinite categories. It is contestability. A firm should be able to see how it was classified, understand what follows, supply contrary evidence, and obtain human review where the consequence is serious. Researchers should publish enough metadata to show who was excluded and how definitions changed.

This is not a technical appendix to fairness. It is how a data system acknowledges that administrative simplicity has been purchased with somebody else's complexity.

What the dashboard cannot see

Aggregation then gives the category visual authority. A dashboard can render a trend precisely while leaving its denominator, selection effects, and unrealized intentions outside the frame.

Dashboards promise current, comparable, decision-ready evidence. The format encourages confidence: clean trend lines, ranked regions, colored indicators. But a dashboard is a theory expressed through selection.

Which firms are captured? How are closures distinguished from transitions? Does a revenue index reflect price increases or real volume? Are rural firms visible? Can the measure detect a household business operating partly in cash? When an average improves, who deteriorated?

The U.S. Census Bureau's Business Formation Statistics provide valuable signals about applications and likely employer businesses. They do not report how many applications will become durable enterprises or why a person filed. The Bureau of Economic Analysis can illuminate small-business contribution. It cannot show the unpriced coordination work of an owner's spouse or the customer lost because a permit took three visits.

This is not a criticism of official statistics. It is a boundary statement. Administrative data is strongest when the administrative event resembles the economic question. Lived experience is strongest when interpretation and sequence matter. Neither should impersonate the other.

Suppose a report says that 20 percent of firms applied for financing. The immediate temptation is to compare that share across places or groups. But what is the denominator? All firms? Firms that needed money? Firms aware of a suitable product? Firms confident enough to apply? Surviving firms at the time of the survey?

Each denominator tells a different story. Applications divided by all firms may look low because many firms are adequately financed. Applications divided by firms that needed credit may reveal discouragement. Approval among applicants can rise even as access worsens if the most uncertain applicants stop applying. This is selection, not pedantry. The people missing from the denominator may be the policy problem.

Survivorship produces a related distortion. Interviews with active owners can illuminate adaptation but cannot directly explain why similar firms ceased trading. Reviews left on a platform describe users who reached the platform, purchased, and cared enough to respond. Program completion rates exclude those who never entered. Procurement award data excludes firms that never found a tender, judged it unsuitable, or abandoned the paperwork. At every stage, the visible group is selected by the system being evaluated.

The honest analyst therefore narrates the funnel. How many firms plausibly encountered the need? How many knew of the pathway? How many considered it? How many began? Where did they stop? What happened afterward? The result is less elegant than a single rate and much more useful.

Business statistics usually privilege realized events: revenue, payroll, applications, closures, shipments. Yet owners act on expectations. They order inventory because they expect demand, decline a lease because they fear a downturn, or postpone hiring because one large contract feels uncertain. The expectation changes the economy before the predicted event occurs.

This creates an awkward measurement problem. Expectations are subjective, noisy, and influenced by the wording and timing of a survey. They are also indispensable. A business can report current stability while abandoning an expansion. A founder can report optimism and still be unable to finance the inventory that optimism implies.

The Census Bureau's Business Trends and Outlook Survey and the Federal Reserve's Small Business Credit Survey collect expectations alongside conditions. Their greatest value may not be in producing an optimism index. It may be in observing the gap between what firms anticipate and what happens next. Which owners systematically underestimate implementation time? When does expected hiring become payroll? Does an anticipated financing need produce an application, discouragement, or adaptation? A longitudinal system can learn where confidence is predictive and where constraints sever intention from action.

That gap matters for support design. If firms repeatedly intend to adopt a technology but do not, another awareness campaign misunderstands the evidence. The blockage may be vendor selection, data cleanup, employee training, or fear of disrupting a working process. If owners expect to bid but withdraw after reading the terms, the issue is not opportunity discovery. An intention is not a soft, inferior form of data. It is the first half of a behavioral sequence.

Recovering the mechanism

The remedy is not anecdote in place of statistics. It is research capable of reconstructing the mechanism between a reported condition and an observed outcome—why the owner acted, which constraint bound, and what changed next.

Imagine a survey in which owners select their greatest challenge. One chooses “access to finance.” Behind that answer are several possible stories.

The first owner applied to a bank and was declined because of credit history. The second did not apply because the process looked futile. The third received credit but at a price that made the opportunity irrational. The fourth has adequate revenue but waits 60 days for customer payment. The fifth needs $8,000, an amount too small for one lender's process and too large for a credit card balance. The sixth does not need debt; she needs customers to pay deposits.

The category identifies a domain. It does not identify an intervention.

This is where qualitative research earns its place. A sequence interview can reconstruct the decision: What happened before the owner sought money? What alternatives were considered? Which document, term, or expectation changed the outcome? What did the person do next? The aim is not to replace measurement with anecdotes. It is to discover the mechanisms that measurement should distinguish.

Take an illustrative neighborhood bakery whose legal entity disappears from an administrative register. One interpretation is failure. Yet several materially different events could sit underneath the same disappearance.

The owner may have sold the assets to an employee who opened a new entity. The family may have closed an unprofitable storefront while keeping a profitable wholesale line under another registration. A health event may have ended a viable business. A landlord's redevelopment may have displaced it. Or demand may genuinely have collapsed. For employment statistics, some of these paths look similar. For economic development, they imply different remedies: succession finance, channel strategy, social insurance, commercial-space policy, or no intervention at all.

Closure interviews are difficult because the subject is sensitive and the owner has moved on. That difficulty does not make the missing information unimportant. It means the evidence system is institutionally biased toward organizations still available to answer. Research partnerships with accountants, community lenders, chambers, and local associations could create ethical ways to understand transitions after the fact, provided participation is voluntary and commercial details are protected.

The point is not to rescue every business from exit. Market economies require entry, adaptation, and exit. The point is to distinguish productive reallocation from preventable destruction. A restaurant closing because customers prefer something else is not the same phenomenon as a profitable restaurant closing because a bridge repair removed access for eighteen months. “Business deaths increased” is a signal. It is not yet an explanation.

Small-business policy is full of plausible associations. Firms that receive advice may survive longer. Exporters may be more productive. Digitally intensive firms may grow faster. Businesses with banking relationships may obtain better terms. The trouble is that capable, ambitious, or better-connected firms are often more likely to seek these things in the first place.

Correlation still matters. It can identify patterns worth investigating. But a venture or policy built on an association should say what causal story it believes. Did advice change the decision, or did firms already likely to succeed select into advice? Did technology raise productivity, or could productive firms afford technology? Did certification create access, or did buyers help favored suppliers become certified?

Randomized trials can answer some questions, but they are neither universally ethical nor sufficient. A well-designed pilot can stagger access, compare alternative messages, or randomize the intensity of support. Natural experiments can use a rule change, threshold, or geographic boundary. Longitudinal comparison can follow similar firms through different sequences. Qualitative process tracing can test whether the mechanism actually appeared. The methods should be chosen for the decision, not for prestige.

The Observatory should therefore grade claims by what the evidence permits. “Participants reported saving time” is not “the service caused productivity growth.” “Awardees later hired” is not “the grant created jobs.” A credible institution does not weaken its argument by naming uncertainty. It prevents a hopeful pattern from becoming a false promise.

Evidence as institutional memory

Once evidence is organized around decisions rather than labels, research can begin to accumulate. But an institution that remembers a firm also acquires power over it, making provenance, correction, reciprocity, and purpose part of the research method.

Organizations often commission studies that end as PDFs. A survey establishes a fact, a workshop adds context, a pilot produces lessons, and a new team begins again eighteen months later. The evidence exists. The institution cannot use it as memory.

The Observatory itself could reproduce this failure. Twelve monthly issues would be easy to publish and difficult to accumulate. Themes must carry forward with provenance: which claim was observed, which was inferred, which remained a question, and which later evidence revised it.

A research memory should preserve disagreement as well as conclusion. If owners say a portal is unusable while administrative data shows high completion, both facts matter. Completion may conceal hours of work. Interviews may overrepresent frustrated users. The contradiction is not noise to remove; it is the next research design.

This is especially important when research becomes venture design. A market narrative can quickly harden into a product premise. The sentence “small firms struggle to find opportunities” may produce another database. A more careful evidence chain might show that discovery is adequate and fit assessment is the real burden. The difference is the difference between a useful experiment and a costly category error.

Better data is not automatically benign. Increased legibility can help firms obtain credit, contracts, advice, and visibility. It can also increase surveillance, enable predatory targeting, or let platforms extract value from relationships they did not build.

A small-business intelligence system must therefore ask: legible to whom, for what purpose, under whose control, and with what right to correction?

Data collected to coordinate pooled purchasing should not become a credit score without explicit consent. A record of declined opportunities should not be treated as evidence of low ambition. A founder's private explanation of cash stress should not become a marketing segment. Purpose limitation is not a technical detail; it is a condition of trust.

There is also a representational risk. Models trained on firms that already succeed in formal systems may define readiness in their image. A capable business with informal records, a nonstandard career path, or a culturally distinct communication style can appear anomalous when the data structure is the problem.

No single dataset can resolve the blind spot. A useful stack would combine four layers.

The first is administrative evidence: formations, payroll, revenue, procurement, trade, and other observable events. The second is structured survey evidence about conditions, behavior, and expectations. The third is longitudinal qualitative evidence that follows decisions across time. The fourth is operating evidence from experiments: what people actually do when a new pathway is available.

The layers should not be blended into false precision. They should challenge one another. If a survey says owners want AI support and a pilot shows they do not return after setup, the discrepancy is the finding. If portal traffic rises while bid participation remains flat, discovery may not be access.

Segmentation should be task-specific. For a hiring question, distinguish employer status, role, revenue stability, and managerial readiness. For a logistics question, distinguish shipment size, product, corridor, and frequency. For technology, distinguish use case, process maturity, data sensitivity, and internal skill. “Small” remains a scope, not an explanation.

The deepest evidence would follow businesses through decisions rather than repeatedly surveying fresh cross-sections. That creates a relationship, not merely a dataset. Participants would see what is collected, receive useful analysis in return, correct their records, and decide which secondary uses are acceptable. The institution would commit to returning findings in language that helps owners act.

Such a compact changes the incentive to participate. Too much research asks small firms to donate time so that an agency, consultant, or platform can become smarter. The owner receives a report months later, if anything. A reciprocal system might return a financing timeline, benchmark an operating process, preserve the evidence used for a bid, or show how similar transitions unfolded without exposing individual peers.

Reciprocity also improves data quality. A business is more likely to update a profile that gives something back. Corrections become part of the research process. Missingness itself becomes interpretable: did a firm leave because it closed, because the service stopped being useful, or because trust failed?

The compact would have to withstand the moment when valuable data attracts another purpose. Could a lender buy access? Could a public agency use it for enforcement? Could CKOS train a commercial model? Consent obtained at enrollment is not a blank cheque. Governance must travel with the data as the institution discovers what the data is worth.

The person behind the denominator

The aim is not to escape averages. It is to return from the average to the person without pretending that either view is complete on its own.

Economies need averages. Institutions need categories. Ventures need addressable markets. The danger begins when these tools become portraits.

An owner does not experience the average approval rate. She experiences one application with a particular balance sheet, lender, purpose, and week of her life. A region does not experience a mean productivity gap. Its firms experience roads, buyers, advice, skills, rents, and relationships.

What would change if research began with the decision rather than the category? We might ask not “What do small businesses need?” but “What must this kind of firm know and be able to do before this kind of opportunity becomes usable?”

That question produces a narrower market and a better intervention.

The small-business data blind spot is not darkness. It is glare from an average held too close to the eye.

Research lineage

Observed evidence. Official and survey sources reveal important differences between employer and nonemployer firms, but broad categories remain heterogeneous. Administrative and experiential evidence answer different questions.

Interpretation. Many support failures originate in category mismatch: an average condition is treated as a customer need. Publication without lineage creates institutional amnesia.

Hypothesis. Decision-specific segmentation and a source-aware research memory can improve both policy and venture design.

Questions carried forward. What is the minimum viable profile for a decision? Which differences change the recommended action? How should owners govern secondary use of their operating data?

Sources and further reading

  1. Federal Reserve Banks, Small Business Credit Survey, methodology and reports through February 2026.
  2. Federal Reserve Banks, 2025 Report on Nonemployer Firms, 2025.
  3. U.S. Census Bureau, Business Formation Statistics, February 2026.
  4. U.S. Census Bureau, Nonemployer Statistics, latest available.
  5. U.S. Bureau of Economic Analysis, Small Business, latest available.
  6. U.K. House of Commons, small-business strategy and evidence, 2025–2026.
  7. Alicia Robb and Adji Fatou Diagne, U.S. Census Bureau, Startup Dynamics: Transitioning from Nonemployer Firms to Employer Firms, Survival, and Job Creation, April 2025.
  8. Adela Luque and Vitaliy Novik, U.S. Census Bureau, Garage Entrepreneurs or Just Self-Employed? An Investigation into Nonemployer Entrepreneurship, October 2024.

Evidence cutoff: February 27, 2026. The financing and business examples are illustrative composites.