No single source of industry market share data tells the complete story. Association programmes miss the non-members and members who do not participate, customs records bundle products under broad tariff codes and include used equipment, and commissioned research relies on sampling. Blending association retail sales data with cleaned customs statistics narrows the uncertainty and reveals the true scale of the market - and monthly collection keeps that picture near-real-time.
Every equipment manufacturer wants to know exactly where they stand in the market. The trouble is, no single source of industry market share data tells the complete story.
Whether you rely on association statistics, customs declarations, commissioned research or your own internal records, each dataset carries blind spots that can lead to overconfident decisions - or missed opportunities entirely. In this post, we explore why market share numbers are inherently incomplete, how blending multiple data sources closes the gap and what this means for OEMs making confident, evidence-based decisions about production planning, territory strategy and competitive positioning.
Every data source has a blind spot
It is tempting to treat a market share figure as settled fact. A number arrives in a report, it looks precise and it becomes the basis for strategic planning. The reality, however, is more nuanced.
Consider association-run statistics programmes - one of the most common sources of industry market share data for equipment manufacturers. These programmes collect retail sales data directly from participating companies and aggregate it into industry totals. The data is granular, timely and operationally relevant. But not every company in an industry participates.
On average, roughly twenty percent of industry players sit outside any given statistics project. Some are not association members at all. Others are members who choose not to take part. The reasons vary widely:
- Smaller companies sometimes assume benchmarking offers no return on investment because they feel they cannot compete with the market leaders.
- Larger companies may perceive a strategic risk in disclosing their sales volumes and stronghold positions.
The result is the same either way: reported market share figures are based on an incomplete participant pool, which means they are structurally inflated relative to the true market.
This is not a flaw unique to association data. Customs import-export statistics carry their own accuracy issues. Commissioned research relies on survey sampling and extrapolation. Internal sales records reflect only your own activity. Every source paints a different picture, and none of them is the whole truth.
Why customs data alone falls short
When companies recognise the participation gap in association statistics, a common next step is to cross-reference against customs import-export data. In principle, this makes sense. Customs data captures all goods crossing a border, regardless of whether the manufacturer or distributor participates in an industry benchmarking programme.
In practice, however, customs data introduces its own complications.
Broad tariff codes and misclassification
Importers and exporters are frequently inaccurate when classifying goods under the correct tariff code. The harmonised system (HS) codes used by customs agencies are often broad enough to capture several product categories under a single heading. A tariff code labelled "handheld power equipment, excluding saws and blades" might include hedge trimmers, blowers and several other product types - making it impossible to isolate the exact category you need.
On top of that, HS codes are only updated every five years, with the next revision not due until early 2027. Until then, if a new product type emerges in your industry, customs data simply cannot track it. There is no mechanism to distinguish a new category from the broad bucket it falls into.
These limitations mean customs data works better as a directional gauge than a precision instrument. It can indicate the approximate size of the total market, but it cannot tell you who is selling what, where or to what broad market segment.
Cleaning the noise
There are ways to improve the usefulness of customs data before you combine it with other sources. A practical approach involves examining three data points for each customs entry:
- the unit count
- the gross weight
- the declared customs value.
By dividing weight and value by unit count, you can derive average weight per unit and average value per unit. These averages make it straightforward to filter out obvious misclassifications. If a customs entry shows a dump truck with a declared value of $17, or heavy equipment with an average weight under 50 kilograms, those rows clearly do not belong.
The thresholds for filtering vary by industry and product type. A $50 average value is entirely reasonable for a hand tool but impossible for an industrial lift truck. Setting these thresholds requires product knowledge, but once established, they remain relatively stable within any given five-year HS code cycle.
This cleaning process removes the most egregious errors. It cannot, however, resolve the subtler miscategorisations baked into broad tariff codes. The result is data that is cleaner than raw customs declarations but still somewhat imprecise at the product level.
Inclusion of used equipment
As a general rule, customs departments require exporters and importers to select the HS code that most closely describes their product. The primary purpose of this classification is to levy the correct tariff or duty - the gathering of statistical information is merely a by-product of that legislative requirement. As it happens, customs tariffs do not distinguish between brand-new units - which are of most interest to OEMs - and second-hand or used items.
Because of this, the figures OEMs often see from a customs-derived data source will generally be higher than those produced by regional data collection programmes run by a local association directly from OEMs and their distribution networks.
The case for blending industry market share data sources
The most reliable picture of market position comes not from choosing the best single source, but from combining multiple imperfect sources and looking for correlations between them.
Association data is granular and operationally specific - it can capture great detail about each retail sale, for example by model code, geography and sometimes customer type. However, it only covers participating companies. Customs data is broad and captures all import activity, but it lacks product-level precision and is prone to classification errors. Each dataset compensates for the other's weaknesses.
When the two are visualised together - for example, as overlaid trend lines on the same graph - patterns emerge. You can see the correlation between association-reported volumes and customs-declared volumes over rolling twelve-month periods. The gap between the two lines represents the portion of the market not captured by association participants. That gap gives OEMs a practical way to estimate their true national market share rather than their share of the participating group alone.
Two imperfect datasets beat one or none
The logic is straightforward:
- A company with no external data is making decisions in the dark.
- A company with one data source can make better decisions, but may be misled by that source's specific blind spots.
- A company that blends two or more sources - understanding the limitations of each - achieves the best available clarity.
This is not about finding a single definitive number. It is about narrowing the range of uncertainty to a point where strategic decisions can be made with confidence. Most OEMs find that the blended approach gives them enough clarity to allocate inventory, plan territory coverage and evaluate competitive positioning with a level of assurance they cannot get from any individual dataset.
When outliers distort the picture
The importance of data blending is perhaps best illustrated by what happens when one source goes wrong. Consider a real scenario involving customs data from a Southeast Asian market. A single import entry recorded over 1.2 million hydraulic excavators arriving in one month - a figure that exceeded the country's total annual domestic market several times over. The error was likely a data entry mistake at the customs level, where a monetary value may have been entered in the unit count field.
That single outlier completely distorted the customs-based market share calculation, making one manufacturer's position appear to have collapsed. Without a second data source to cross-reference against, the company might have launched an urgent strategic review chasing a decline that never actually happened. Identifying and removing the outlier through standard filtering restored a picture that aligned with the association data and with the company's own operational experience on the ground.
Why real-time data changes the equation
Beyond accuracy, there is the question of timeliness. Industry market share data is only useful if it reflects what is happening now, not what happened six or twelve months ago.
Association-based statistics programmes that collect retail sales data monthly - with submissions typically due around the tenth of the following month - provide a near-real-time view of market activity. This is fundamentally different from commissioned research, which typically delivers annual or biannual snapshots, or customs data, which may lag by several months depending on the reporting country.
Monthly frequency also enables something that annual reports cannot: the ability to spot trends early. A gradual shift in regional demand, a new competitor gaining traction or a seasonal pattern that departs from historical norms - these signals appear in monthly data well before they surface in an annual study.
There is also the matter of adaptability. When a new product type enters the market, a monthly data collection programme can begin tracking it within days, as soon as enough participants agree to report on it. Customs data, by contrast, remains locked into its existing classification structure until the next HS code revision - which could be years away. For industries where product innovation moves faster than regulatory classification, this difference is significant.
The give-and-take model behind reliable data
The quality of any collaborative data exchange depends on who participates and how committed they are to submitting accurate, timely data. This is why the structure of the exchange matters as much as the data itself.
PowerStats operates on a give-and-take model. Participation is not a transaction where data can be purchased. Every participant must submit their own retail sales data on time, every time, in order to receive access to the aggregated industry totals and averages. This structure ensures that all participants have a vested interest in data quality, because the insights they receive are only as good as the inputs from the group.
This model is different from both data aggregators and research firms:
- Aggregators purchase or collect data and then sell packaged reports to anyone willing to pay.
- Research firms conduct surveys of a representative sample and extrapolate to estimate the total market.
The PowerStats model creates a closed-loop exchange where real operational data flows continuously between participants through a neutral facilitator - producing a level of detail and frequency that neither aggregators nor research firms can match.
Transparency is built into the process. All participants know who else is in the project, which allows each company to make its own assessment of how representative the participant group is. If a major player is absent, participants can factor that into their interpretation. If the participant pool is strong, confidence in the data is correspondingly high.
AI cannot replace the source
A question that comes up increasingly is whether artificial intelligence will disrupt the industry data space. The short answer is: not in the way most people expect.
The retail sales data that powers market share statistics - often detailed to the level of model codes, customer types, postcodes and monthly volumes - is commercially sensitive information that does not exist in any public dataset. No AI system can generate, extrapolate or forecast this data because there is nothing publicly available to train on or scrape. The only way to obtain it is to collect it directly from the companies that hold it.
What AI can do is enhance how participants interact with the data once it has been collected and processed. Natural language interfaces, automated trend detection and AI-generated summaries can compress the time between data availability and strategic insight - turning what was previously weeks of analyst processing into near-instant delivery.
Key takeaways
- No single source of market share data is complete or perfectly accurate - every dataset has structural blind spots that must be understood.
- Blending association retail sales data with customs statistics narrows the range of uncertainty and reveals the true scale of the market.
- Monthly data collection provides near-real-time visibility that annual research studies and lagging customs reports cannot match.
- Collaborative data exchange built on a give-and-take model produces richer and more reliable insights than purchased reports or one-off surveys.
- AI will enhance how industry data is consumed and interpreted, but it cannot replace the need to collect operational data directly from the source.
Frequently asked questions
Why is industry market share data never fully accurate?
Every source carries structural blind spots. Association programmes only cover participating companies - that naturally don't include non-member companies and members who don't join in - while customs data suffers from broad tariff codes and misclassification and commissioned research relies on survey sampling and extrapolation. The practical answer is to blend sources rather than search for a single perfect number.
How do you combine association data with customs data?
Clean the customs entries first - dividing gross weight and declared value by unit count exposes obvious misclassifications - then overlay the two sources as trend lines across rolling twelve-month periods. The gap between the lines represents the portion of the market not captured by association participants, which lets an OEM estimate true national market share rather than share of the participating group alone.
Why do customs figures run higher than association statistics?
Customs tariffs do not distinguish between brand-new and second-hand units, so used equipment is counted alongside new sales. Broad HS codes can also sweep several product categories into one heading. Both effects push customs-derived market figures above those produced by regional collection programmes that gather data directly from OEMs and their distribution networks.
Can AI replace industry market share data collection?
No. The retail sales data behind market share statistics is commercially sensitive and does not exist in any public dataset, so there is nothing for an AI system to train on or scrape. AI improves how the data is consumed once collected - summaries, trend detection and natural language interfaces - but the data itself must still come directly from the companies that hold it.
Make confident decisions with better data
Understanding your true market position starts with acknowledging that no single data source has all the answers. The companies that gain the clearest competitive intelligence are those that combine multiple sources, understand the limitations of each and use the correlations between them to make sharper strategic decisions.
If your organisation is ready to move beyond single-source market share estimates and explore what a blended, real-time data approach could look like, contact PowerStats to learn how we can support your industry benchmarking needs.



