How Hiasynth is built, what it guarantees, and how that was tested
Methodology
Audience: technical users and decision-makers deciding whether to trust Hiasynth with a real decision.
This page covers what the model guarantees, how those guarantees were tested, and where it breaks. It argues against us wherever the evidence does. Construction detail is not published. If you need it for procurement, due diligence or an audit, email hello@hiasynth.co.
The problem
Market analysis has two sources of population truth and neither one is enough.
Official statistics are large, exact and public, and they only publish margins. Age by sex per area. Employment by sex per area. You can answer how many. You cannot answer how those things go together, because the cross-tabulations were never published.
Survey microdata has the joint structure. It knows how trust moves with income, how education moves with values. It also has a few thousand respondents per country and no idea which street any of them live on.
So you can have the correlations or you can have the geography. Every market intelligence product you have used has quietly picked one.
What we do about it
Model the shape, rake to the truth.
The joint structure comes from survey microdata: a density model that learns how attributes move together across a population. The levels come from official statistics: the output is raked until aggregating it reproduces the published counts, region by region, cell by cell.
The result is a population where the totals are the statistical office's totals, and the combinations are the ones the surveys observed. Neither source is right on its own. That is the whole idea.
Every person is then placed. Not in a region, in a square kilometre, inside a household, with the household's income and dwelling and car and heating bill, in a specific climate. Context is what makes a persona mean anything.
What it guarantees
Counts. Aggregate the population in any built region and you get the number the national statistics office published, within 0.1%. This is a build gate. A build that misses it does not ship.
Combinations. Sampled people carry realistic combinations of attributes. A high-income, low-education, high-trust person turns up as often as the survey says such people turn up.
Budgets. Every zone's population is exact. No rounding drift, no zone quietly gaining or losing people.
Coherence. Nobody is lost and nobody is counted twice. Add up the households in a region and you get its population, including the people official statistics keep outside household counts: students in halls, residents in care homes. Size a market by households or by people and the two answers agree.
Privacy by construction
Every synthetic person is a fresh draw from a fitted distribution. Their attributes are generated in that moment. Two people conditioned on identical survey structure still come out different.
Hiasynth is anonymous by construction. There is no real person behind any row, so there is nothing to re-identify. It is a property of how the thing is built.
No input contains identifiable individuals. No output does either.
Validation
Anchors
Population and household totals land within 0.1% of the national statistics office figure for every built country. Where Eurostat and the national office disagree, and they do, by up to 7% on some country totals, we anchor to the national office and use Eurostat only for the regional distribution.
Marginal reproduction is gated on KL divergence below 0.1 against Eurostat, per country, per build.
Held-out checks
The interesting tests are the ones against data the build never saw.
France, secondary dwellings. The model derives a 10.3% secondary-dwelling share from the gap between housing stock and resident population. INSEE publishes 10%.
Netherlands, neighbourhood education. Ranking Dutch neighbourhoods by education level against CBS neighbourhood data gives a Spearman correlation of 0.59. National distributions match OECD: 45% tertiary, 80.5% employment.
Regional gradients. The model resolves variation that national averages erase. French heat pump penetration ranges from 2.6% to 35% across regions, matching the INSEE IRIS gradient. Mean household size comes out at 2.31 in Paris against 1.76 in Burgundy, 2.78 in Małopolska against 2.34 in Zachodniopomorskie. These are not tuned. They fall out of the placement.
Still open
The headline external validation is a held-out attitudinal comparison against a European survey programme that was never an input to the build. It is specced and not yet run. We declare known defects before the run and publish the results either way.
We would rather say that than imply it is done.
Known limits
Absolute income levels. Euro figures run low in high-income countries. Relative position, quintiles, rankings and within-country comparison are sound. Treat cross-country absolute euro figures as directional until the next release. This is a known defect with a known cause, and it is being fixed.
Children. Everyone under 15 is in the population, placed in their household and counted in every total. Their profile is shorter, because the survey evidence behind values, attitudes and behaviour starts at 15. Extending it downward is planned for v3.
Attitudinal coverage. The attitudinal structure comes from a survey programme covering 28 countries. The build covers 38. The countries without a national sample inherit a pan-European structure conditioned on their own demographics, geography and macro context. That is a reasonable estimate and it is weaker than a measured one, so the data flags it.
Grain versus detail. Aggregate to a region or a country and the numbers are precise. Cut to one square kilometre and stack five filters and you are reading an estimate. The model is a statistical construction, and it stops being able to tell you things at some depth. It says so when it gets there.
Occupation and industry gaps. A handful of Western Balkan countries have no European labour force coverage for occupation and industry. Those columns come back empty.
Time. Reference year 2024. A snapshot. Backwards extension and forward projection with fertility, migration and mortality are architected and not shipped. When they ship, they ship with a backtest.
Where the data comes from
Every input is a named, published dataset with a publisher, a vintage and a licence we are entitled to use. The list is on the sources page, and every attribute in the data carries its own source string, so a figure is traceable to its origin.