Every dataset behind Hiasynth, with publisher, role, vintage and licence
Sources
Every number Hiasynth returns traces back to a published dataset with a named publisher, a stated vintage and a licence we are entitled to use. If you cannot check where a figure came from, you should not stake a decision on it. So this page exists to be checked.
What we use, and what for
Nothing here comes from web scraping, ad-tech exhaust, purchased consumer files or panel providers. Every input is an official statistical product, a freely available academic survey, an open geospatial layer, or a commercial product we hold a licence for.
Which source does which job, the dataset codes, the extraction dates and the vintage each build was frozen against are not published. If you need them for procurement, due diligence or an audit, email hello@hiasynth.co and we will send the set that covers your case.
On the European Social Survey. We use rounds 1 to 11. The data is freely available to download from the ESS data portal, and the same source data is there for anyone who wants to check what we did with it. We are not affiliated with, endorsed by or authorised by ESS ERIC, and nothing on this page should be read as their view. Each round is cited with its own DOI:
European Social Survey European Research Infrastructure (ESS ERIC) (2025)
ESS11 - integrated file, edition 4.1 [Data set]. Sikt - Norwegian Agency for
Shared Services in Education and Research. https://doi.org/10.21338/ess11e04_1
Traceable one number at a time
Provenance lives in the data itself. Every attribute carries a source string naming the table, the wave and any assumption applied to it, so a figure is traceable straight to its origin.
The one layer that is not a dataset
Cultural identity attributes, things like religion, dietary practice and calendar, are our own research rather than an extract from anyone's database. The resulting value is an informed estimate, not a lookup.
The distinction that matters is between learning from published work and copying it. Frameworks, categories and structure - which identities exist, which dimensions matter, how they relate - are knowledge, and reading the research is how anyone acquires it. The numbers attached to that structure are ours: a synthesis built from that reading and from open statistics. Where a published figure shaped an estimate, it did so as influence on our own work, not as a value lifted into a table.
No premium or restricted source appears raw in this layer, and none is cited as though it does. That cuts both ways. We do not claim a source we only learned from, and we do not ship a number we were not entitled to.
Coverage is uneven and we would rather say so than smooth over it. Real identity structure exists for some countries and not others. Where it does not, the layer is suppressed: the attributes come back null and stamped as suppressed, never quietly defaulted to the majority. A plausible-looking cultural profile for a country we hold no seed data on is invention, and it would be the kind that is hardest to catch, because it looks exactly like the real thing.
We say so here rather than list it in the table as though it came out of a database. It is the one part of the build where the honest description is "we worked it out", and you should weigh it accordingly. The method is written up in full and we will send it on request.
Where we draw the line
We do not estimate a country into existence. Four countries sit outside the current build: Ukraine, Bosnia and Herzegovina, Moldova and Kosovo. None has a standard European regional geography to build against, and in Ukraine's case wartime displacement means any zone-level structure for 2024 would be invention. A synthetic population that fabricates on demand is worse than no population at all, so the tools say the country is unavailable instead.
We do not blend in data we cannot name. If a source cannot appear on this page with a publisher and a licence, it does not go into a build. Several well-known survey programmes are permanently excluded because their licences do not permit commercial use. Their microdata is in no build. Wanting the coverage does not change that.
We do not use personal data. No input contains identifiable individuals and no output does either. See methodology for why that is structural rather than procedural.
Vintages and drift
Each build is frozen against a stated reference year, currently 2024, and carries a macro context vector so cross-year comparisons stay honest. Statistical offices revise their published tables. When a revision lands it is picked up in the next version of the population, leaving a shipped one exactly as you queried it. Which version incorporated which revision is recorded in the release lineage.
Attribution
If you publish analysis derived from Hiasynth, cite Hiasynth and the reference year of the build you queried:
Hiasynth, Humanity (reference year 2024). https://hiasynth.co
Several inputs carry their own attribution requirements. Where your publication reproduces figures deriving substantially from a single upstream source, cite that source too. Ask and we will tell you which ones apply to a given query.
Corrections
If you think a figure is wrong, we would rather hear it: hello@hiasynth.co. Send the query and the result. Confirmed errors are fixed in the next release and noted in its entry.