The National Genomics Platform is the shared national infrastructure where Sweden's seven Genomic Medicine Centres store genomic data and associated health data generated in clinical care. It is built by Genomic Medicine Sweden and operated by Västra Götalandsregionen. For researchers, the platform is a route to clinically generated genomic data that would otherwise have to be requested region by region, from separate systems, in separate formats.
Access is possible but it is not open. The healthcare region that generated the data is responsible for its data in the NGP, and every release is a formal disclosure decision made by that region. This page sets out what the platform holds, how to find out whether the cohort you need exists, and the sequence of approvals you need before any data moves.
At a glance
Operated by
Västra Götalandsregionen, on behalf of Genomic Medicine Sweden
Data controller
Each healthcare region, for its own data
Data types
Genomic raw data and analysed data, plus associated health and sample metadata
Genomic sequencing in Swedish healthcare happens in the seven Genomic Medicine Centres, each attached to a university hospital region. Historically each centre kept its data in its own laboratory and storage systems, which meant negotiating separately with each for any study spanning more than one region.
The NGP consolidates that storage without consolidating control. Genomic data and metadata are transferred from local laboratory information systems over Sjunet or an encrypted connection into a storage area reserved for that region.
Researchers authenticate through SWAMID.
The seven Genomic Medicine Centres are GMC Norr (Umeå), GMC Uppsala, GMC Karolinska (Stockholm), GMC Örebro, GMC Sydöst (Linköping), GMC Väst (Gothenburg) and GMC Syd (Lund and Malmö).
What data the platform holds
The platform holds two broad classes of information.
Genomic data, covering inherited and acquired genetic variants, stored as raw unprocessed data, as analysed data describing identified genetic changes, or both, together with quality metrics from the laboratory and bioinformatic processes.
Associated data about the patient or research subject, which can include identity, diagnosis, the clinical question, phenotypic information, and details of the sample and the analysis performed.
Note that directly identifying information is within scope. The NGP is not a de-identified archive, and which form of data a region can release to a given project depends on that project's ethical approval.
Assay types present across the diagnostic areas include:
Whole genome sequencing, short read, in rare and hereditary disease, acute leukaemia and childhood cancer
Whole exome sequencing in rare disease
Targeted panels, including the GMS560 solid tumour panel covering roughly 500 genes, and myeloid and lymphoid malignancy panels
RNA and whole transcriptome sequencing in solid tumours, under evaluation for rare disease
Long read sequencing, under evaluation for repeat expansions and complex structural variants
16S amplicon and shotgun metagenomics in clinical microbiology
Retrospective collection covers cancer, rare diagnoses and microbiology from 2019 onward.
What the platform does not hold. Medical images and clinical records as such are not stored. Electronic informed consent is not yet implemented as a data object, so consent is handled through each study's ethics process rather than queried in the platform. Coverage is uneven: national sequencing volumes are far larger than what has so far been uploaded, so national activity figures are not a guide to NGP holdings.
Finding out what exists before you apply
There is no self-service catalogue of NGP holdings for researchers today. Building one, as a searchable metadata catalogue, is an active development goal. Until it exists, use these routes.
Feasibility portal
Aggregate counts by diagnosis and molecular criterion, per region. Designed for study sizing, discloses no confidential information, and currently in user testing.
GA4GH Beacon
The NGP participates in the GA4GH Beacon network, which answers whether a genetic variant is present in a dataset. Beacons return presence and count information rather than records.
The Genomic Medicine Centre in the region
For anything the automated surfaces cannot answer, the GMC is the authoritative source on what its region holds and in what state.
National activity figures
Genomic Medicine Sweden publishes annual inventories of NGS-based analyses performed in Sweden. These describe national clinical activity, not NGP holdings, and the difference is currently large. Use them for context, not for cohort estimates.
Who controls the data
This is the part that determines how much work an application takes, so it is worth being precise about.
Each region is responsible and controls its data, its content and what is extracted from it. Västra Götalandsregionen operates the platform and hosts the servers, but operating it confers no access to other regions' information; VGR acts as a data processor under agreement with each region. Only the region itself, and the users that region authorises, can reach that region's data.
For research, the mechanism works like this:
A region decides that specific data may be disclosed to a specific approved project.
The region applies a metadata tag to that data. The tag is applied by the region where the data owner sits, and by no one else.
Tagged data becomes visible in a shared area of the platform known as GMC Joint, which is reachable only by the users belonging to that project.
Data shared this way is treated in law as a disclosure under the Public Access to Information and Secrecy Act (offentlighets- och sekretesslagen).
The practical consequence: ethical approval alone does not give you data.
Every region holding data you need makes its own decision, and a multi-region study means multiple parallel decisions.
Two approvals in series. National ethical approval is granted once. A separate disclosure decision is then needed from every region holding data you want, and only that region can tag its data for your project.
Three areas already have national ethical approvals and joint controllership agreements in place: rare disease, haematology and childhood cancer. Further approvals are in progress, for example for microbiology. If a study falls inside one of the established areas, the path is usually shorter, because the framework agreements already exist.
Can you use NGP data for your project?
Work through these before you invest in an application.
You have a Swedish research principal (forskningshuvudman). Ethical approval is granted to a research principal, not to an individual researcher.
Your research question needs clinically generated genomic data. If a research cohort or an existing open dataset would answer it, that route is faster. See Swedish research cohorts on this portal.
You can name the regions. Because disclosure is regional, you need to know which Genomic Medicine Centres generated the data you want. A study restricted to one or two regions is significantly faster than a national one.
Your ethical application describes the platform. The application must describe where data will be processed and by whom. Naming the NGP, its processing environment, and the regions involved at the outset avoids an amendment later.
You are not seeking access for commercial product development. This page covers the academic and clinical research route. Industry collaborations run through a separate process not covered on this page.
How to apply for access, step by step
Three phases: prepare, obtain permission, obtain data. Steps within a phase can often run in parallel; the phases cannot.
Phase 1. Prepare
Step 1. Scope the cohort and confirm it exists.
Define the diagnostic area, assay type, time period and regions. Then check feasibility before committing. The NGP feasibility portal returns aggregate counts by diagnosis and molecular criterion per region and discloses no confidential information. For anything the portal cannot answer, contact the Genomic Medicine Centre in the relevant region, or Genomic Medicine Sweden.
Step 2. Confirm your research principal and legal basis.
Establish which organisation is the research principal, and settle the GDPR basis for processing. Your research support office or data protection officer should be involved from this point, not later.
Describe explicitly: the data categories you need, whether you require directly identifying or pseudonymised data, the regions you will approach, where processing will happen, and who will have access. Changes later require an amendment application, with a decision within 35 days.
Step 4. Apply for biobank approval, if you need samples.
Step 5. Request disclosure from each holding region.
After ethical approval, and where samples are involved biobank approval, approach each region holding data you need, through its Genomic Medicine Centre. Each region carries out a harm assessment (menprövning) under chapter 25 of offentlighets- och sekretesslagen and decides independently whether it may release the data. Supply the ethical approval and the data specification.
Step 6. Put the agreements in place.
Rare disease, haematology and childhood cancer already have national frameworks, which shortens this step considerably.
Phase 3. Obtain data
Step 7. The region tags the data.
Once a region has approved disclosure, it applies the metadata tag that makes the specified data visible in GMC Joint to the users named in your project. Nothing moves until this happens, and only the holding region can do it.
Step 8. Get accounts and access the data.
Project users authenticate to the NGP Research Portal using SWAMID. From there you can reach the tagged data. Analysis can run inside the platform on NGPcompute, reached through Open OnDemand, or programmatically through the open-source NGPIris client. Export outside the platform is governed by the terms of your approvals and agreements, not by technical capability.
Analysing data inside the platform is normally simpler than exporting it. Whole genome data is large, and every export creates a new set of security and agreement obligations for your own organisation. Plan for in-platform analysis unless you have a specific reason not to.
Working with the data once you have access
Where analysis happens. NGPcompute provides elastic cloud compute alongside the storage layer, with encryption keys held by a dedicated hardware security module owned by the platform. Users reach it through Open OnDemand, which offers both a graphical interface and a command line.
Programmatic access.NGPIris is the open-source Python client for the platform, providing upload, download, listing and search against both the storage and the index layers. It is available on PyPI and on GitHub.
Analysis pipelines. Genomic Medicine Sweden develops and publishes its diagnostic pipelines as open source, including nf-core/raredisease, Tomte and nallo for rare disease, Twist Solid and the GMS560 panel workflow for solid tumours, and JASEN and gms_16S for microbiology. Using the same pipeline as the originating laboratory makes your results directly comparable to the clinical result.
Metadata standards. The indexing layer can express extracted metadata as HL7 FHIR, openEHR, SNOMED CT and DCAT-AP.
Data you may need alongside NGP
Genomic data rarely answers a clinical research question on its own. These are separate applications with separate timelines, and they are worth starting in parallel rather than in sequence.
National health registers from Socialstyrelsen, for diagnoses, prescribed drugs, procedures and causes of death. Expect several months and a fee running from tens of thousands of SEK.
Register data from Statistics Sweden, for socioeconomic and demographic variables, delivered into the MONA remote analysis environment.
National quality registers, including the cancer quality registers held on the INCA platform by Regionala cancercentrum. A national genomics module linking NGP data into the cancer registers and the individual patient overview (IPÖ) is under development.
Biobank samples through Biobank Sverige, where your study needs material rather than data.
Existing research cohorts, which may already hold what you need with a shorter access path. See Swedish research cohorts on this portal.
Contact
For questions about NGP data, access or the platform itself, contact Genomic Medicine Sweden at info@genomicmedicine.se, or raise a request through the NGP service desk.
For questions about the Precision Medicine Portal, see our contact page.