The MII core data set is divided into modules:
The core dataset (CDS) module Person contains basic information about a person in the context of the healthcare system. The specification defines the data models that allow the recording of information on patients or probands within the MII using the HL7-FHIR IT standard. This includes attributes for identification, i.e. unique IDs such as health insurance numbers, organisation-internal patient identifiers from the local identity management of the sites or identifiers for test subjects within studies. Furthermore, demographic information such as name, date of birth, gender and address are included. These attributes can be used for external record linkage and to identify and link patient data from different sources. This can be relevant for personalised, cross-institutional analyses. Demographic information enables descriptive analyses by region, gender and age. In addition, the CDS-module “Person” contains information on vital status and, if applicable, the cause of death coded with ICD-10. The contents of other core data set modules refer to the Person module.
The Case module records the contacts of patients in healthcare facilities. Contacts include information on patient visits along the entire treatment process, such as inpatient stays, outpatient visits and virtual, telemedicine contacts. Inpatient contacts can be assigned to a hospital, within a hospital to corresponding wards and departments, and below that to an operating theatre (OR) or specific functional departments for carrying out diagnostic or therapeutic measures. As part of this, the specialist discipline or department can also be documented using the specialist department code of the German Hospital Federation (DKG). The Case module also records the times and periods in which a contact took place, who was involved, which main or secondary diagnoses were relevant during a contact and which procedures were carried out. The Case module refers to the modules Person, Diagnosis and Procedure of the core dataset in order to link the clinical data of those core dataset modules to the treatment case.
The basis for the provision of patient data in the MII for secondary data use is the MII Broad Consent. This is a standardised informed consent form for the scientific reuse of patient data collected in the context of medical care. The Broad Consent was developed by the MII's Consent Working Group and agreed with the responsible data protection officers.
This technical implementation of the MII Broad Consent is based on the generic HL7 specifications and has been adapted to the requirements of the MII.
The Diagnosis module describes the diagnoses of patients in the healthcare system and the characteristics of these diagnoses. The use of various terminology systems for coding diagnosis information is supported. This includes the official classification for coding diagnoses in Germany ICD-10-GM as well as the use of the comprehensive healthcare terminology SNOMED CT for semantic interoperability in the electronic exchange of diagnostic data. Furthermore, the module contains specifications for the coding of rare diseases using Orpha codes and the coding of tumour diseases using the International Classification of Diseases for Oncology ICD-O-3. The terminology systems for diagnoses are complemented by extensive data elements for the contextualisation of a diagnosis, such as the determination of the affected body part of a disease, the times and periods of occurrence or subsidence of the disease-specific symptoms as well as the determination and documentation date of a diagnosis. A diagnosis is assigned to a patient via a reference to the Person module.
The Procedure module describes treatments and measures that patients have undergone during healthcare. The operation and procedure code (OPS) is used to code operations and medical procedures in inpatient care, outpatient surgery and selected drug therapies. The procedure codes from the comprehensive SNOMED CT terminology can also be used for coding. In addition, the module contains data elements to map structured information on anatomical target body sites of procedures, to record the intention to perform and also to record the date of performance and documentation. The connection with the question of who a procedure was performed on is realised by reference to the Person module.
Laboratory tests are performed for almost all hospital in-patients. The test results are usually stored centrally. The following items of information for each patient and case should be transferred to the corresponding data integration centre (DIC) each time a laboratory test is conducted:
- The test/analysis designation, with unique test identifier (using the LOINC standard coding system)
- The date the test was conducted
- The result (measured value) with standardised unit of measurement
- Interpretation: indication whether result represents a pathological value (recommended)
- Type of scale (optional)
- Reference range (recommended)
- Source laboratory (optional)
It is also possible to model other clinical tests, microbiological tests and vital signs that go beyond the scope of conventional clinical/chemical and haematological laboratory data.
The Medication module enables the documentation of medication orders and administration as well as medication plans. Prescribing and administering pharmaceuticals are core processes of routine healthcare and take place at all MII sites. However, the proportion of digitally documented prescriptions and administrations varies between the sites in terms of the degree of structuring, the populations and medications covered. Medication data is of central importance for a variety of issues, for example in pharmacovigilance or as an inclusion/exclusion criterion for study collectives.
The following types of medication documentation can be distinguished:
- Medication in hospital (mainly inpatient/partial inpatient)
- Admission and discharge medication
- Outpatient medication
- Self-medication (OTC)
- Medication in the context of clinical studies
- Medication documentation for the nationally standardised medication plan
Medication details can range from the simple documentation of the administration of a medicinal product to the detailed, structured recording of individual doses with coding of the active ingredient, pharmaceutical form, route of administration and dose in accordance with internationally established standards.
The Rare Diseases module covers the documentation of the diagnostic process for rare diseases for the estimated 4 million people affected in Germany. The module enables the structured recording of clinical and genetic diagnoses, phenotyping, family history, interdisciplinary examinations at specialized centers, and treatment recommendations. It builds on existing data specifications such as ERDRI CDS (The European Rare Disease Registry Infrastructure Common Data Set), NARSE (National Registry for Rare Diseases), and MVGenomSeq (Model Project for Genome Sequencing), for which it also provides mappings of the individual data elements.
The module deliberately makes use of pre-modeled elements from other FHIR profiles in the MII core dataset (e.g., the Molecular Genetic Report module, the Diagnosis module, etc.).
The PROs, PROMs and Derived Metrics module standardises the collection and analysis of patient-reported health data (Patient-Reported Outcomes) through FHIR-based specifications. It contains guidelines on questionnaire design with regard to various areas of application (presentation, collection, calculation, and conversion of questionnaire content into other FHIR resources). In addition, frequently used, validated questionnaires such as the PHQ-9, PROMIS-29, EQ-5D-5L and EORTC QLQ-C30 are made centrally available for use in data collection or as a common harmonisation mapping, and strategies for instrument-independent secondary data use are explained.
The Molecular Tumour Board (MTB) module enables all data used in the context of personalised oncology within molecular tumour boards to be mapped in a standardised manner. Patients undergo a complex patient journey, meaning that the dataset not only takes into account diagnostic and therapeutic measures at different points in time, but also the respective delivery of these by inpatient and outpatient care providers.
The MTB module is intended to complement the existing Oncology, Molecular Genetics Report and Pathology Report modules, and relies on cross-references to minimise redundancy as far as possible. In particular, it provides additional information from the fields of clinical and (molecular) pathological data, such as complex biomarkers.
In this way, the MTB module complements the existing MII core data set modules for the ‘Personalised Oncology Application Profile’, which was initiated as part of PM4Onco. This is based on version 2.0 of the German Network for Personalised Medicine (DNPM) dataset. In addition, a comparison was carried out with version 1.2.2 of the dataset from the National Network for Genomic Medicine in Lung Cancer (nNGM).
The Document module enables metadata relating to clinical documents of any kind – including not only text but also images and videos – to be structured and recorded for any purpose. Using this profile facilitates both the internal and external use of documents. Document types are clearly categorized. Document relationships, document status, document discoverability, corpus navigation and document archiving are coordinated according to a standardised scheme. The module also enables the creation of document references linked to the Case and Person modules. Using an NLP (Natural Language Processing) extension, the processing status can be mapped in relation to NLP procedures, such as de-identification or annotations. The module can be used seamlessly in conjunction with the Gematik/ISiK and KBV/MIO standards.
The Oncology module fully replicates the Common Basic Oncology Dataset (oBDS) in its 2021 version, including the four organ-specific modules, and enables this data to be used in a standardised manner within the data integration centres. The module includes a detailed description of the primary diagnosis and histology, organ-specific data points such as the PSA level and biopsy information for prostate cancer or receptor status for breast cancer, details of treatments administered, including complications and side effects, as well as cancer-specific parameters such as the TNM classification, progression staging and tumour board recommendations. In addition to the oBDS, the module provides further categorisation, for example regarding medication or the ‘Further Classification’ profile, which has been supplemented with commonly used staging and grading systems such as ELN, IPI or FIGO.
The Imaging module describes the general structure of radiological findings and provides a data model for the analysis and mapping of imaging data, as well as report data for all common radiological modalities. It incorporates information from DICOM headers and maps radiological reports in accordance with the DIN 25300-1 standard of the German Radiological Society (DRG). The structured capture of this data supports the further development of diagnostic methods and promotes patient-centred care. The core data set module is flexible and integrates both unstructured free-text reports and semi- or structured reports, thereby supporting both historical and modern report formats. Particular attention is paid to the traceability of clinically relevant entities such as tumour diseases. In addition, modality-specific attributes are linked to the report descriptions to enable deeper technical insights and patient selection for downstream analyses.
The Microbiology module enables the documentation of the presence of microorganisms in isolates and samples using various diagnostic methods: culture, microscopy, molecular diagnostics, serology (part of immunology) and immunology. In addition, tests that can further characterise the properties of the microorganisms, such as the resistance mechanism or virulence factor, are also described. Phenotypic susceptibility tests including all methods are included as well as predictive susceptibility based on molecular diagnostics. The MRGN classification (multi-resistant gram-negative bacteria) of the Robert Koch Institute can also be represented with this model.
The Medical Research Projects module describes characteristics of studies and other medical research projects, i.e. for the investigation of experimental clinical and epidemiological hypotheses through structured data collection and processing, mostly of human subjects. This version of the module supports the recording of attributes for identifying and managing a research project as well as for recording the basic characterising attributes (study register). In addition, structured inclusion and exclusion criteria can be defined, which can be used to decide, at least semi-automatically, whether an individual with their intrinsic characteristics belongs to the target population or not (feasibility).
The Pathology Findings module describes the general structure of a structured pathology report based on the IHE PaLM (Pathology and Laboratory Medicine) profile APSR 2.0 (Anatomic Pathology Structured Report). These pathology reports are mostly available in text form. The results of clinically requested examinations are collected and documented with a synoptic (summarising) evaluation, such as histological and cytomorphological (tissue and cell morphological) and molecular examinations. The freely formulated examination results can also be supplemented (semantically annotated) by structured coding. It is important to note that each structured code must also be readable as text, but not all text information must be coded. Furthermore, information on the sample, the examination order, image data and the reason for the examination can also be recorded.
The Molecular Genetic Report module specifies the reporting of variants detected by sequence analysis that are present in a patient's sample material. Genetic tests provide information on causal relationships between structural variants or changes in the genome and potential diseases as well as possible therapies. In addition to the description of the sequence variants and region(s) analysed, the contents of the report may include request information such as the patient's and family members' medical histories and a reference to previously performed tests. The report may also contain information on billing codes and therapeutic and diagnostic implications derived from the variant information.
The Intensive Care Medicine module specifies ICU data for primary and secondary use. The special nature of this data is not only in the severity of the patient's illness, but also in the fine-grained data collection in special documentation systems and the comparatively high density of fully and partially structured data. Furthermore, intensive care data is of great importance in the context of local and national pandemic management and pandemic-related research.
The Biosample Data module describes both superordinate collections/biobanks and individual samples, including information on their collection, composition, processing and storage. Biosamples are collected, processed and stored in both clinical and population-based biobanks in order to provide high-quality samples for scientific projects. The different collections in a biobank as well as the individual samples must be described in a structured manner to facilitate the retrieval of samples and their appropriate use. Sample-specific data should include information on sample type, quantity, collection, pre-analytical processing and storage. Clinical data on the sample is explicitly not covered by this module, but should be provided via the modules intended for the respective data type.
Contact:
Do you have any questions or suggestions regarding the further development of the core data set or individual core data set modules? Would you like to get involved in the work on the core data set?
Please contact us by email at office@medizininformatik-initiative.de (keyword: core data set).