ISARIC Data Schema

This is a brief guide to the isaricanalytics.isaric_data_schema library, which supports direct transformations of raw, custom clinical datasets into the standardised ISARIC data schema format, which can be analysed and visualised without much effort.

Before proceeding further, it is recommended to consult the ARC ISARIC data schema guide for more information on the core table (one-row-per-patient) and long table (multiple-rows-per-patient) output formats, and the required columns within each table, and also the ARC parser guide on how to write the parser file.

The Schema Transformation Pipeline

The transformation pipeline is implemented as a single function transform_to_isaric_data_schema(), which can produce either the basic ISARIC data schema outputs, or a VERTEX-style outputs, as described below.

Basic Schema Outputs

To produce the basic ISARIC data schema outputs call the transform_to_isaric_data_schema() function with three required arguments:

  • parser_file - the (relative or absolute) file path of the .toml parser file that describes how columns in the source dataset map to columns in the core or long table versions of the ISARIC data schema. If both core and long table outputs are required the parser file will need to define both sets of mappings.

  • data_file - the (relative or absolute) file path of the data file containing the raw data.

  • arc_version - the ARC release version/tag, which can also be obtained from the new ARC Python package using the isaric-arc:arc.arc_core.get_arc_versions() function.

Provided you have cloned the ARC repository side-by-side with your project, then the following snippet shows the transform_to_isaric_data_schema() function at work:

>>> tables = transform_to_isaric_data_schema('../ARC/docs/examples/example_parser.toml', '../ARC/docs/examples/example_data.csv', 'v1.6.1')
[covid-study] parsing example_data.csv: 100%|███████████████████████████████████████████| 5/5 [00:00<00:00, 51.04it/s]
[covid-study] validating core table: 5it [00:00, 52038.51it/s]
[covid-study] validating long table: 109it [00:00, 128819.14it/s]
2026-10-06 16:36:32 [INFO] arc.arc_core: version: v1.6.1

>>> tables['core']
  subjid       siteid  ... adtl_valid                                         adtl_error
0   C001  SITE-GBR-01  ...       True                                                NaN
1   C002  SITE-DEU-01  ...       True                                                NaN
2   C003  SITE-USA-01  ...       True                                                NaN
3   C004  SITE-GBR-02  ...      False  data must contain ['subjid', 'siteid', 'datase...
4   C005  SITE-ESP-01  ...       True                                                NaN

[5 rows x 13 columns]

>>> tables['long']
           date   dataset_id attribute_status subjid  ... adtl_valid attribute_unit value_num duration
0    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
1    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
2    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
3    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
4    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
..          ...          ...              ...    ...  ...        ...            ...       ...      ...
104  2023-01-21  COVID-STUDY              VAL   C005  ...       True         10^9/L       9.2      NaN
105  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
106         NaN  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
107  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
108  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN

[109 rows x 12 columns]

The core table contians individual patient records, that is, with each row representing a distinct patient record, while the long table contains multiple records per patient, with each record representing an observation about a patient.

VERTEX-style Outputs

If you’re a VERTEX user, and interested in writing insight panels for a VERTEX dashboard then with the additional optional argument of as_vertex_data=True the transform_to_isaric_data_schema() function will return a dictionary of the main VERTEX dashboard input file dataframes (df_map, daily, and dictiomary) as shown in the snippet below, that can be used to set up a compatible VERTEX project.

>>> vertex_project_data = transform_to_isaric_data_schema('../ARC/docs/examples/example_parser.toml', '../ARC/docs/examples/example_data.csv', 'v1.6.1', as_vertex_data=True)
[covid-study] parsing example_data.csv: 100%|██████████████████████████████████████████| 5/5 [00:00<00:00, 530.95it/s]
[covid-study] validating core table: 5it [00:00, 36986.81it/s]
[covid-study] validating long table: 109it [00:00, 180774.67it/s]
>>>
>>> vertex_project_data.keys()
dict_keys(['df_map', 'daily', 'dictionary'])
>>>
>>> vertex_project_data['df_map']
  subjid       siteid  ... adtl_valid                                         adtl_error
0   C001  SITE-GBR-01  ...       True                                                NaN
1   C002  SITE-DEU-01  ...       True                                                NaN
2   C003  SITE-USA-01  ...       True                                                NaN
3   C004  SITE-GBR-02  ...      False  data must contain ['subjid', 'siteid', 'datase...
4   C005  SITE-ESP-01  ...       True                                                NaN

[5 rows x 13 columns]
>>>
>>> vertex_project_data['daily']
           date   dataset_id attribute_status subjid  ... adtl_valid attribute_unit value_num duration
0    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
1    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
2    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
3    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
4    2023-01-10  COVID-STUDY              VAL   C001  ...       True            NaN       NaN      NaN
..          ...          ...              ...    ...  ...        ...            ...       ...      ...
104  2023-01-21  COVID-STUDY              VAL   C005  ...       True         10^9/L       9.2      NaN
105  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
106         NaN  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
107  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN
108  2023-01-21  COVID-STUDY              VAL   C005  ...       True            NaN       NaN      NaN

[109 rows x 12 columns]
>>>
>>> vertex_project_data['dictionary']
              Form             Section  ... Branch                                   Question_english
0     presentation                 NaN  ...                   Participant Identification Number (PIN)
1     presentation  INCLUSION CRITERIA  ...         Suspected or confirmed infection, condition, o...
2     presentation  INCLUSION CRITERIA  ...         Is the suspected or confirmed infection, condi...
3     presentation  INCLUSION CRITERIA  ...                                Case classification status
4     presentation  INCLUSION CRITERIA  ...                                        Reason for testing
...            ...                 ...  ...    ...                                                ...
1752    withdrawal          WITHDRAWAL  ...                                        Date of withdrawal
1753    withdrawal          WITHDRAWAL  ...                                     Reason for withdrawal
1754    withdrawal          WITHDRAWAL  ...         Did the participant withdraw from active parti...
1755    withdrawal          WITHDRAWAL  ...         Did the participant withdraw consent to use da...
1756    withdrawal          WITHDRAWAL  ...         Did the participant withdraw consent to use sa...

[1757 rows x 38 columns]

These VERTEX project data files should be exported to CSVs via the pandas.DataFrame.to_csv() method (without the dataframe index, i.e. with index=False) first in order to set up the VERTEX project.