- Introduction
- Data Coverage
- Further Details
- Request Access
This synthetic dataset includes 16,276 patients admitted for drug overdose from 2016 to 2022, featuring comprehensive patient demographics, comorbidities coded by ICD-10 and SNOMED-CT, and detailed admission data from the index event onward. Information on clinical outcomes, primary diagnoses, psychiatric referrals, and all treatments (e.g., fluids, blood products, procedures) is included.
The dataset was generated using the SDV package’s HMA1 synthesizer. The real data was pre-processed, with metadata defining schema, primary/foreign keys, and inter-table relationships, guiding the synthesizer in learning data structure and dependencies. This approach produced synthetic data that mirrors the original’s statistical properties, supporting privacy-preserving analysis and model training.
Geography: The West Midlands has a population of 6 million & includes a diverse ethnic & socio-economic mix. UHB is one of the largest NHS Trusts in England, providing direct acute services & specialist care across four hospital sites, with 2.2 million patient episodes per year, 2750 beds & > 120 ITU bed capacity. UHB runs a fully electronic healthcare record (EHR) (PICS; Birmingham Systems), a shared primary & secondary care record (Your Care Connected) & a patient portal “My Health”.
Data set availability: Data access is available via the PIONEER Hub for projects which will benefit the public or patients. This can be by developing a new understanding of disease, by providing insights into how to improve care, or by developing new models, tools, treatments, or care processes. Data access can be provided to NHS, academic, commercial, policy and third sector organisations. Applications from SMEs are welcome. There is a single data access process, with public oversight provided by our public review committee, the Data Trust Committee. Contact pioneer@uhb.nhs.uk or visit www.pioneerdatahub.co.uk for more details.
Available supplementary data: Matched controls; ambulance and community data. Unstructured data (images). We can provide the dataset in OMOP and other common data models and can build synthetic data to meet bespoke requirements.
Available supplementary support: Analytics, model build, validation & refinement; A.I. support. Data partner support for ETL (extract, transform & load) processes. Bespoke and “off the shelf” Trusted Research Environment (TRE) build and run. Consultancy with clinical, patient & end-user and purchaser access/ support. Support for regulatory requirements. Cohort discovery. Data-driven trials and “fast screen” services to assess population size.

