The use of saturated count models for synthesis of large confidential administrative databases

Jackson, James and Mitra, Robin and Francis, Brian and Dove, Iain (2022) The use of saturated count models for synthesis of large confidential administrative databases. PhD thesis, Lancaster University.

[thumbnail of 2022jacksonphd]
Text (2022jacksonphd)
2022jacksonphd.pdf - Published Version

Download (1MB)

Abstract

Synthetic data sets are being increasingly used to protect data confidentiality. In the three decades since they were first introduced, methods for synthetic data generation have evolved, but mainly within the domain of survey data sets. As greater interest is being taken in utilising administrative data for statistical purposes, there is inevitably greater interest in creating synthetic administrative databases. Yet there are characteristics of these databases that require special attention from a synthesis perspective, such as their size and the presence of structural zeros. This thesis, through the fitting of saturated models in conjunction with overdispersed count distributions, presents a mechanism that allows large administrative databases to be synthesized efficiently. This thesis also proposes a concept of satisfying risk and utility metrics a priori - that is, prior to synthetic data generation - using the synthesis mechanism’s tuning parameters, allowing a more formalized approach to synthesis. The methods are demonstrated empirically throughout, primarily through synthesizing a database that can be viewed as a close substitute to the English School Census.

Item Type:
Thesis (PhD)
Subjects:
?? synthetic datastatistical disclosure controlcount distributionstabular data ??
ID Code:
181728
Deposited By:
Deposited On:
19 Dec 2022 12:50
Refereed?:
No
Published?:
Published
Last Modified:
17 Feb 2024 00:24