Radwan, Dalya, Wall, Julie ORCID: https://orcid.org/0000-0001-6714-4867 and Zolgharni, Massoud
ORCID: https://orcid.org/0000-0003-0904-2904
(2026)
A pipeline for evaluating context-aware privacy risk in
synthetic healthcare narratives.
In: International Conference on AI in Healthcare, 26-28 Aug 2026, London.
(In Press)
|
PDF (Will be made OA based on IRR on publication)
A Pipeline for Evaluating Context-Aware Privacy Risk in Synthetic Healthcare Narratives - Dalya Radwan - 1.0 Final.pdf - Accepted Version Restricted to Repository staff only Available under License Creative Commons Attribution. Download (446kB) | Request a copy |
Abstract
Healthcare providers increasingly rely on synthetic data when privacy and governance requirements restrict access to real-world data, particularly in large-scale healthcare systems. While prior research has focused on structured data, support for unstructured healthcare narratives remains limited, especially where identifiable information is conveyed indirectly through context, events, and relationships. This paper presents a privacy-aware synthetic text pipeline for experimentation on unstructured narratives. The pipeline uses anonymised real-world text as a base and injects synthetic identifiers through a prompt-based template, enabling controlled manipulation of identity-related content. We demonstrate the pipeline on a large patient-feedback dataset from the NHS and evaluate it through metrics of semantic preservation, fluency, data utility, and privacy-risk indicators. The results suggest that context-aware exposure is influenced not only by identifier quantity but also by contextual integration. The paper contributes: (1) a technological artefact for controllable narrative synthesis with fabricated identifiers, (2) an evaluation approach linking utility with context-aware privacy-risk indicators, and (3) a practical framework for privacy-oriented testing in access-restricted healthcare settings without real identities. The proposed context-aware exposure metric is intended as an ex-ploratory indicator of contextual disclosure potential rather than a validated measure of re-identification risk.
| Item Type: | Conference or Workshop Item (Paper) |
|---|---|
| Keywords: | synthetic healthcare data; unstructured text; large language models; privacy-risk assessment; narrative data. |
| Subjects: | Computing > Information security Computing Medicine and health |
| Date Deposited: | 17 Jul 2026 |
| Dates: | Date Publication status 22 May 2026 Accepted |
| School, department or research centre: | CAINT (The Centre for AI and Natural Language Technologies) School of Computing and Engineering |
| Keywords: | synthetic healthcare data; unstructured text; large language models; privacy-risk assessment; narrative data. |
| URI: | https://repository.uwl.ac.uk/id/eprint/15240 | Sustainable Development Goals: | Goal 9: Industry, Innovation, and Infrastructure |
Actions (admin access)
![]() |
Lists
Lists