A pipeline for evaluating context-aware privacy risk in synthetic healthcare narratives

Radwan, Dalya, Wall, Julie ORCID logoORCID: https://orcid.org/0000-0001-6714-4867 and Zolgharni, Massoud ORCID logoORCID: https://orcid.org/0000-0003-0904-2904 (2026) A pipeline for evaluating context-aware privacy risk in synthetic healthcare narratives. In: International Conference on AI in Healthcare, 26-28 Aug 2026, London. (In Press)

[thumbnail of Will be made OA based on IRR on publication] PDF (Will be made OA based on IRR on publication)
A Pipeline for Evaluating Context-Aware Privacy Risk in Synthetic Healthcare Narratives - Dalya Radwan - 1.0 Final.pdf - Accepted Version
Restricted to Repository staff only
Available under License Creative Commons Attribution.

Download (446kB) | Request a copy
Official URL: https://aiih.cc/

Abstract

Healthcare providers increasingly rely on synthetic data when privacy and governance requirements restrict access to real-world data, particularly in large-scale healthcare systems. While prior research has focused on structured data, support for unstructured healthcare narratives remains limited, especially where identifiable information is conveyed indirectly through context, events, and relationships. This paper presents a privacy-aware synthetic text pipeline for experimentation on unstructured narratives. The pipeline uses anonymised real-world text as a base and injects synthetic identifiers through a prompt-based template, enabling controlled manipulation of identity-related content. We demonstrate the pipeline on a large patient-feedback dataset from the NHS and evaluate it through metrics of semantic preservation, fluency, data utility, and privacy-risk indicators. The results suggest that context-aware exposure is influenced not only by identifier quantity but also by contextual integration. The paper contributes: (1) a technological artefact for controllable narrative synthesis with fabricated identifiers, (2) an evaluation approach linking utility with context-aware privacy-risk indicators, and (3) a practical framework for privacy-oriented testing in access-restricted healthcare settings without real identities. The proposed context-aware exposure metric is intended as an ex-ploratory indicator of contextual disclosure potential rather than a validated measure of re-identification risk.

Item Type: Conference or Workshop Item (Paper)
Keywords: synthetic healthcare data; unstructured text; large language models; privacy-risk assessment; narrative data.
Subjects: Computing > Information security
Computing
Medicine and health
Date Deposited: 17 Jul 2026
Dates:
Date
Publication status
22 May 2026
Accepted
School, department or research centre: CAINT (The Centre for AI and Natural Language Technologies)
School of Computing and Engineering
Keywords: synthetic healthcare data; unstructured text; large language models; privacy-risk assessment; narrative data.
URI: https://repository.uwl.ac.uk/id/eprint/15240
Sustainable Development Goals: Goal 9: Industry, Innovation, and Infrastructure

Actions (admin access)

View Item

Menu