Hybrid Natural Language Processing Framework for Quality Assessment and Normalization of Clinical Data with Human-in-the-Loop Validation (#1045)
Read ArticleDate of Conference
July 15-17, 2026
Published In
"Engineering without Borders: Artificial Intelligence, Knowledge, Innovation, and Alliances for a Future from the Americas"
Location of Conference
Santiago (Chile)
Authors
Madrid, Marcio
Agudelo-Santos, Carlos
Giacaman, Laura
Simons, Perla
Madrid, Melania
Argueta, Edil
Abstract
Data quality is a problem in healthcare information systems, especially when the capture of diagnostic information is not standardized. In this article, a hybrid computational framework for clinical data normalization is designed and evaluated, combining deterministic text processing methods with fuzzy similarity algorithms and a human-in-the-loop validation mechanism. The proposed system was implemented and tested on 2,777 outpatient care records from the Villa Nueva Health Center in Honduras between January and December 2024. The proposed multi-level normalization pipeline reduced the cardinality of unique diagnoses by 55.56% (from 90 to 40 categories) and geographic locations by 30.07% (from 153 to 107 categories), with 99.46% of mappings achieved by exact match (Level 1). Diagnostic entropy decreased from 3.06 to 2.55 bits (16.65%), while the Herfindahl–Hirschman index increased by 20.91%, indicating greater uniformity in the frequency distribution. The ranking stability analysis yielded Jaccard indices of 0.667 and Kendall’s τ coefficients of 0.929 for the top 10 diagnoses. Thirteen ambiguous cases (0.47%) were identified that required manual review, demonstrating the feasibility of the semi-automated method. The temporal drift analysis showed an average vocabulary stability of 43% (Jaccard) between successive months. The proposed architecture is replicable, scalable, and exportable as a programd pipeline, providing a practical solution for data governance in epidemiological surveillance systems.