<< Back

Clinical Data Engineering Pipeline for Enhanced Quality, Normalization, and Analytical Usability (#1224)

Read Article

Date of Conference

July 15-17, 2026

Published In

"Engineering without Borders: Artificial Intelligence, Knowledge, Innovation, and Alliances for a Future from the Americas"

Location of Conference

Santiago (Chile)

Authors

Zablah, Isaac

Hernandez, Edwin

Garcia, Fiama

Zuniga, Antonieta

Garcia Loureiro, Antonio

Abstract

Healthcare data fragmentation and heterogeneity pose significant challenges for reliable clinical analytics and artificial intelligence applications. This study presents a systematic data engineering pipeline designed to improve data quality, semantic normalization, and analytical readiness in clinical environments. The pipeline integrates data validation, cleansing, standardization using HL7 FHIR standards, and quality assessment modules. We evaluated the pipeline using a simulated heterogeneous clinical dataset comprising 50,000 patient records with intentionally introduced quality defects representing real-world data challenges. Quality metrics including completeness, consistency, validity, and semantic coherence were measured before and after pipeline processing. Results demonstrated substantial improvements: completeness increased from 73.2% to 98.5% (p<0.001), data consistency improved from 68.7% to 96.3% (p<0.001), duplicate records were reduced from 8.3% to 0.2%, and semantic standardization reached 97.8% conformance with FHIR resources. The pipeline successfully transformed fragmented clinical data into analysis-ready datasets suitable for advanced analytics and machine learning applications. These findings suggest that systematic data engineering approaches can significantly enhance the reliability and interoperability of clinical information systems, thereby supporting evidence-based decision-making and precision medicine initiatives.

Read Article