Skip to main content
Natalia GarcíaNG
Open to opportunities

Natalia García

@nataliagarca

Data engineer and data scientist building ETL, NLP, and MLOps pipelines for clinical decision support.

Spain
Message

What I'm looking for

I’m looking for a team building production-grade medical data and AI, where I can own scalable ETL, deploy ML models with strong observability (MLflow/EvidentlyAI), and collaborate with researchers to turn clinical data into reliable decisions.

I build production-grade ETL and data pipelines for healthcare and hybrid cloud environments. I designed and maintained large-scale pipelines with Spark, Hive, Iceberg, Hadoop, and CI/CD to keep clinical data flowing reliably.

I specialize in applied NLP and end-to-end medical analytics—developing extraction and normalization for medical text, entity extraction, knowledge graph generation, and clinical text and image anonymization, summarization, and RAG. I also work with normalized EHR and genomics datasets to train and deploy machine learning (XGBoost, Logistic Regression, SVC, Random Forest) and deep learning models (LSTM, GRU) that support faster medical decision-making.

I bring a strong MLOps mindset to reliability and monitoring, implementing model logging and data drift tracking with MLflow and EvidentlyAI as part of an observability-first framework. Earlier, as a research assistant, I developed bioinformatics pipelines, automated literature mining workflows, and built computer vision tools for cell classification—reducing manual workload by 30%—while publishing protein sequence analysis work for drug repurposing.

Experience

Work history, roles, and key accomplishments

CC
Current

Data Engineer / Data Scientist

CGM Clinical

Aug 2024 - Present (1 year 11 months)

Designed and maintained large-scale ETL pipelines using Spark, Hive, Iceberg, Hadoop, and CI/CD in hybrid cloud environments. Built NLP and ML/DL workflows for medical text extraction/normalization, knowledge graph generation, anonymization and RAG, and trained models for EHR/genomics-driven medical decision support with MLflow/EvidentlyAI observability and MLOps.

SC

Research Assistant Bioinformatician

Spanish Research Council (CSIC)

Sep 2023 - Jul 2024 (10 months)

Developed bioinformatics pipelines for gene expression and epigenomic sequencing data analysis. Automated scientific literature mining workflows to support hypothesis generation for drug discovery.

CT

Biomedical Research Assistant

CTB-UPM

Jan 2022 - Jul 2022 (6 months)

Designed and implemented integration of new biomedical protein and gene sequence features in a project database. Developed semi-automated extraction workflows from NCBI public databases and used bioinformatics alignment and sequence embedding similarity methods to support drug target sequence-based drug repurposing hypotheses.

Education

Degrees, certifications, and relevant coursework

Universidad Politécnica de Madrid logoUM

Universidad Politécnica de Madrid

Master of Science, Computational Biology

2022 - 2023

Grade: GPA: 3.78

MSc in Computational Biology (Data Science specialization) at Universidad Politécnica de Madrid, including protein sequence analysis in the context of drug repurposing.

Universidad Politécnica de Madrid logoUM

Universidad Politécnica de Madrid

Bachelor of Science, Biotechnology

2018 - 2022

BSc in Biotechnology (Bioinformatics specialization) at Universidad Politécnica de Madrid.

Tech stack

Software and tools used professionally

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan