Natalia García
@nataliagarca
Data engineer and data scientist building ETL, NLP, and MLOps pipelines for clinical decision support.
What I'm looking for
I build production-grade ETL and data pipelines for healthcare and hybrid cloud environments. I designed and maintained large-scale pipelines with Spark, Hive, Iceberg, Hadoop, and CI/CD to keep clinical data flowing reliably.
I specialize in applied NLP and end-to-end medical analytics—developing extraction and normalization for medical text, entity extraction, knowledge graph generation, and clinical text and image anonymization, summarization, and RAG. I also work with normalized EHR and genomics datasets to train and deploy machine learning (XGBoost, Logistic Regression, SVC, Random Forest) and deep learning models (LSTM, GRU) that support faster medical decision-making.
I bring a strong MLOps mindset to reliability and monitoring, implementing model logging and data drift tracking with MLflow and EvidentlyAI as part of an observability-first framework. Earlier, as a research assistant, I developed bioinformatics pipelines, automated literature mining workflows, and built computer vision tools for cell classification—reducing manual workload by 30%—while publishing protein sequence analysis work for drug repurposing.
Experience
Work history, roles, and key accomplishments
Data Engineer / Data Scientist
CGM Clinical
Aug 2024 - Present (1 year 11 months)
Designed and maintained large-scale ETL pipelines using Spark, Hive, Iceberg, Hadoop, and CI/CD in hybrid cloud environments. Built NLP and ML/DL workflows for medical text extraction/normalization, knowledge graph generation, anonymization and RAG, and trained models for EHR/genomics-driven medical decision support with MLflow/EvidentlyAI observability and MLOps.
Research Assistant Bioinformatician
Spanish Research Council (CSIC)
Sep 2023 - Jul 2024 (10 months)
Developed bioinformatics pipelines for gene expression and epigenomic sequencing data analysis. Automated scientific literature mining workflows to support hypothesis generation for drug discovery.
Biomedical Research Assistant
CTB-UPM
Jan 2022 - Jul 2022 (6 months)
Designed and implemented integration of new biomedical protein and gene sequence features in a project database. Developed semi-automated extraction workflows from NCBI public databases and used bioinformatics alignment and sequence embedding similarity methods to support drug target sequence-based drug repurposing hypotheses.
Education
Degrees, certifications, and relevant coursework
Universidad Politécnica de Madrid
Master of Science, Computational Biology
2022 - 2023
Grade: GPA: 3.78
MSc in Computational Biology (Data Science specialization) at Universidad Politécnica de Madrid, including protein sequence analysis in the context of drug repurposing.
Universidad Politécnica de Madrid
Bachelor of Science, Biotechnology
2018 - 2022
BSc in Biotechnology (Bioinformatics specialization) at Universidad Politécnica de Madrid.
Availability
Location
Authorized to work in
Website
natpod.github.ioPortfolio
github.com/Natpod/SMLJob categories
Interested in hiring Natalia?
You can contact Natalia and 90k+ other talented remote workers on Himalayas.
Message NataliaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
