peek User
@shawnpeek
Senior Data Engineer building HIPAA-aware batch/streaming pipelines and ML feature systems.
What I'm looking for
I’m a data engineer with 8 years building batch and streaming pipelines—11 years total in data—across energy, healthcare, and analytics SaaS. I’m deep on AWS (Glue, Lambda, Step Functions, Kinesis/MSK), Airflow, and dbt, integrating multi-source telemetry and operational data.
At Copay Health AI, I unified FHIR R4 EHR, practice-management, and payment-gateway data into a single HIPAA-compliant pipeline with strict PHI/payment isolation. I built the feature pipeline and retraining loop behind ML scoring, scoring ~25K patient accounts nightly and holding the collectability model at AUC ~0.81 through payer-mix shift while lifting recovered payments ~15% over a rules-based baseline.
I also modernized performance-critical workloads by redesigning a 30-minute batch scoring job into a streaming consumer on AWS MSK and Kinesis—cutting scoring-path lag to under 2 minutes on live payment and claims events. For “agentic billing,” I grounded Navi on curated rules and policy data using pgvector for RAG.
Before that, at IDARE, I engineered the digital-twin data pipeline for energy assets—connecting, transforming, and loading multi-source sensor/SCADA and asset-registry data into one queryable model at ~2M sensor readings/day. I’m Python-first and reliability-obsessed, with the security and compliance mindset needed for HIPAA- and PCI-aware systems.
Experience
Work history, roles, and key accomplishments
Senior Data Engineer
Copay Health AI
Apr 2022 - Jun 2026 (4 years 2 months)
Unified FHIR EHR, practice-management, and payment data into a HIPAA-compliant pipeline with PHI/payment isolation, and built ML feature pipelines for nightly patient scoring and drift monitoring. Also delivered batch-to-streaming scoring on AWS (MSK/Kinesis) and landed outputs into Azure services for RBAC-enforced patient services and an agentic billing assistant grounded in curated data.
Data Engineer
Idare
Jul 2018 - Apr 2022 (3 years 9 months)
Built data-pipeline capabilities for IDARE's digital-twin platform by connecting, transforming, and loading multi-source sensor and SCADA data into a queryable model. Owned ingestion and time-series pipelines for predictive-maintenance and 3D asset digital twins, and orchestrated ETL and feature datasets for AutoML analytics and GIS dashboards.
Data Analyst
ActivTrak
Jun 2015 - Jun 2018 (3 years)
Built productivity dashboards, benchmarks, and reports for ActivTrak's workforce-analytics platform from desktop-activity telemetry. Modeled and queried activity event data in SQL to classify application and website usage, and automated recurring reporting by defining productivity-metric definitions.
Education
Degrees, certifications, and relevant coursework
University of Houston
Bachelor of Science, Computer Science
2010 - 2014
Earned a B.S. in Computer Science from the University of Houston from 2010 to 2014.
Tech stack
Software and tools used professionally
Apache Spark
AWS Glue
AWS Step Functions
GitHub
Kubernetes
GitHub Actions
Pandas
dbt
DB
MySQL
PostgreSQL
MongoDB
Databricks
Redis
Terraform
Azure DevOps
scikit-learn
Kafka
FastAPI
AIOHTTP
asyncio
Grafana
Prometheus
SQLAlchemy
Linux
Amazon Kinesis
Azure Functions
pytest
Airflow
navi
SQL
XGBoost
Pydantic
Cosmos
Bash
pgvector
Agentic
Loops
Availability
Location
Authorized to work in
Social media
Job categories
Skills
Interested in hiring peek?
You can contact peek and 90k+ other talented remote workers on Himalayas.
Message peekGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
