Joy Miao
@joymiao
Senior Data Engineer building reliable cloud data platforms and analytics-ready products.
What I'm looking for
I’m a hands-on Senior Data Engineer with 6+ years designing, building, and operating production-grade cloud data platforms. I focus on distributed batch and streaming pipelines that deliver analytics-ready data products. My strength is turning complex source data into governed, reusable datasets that scale across teams.
At Bell Mobility, I architected a reusable, configuration-driven Source-to-Raw ingestion framework on GCP, standardizing production ingestion from RDBMS and Salesforce Bulk API sources using Dataflow and Serverless Dataproc/PySpark. I also delivered a reusable GCS-to-BigQuery Raw-to-Standardized framework that standardized column names, data types, null handling, and timestamps, reducing tenant onboarding from multiple weeks to approximately one week.
I led a versioned data contract capability used by 10+ tenant teams, with CI-based static validation of tenant-authored YAML contracts, Cloud Functions registration in Firestore, and runtime enforcement within the Raw-to-Standardized pipeline. For high-volume Salesforce and RDBMS sources, I evaluated and delivered CDC patterns covering incremental ingestion, watermark management, deduplication, recovery, and full-reload paths—balancing latency, scalability, cloud cost, and operational resilience. I raised team engineering standards through mentoring and technical design/code reviews across Spark and SQL performance, reusable component design, failure handling, and production readiness.
Earlier, I migrated 5+ years of telecom data from Hadoop and Amazon S3-based environments to GCS and BigQuery, automating backfills supporting approximately 2 TB/day of raw CDR binary data and 3.5 TB/day of decoded JSON, and reducing GCS storage costs by 89%. I designed and operated event-driven streaming pipelines with Spark Structured Streaming and Confluent Kafka for near-real-time telecom telemetry, plus a Dataplex-based metadata and governance solution that reduced configuration by approximately two-thirds. I also built centralized observability with Logstash/Elasticsearch/Kibana and delivered Terraform changes through CI/CD workflows to keep production releases dependable.
Experience
Work history, roles, and key accomplishments
Data Engineer II / Cloud Data Engineer
Bell Mobility
Jan 2025 - May 2026 (1 year 4 months)
Architected reusable, configuration-driven Source-to-Raw and Raw-to-Standardized ingestion frameworks on GCP, producing analytics-ready datasets and standardizing formats. Built a versioned data contract capability for 10+ tenant teams and led CDC pattern delivery for high-volume Salesforce and RDBMS sources.
Data Engineer / Data Scientist
Bell Mobility
Jan 2020 - Jan 2025 (5 years)
Migrated telecom datasets from Hadoop and S3 to GCS and BigQuery, operating large-scale backfills and improving cost efficiency. Designed and ran streaming pipelines (Spark + Confluent Kafka), enterprise analytics data products, metadata governance (Dataplex), and observability/alerting using ELK; also delivered Terraform changes via CI/CD and built Kubeflow-based model-training workflows.
Education
Degrees, certifications, and relevant coursework
University of Ottawa
Master of Applied Science (MASc), Electrical & Computer Engineering
2016 - 2019
Completed a Master of Applied Science (MASc) in Electrical & Computer Engineering from 2016 to 2019.
Southwest Jiaotong University
Bachelor of Engineering, Electrical Engineering and Automation
2012 - 2016
Completed a Bachelor of Engineering in Electrical Engineering and Automation from 2012 to 2016.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Joy?
You can contact Joy and 90k+ other talented remote workers on Himalayas.
Message JoyGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
