Sana Farooq
@sanafarooq1
Data Engineer and Data Scientist building scalable ETL pipelines, distributed systems, and predictive models for enterprise impact.
What I'm looking for
I’m a results-driven Data Engineer and Data Scientist with 5+ years designing scalable ETL pipelines, distributed data systems, and predictive models across healthcare, fintech, and government domains. At Punjab Information Technology Board (PITB), I designed and implemented Trino- and Airflow-based ETL pipelines for OSSP, integrating data from 10+ major government departments and processing terabyte-scale datasets. I also deliver dimensional data models for consistent reporting, optimize Big Data workflows, and implement governance with OpenMetadata to ensure data integrity, lineage, and metadata management.
I build analytical and machine learning solutions that support data-driven decision-making, from feature engineering and predictive modeling to large-scale data preprocessing and validation. I bring additional production experience from Data Scientist work at Algo and freelance Data Engineering for healthcare and fintech clients—delivering reliable workflow automation, optimized pipelines, and dashboards using Apache Superset. I’m Microsoft Certified (Azure Data Engineer Associate DP-700), and I’m passionate about enabling data-driven decisions at scale through robust, future-proof cloud data platforms.
Experience
Work history, roles, and key accomplishments
Data Engineer / Data Scientist
Punjab Information Technology Board (PITB)
Jun 2024 - Present (2 years 2 months)
Designed and implemented scalable ETL pipelines for OSSP using Trino and Apache Airflow to integrate data from 10+ government departments. Built ClickHouse analytics, star-schema dimensional models, predictive machine learning models, and OpenMetadata-based data governance for lineage and metadata management.
Data Scientist (Python)
Algo
Jan 2024 - Apr 2024 (3 months)
Performed exploratory data analysis and built predictive machine learning models to improve operational decision-making. Implemented data preprocessing, feature engineering, and validation pipelines and collaborated to integrate solutions into production workflows.
Designed and deployed end-to-end ETL pipelines for a healthcare client, and implemented Apache Airflow for workflow orchestration and automation. Built Apache Superset dashboards for KPI visualization and provided data pipeline optimization for a fintech client.
Data Engineer (Python)
PrivySol
Jan 2020 - May 2022 (2 years 4 months)
Built and maintained ETL processes to improve data flow and integrity for healthcare datasets supporting patient record accuracy and regulatory reporting. Developed machine learning models for patient readmission prediction and resource allocation optimization while managing scalable Big Data pipelines.
Freelance Web Developer
Upwork & Fiverr
Feb 2016 - Dec 2018 (2 years 10 months)
Developed responsive websites for international clients using HTML, CSS, JavaScript, React, and Node.js. Built e-commerce platforms with SEO optimization and integrated payment gateways.
Education
Degrees, certifications, and relevant coursework
FAST-NUCES
Master of Science in Data Science, Data Science
2023 - 2025
Master of Science in Data Science at FAST-NUCES in Lahore from 2023 to 2025 (expected).
University of Management and Technology (UMT)
Bachelor of Science in Software Engineering, Software Engineering
2016 - 2020
Bachelor of Science in Software Engineering at University of Management and Technology (UMT) in Lahore from 2016 to 2020.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Sana?
You can contact Sana and 90k+ other talented remote workers on Himalayas.
Message SanaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
