Anna Iqbal
@annaiqbal
Senior data engineer architecting multi-cloud lakehouses and low-latency streaming backbones.
What I'm looking for
I’ve built and deployed end-to-end Medallion Lakehouse platforms at Dataquest using Databricks, Delta Engine, Unity Catalog, Delta Live Tables, and PySpark—migrating 600+ TB of legacy data into structured Delta Lake and Apache Iceberg tables.
I run production pipelines across Apache Airflow and Azure Data Factory, orchestrating 150+ daily dynamic DAGs for financial and transactional workloads with a 99.95% operational SLA. I’ve also re-engineered slow SQL and Spark jobs on AWS EMR and Google Cloud Dataflow to cut cloud compute spend by 28%.
For streaming, I’ve engineered Kafka-based event pipelines with exactly-once semantics (plus Azure Event Hubs and Confluent Cloud) and transformed operational logs into analytics-ready layers in Snowflake and Amazon Redshift with DBT, Jinja, incremental materializations, and custom macros.
I’m hands-on with governance and reliability—Great Expectations and dbt tests, RBAC and PII masking via Microsoft Purview and Apache Atlas, plus IaC with Terraform and AWS CloudFormation. I also lead code review and testing standards for a team of four, and I’ve applied FinOps to save $140,000 annually through automated compute spin-downs and storage tiering.
Experience
Work history, roles, and key accomplishments
Lead Data Engineer & Data Architect
Dataquest
May 2023 - Present (3 years 3 months)
Architected and deployed Medallion Lakehouse Architecture using Databricks, Delta Engine, and PySpark, migrating 600+ TB of legacy storage into Delta Lake and Apache Iceberg. Designed enterprise pipelines in Apache Airflow and Azure Data Factory, orchestrating 150+ daily DAGs with 99.95% SLA.
Senior Data Engineer
Orderly Health
Sep 2019 - Apr 2023 (3 years 7 months)
Constructed real-time event-driven streaming pipelines using Apache Kafka, Confluent Cloud, and Azure Event Hubs, processing 20,000+ messages per second with exactly-once semantics. Developed scalable data transformations in Snowflake and Amazon Redshift using DBT and Jinja templating.
Data Engineer
Lendbuzz
Nov 2016 - Aug 2019 (2 years 9 months)
Built and scaled operational data ingestion pipelines using Python, SQL, and AWS Lambda. Migrated on-premise scripts to Apache Airflow, reducing pipeline failure points by 40%. Authored analytical queries and stored procedures across Oracle, IBM DB2, and MariaDB.
Education
Degrees, certifications, and relevant coursework
University of New South Wales
Bachelor of Science, Computer Science
2012 - 2016
Bachelor of Science in Computer Science from the University of New South Wales from 2012 to 2016.
Tech stack
Software and tools used professionally
Amazon Redshift
Airbyte
Fivetran
Apache Spark
AWS Glue
Apache Flink
Microsoft Azure
Google Cloud Platform
Amazon S3
AWS Step Functions
GitHub
GitLab
Kubernetes
Jenkins
GitHub Actions
GitLab CI
Salesforce
PySpark
dbt
DB
MySQL
PostgreSQL
MongoDB
Microsoft SQL Server
MariaDB
Cassandra
Hadoop
IBM DB2
YugabyteDB
CockroachDB
Databricks
Terraform
AWS CloudFormation
Java
Kafka
Amazon DynamoDB
GraphQL
Google Cloud Dataflow
Milvus
AWS Lambda
Serverless
Azure Functions
Amazon Aurora
Azure SQL Database
pytest
Airflow
Apache Beam
Google BigQuery
Amazon Athena
SQL
Azure Cosmos DB
Dagster
Apache Iceberg
Weaviate
Pinecone
Feast
DataHub
Delta Lake
Great Expectations
Apache Hudi
Collibra
dbt Cloud
Cosmos
Bash
Apache Parquet
Deequ
OpenLineage
Column
Unity Catalog
Factory
Beam
Microsoft Purview
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Anna?
You can contact Anna and 90k+ other talented remote workers on Himalayas.
Message AnnaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
