Anusha Shrestha
@anushashrestha
Senior Data Engineer specializing in scalable lakehouse, streaming ETL, and AI-powered data products.
What I'm looking for
I’m a results-driven Senior Data Engineer with 6+ years of experience designing scalable data platforms and ETL/ELT pipelines across AWS, Azure, and GCP. I build cloud-native lakehouse architectures and modern data warehousing solutions using Python, PySpark, Scala, SQL, Apache Spark, Databricks, Delta Lake, and Snowflake—supporting analytics, governance, and self-service reporting.
I also lead end-to-end real-time and near real-time streaming systems with Kafka, Spark Structured Streaming, and cloud messaging services, delivering sub-second analytics and low-latency operational insights. I’ve optimized performance and cost (including reducing processing times by 60% and cloud costs by 40%), automated infrastructure and deployments with Terraform/Docker/Kubernetes and CI/CD, and implemented enterprise data quality, lineage, observability, and security aligned with GDPR, HIPAA, and SOC 2. Lately, I’ve been developing Generative AI and RAG solutions (Azure OpenAI, LangChain, LangGraph, vector databases) to automate document processing, knowledge extraction, and contextual question answering.
Experience
Work history, roles, and key accomplishments
Senior Data Engineer
Designed and optimized enterprise ETL/ELT pipelines using PySpark, Databricks, Delta Lake, and Snowflake to process healthcare datasets for analytics and regulatory reporting. Built real-time streaming and Lakehouse architectures, implemented data governance, and developed RAG solutions using Azure OpenAI and LangChain/LangGraph to support healthcare document processing.
Developed scalable PySpark/Scala/Spark SQL applications and enterprise ETL/ELT pipelines to process large financial and customer datasets, improving processing performance by 35%. Built Medallion Lakehouse pipelines, near real-time streaming solutions, optimized analytics workloads, and supported self-service reporting for 300+ business users.
ETL Developer
Change Healthcare
Assisted with ETL workflow development and maintenance using AWS Glue and PySpark to process structured and semi-structured data for analytics. Supported AWS migration of legacy ETL processes, data loading to S3/RDS/Snowflake (via Matillion), and troubleshooting of Kafka and Spark Streaming pipelines.
Education
Degrees, certifications, and relevant coursework
The University of Texas at Arlington
Bachelor of Science in Business Analytics, Business Analytics
Earned a Bachelor of Science in Business Analytics at the University of Texas at Arlington.
Tech stack
Software and tools used professionally
Amazon Redshift
Matillion
Azure Synapse
Apache Spark
AWS Glue
Apache Flink
Talend
AWS IAM
Amazon S3
AWS Step Functions
GitHub
GitLab
Bitbucket
Azure Repos
Kubernetes
Jenkins
GitHub Actions
GitLab CI
NumPy
Pandas
PySpark
dbt
Sqoop
MySQL
PostgreSQL
MongoDB
Cassandra
Hadoop
HBase
Databricks
Terraform
Azure DevOps
Jira
Java
JSON
MATLAB
XML
Apache Flume
Kafka
Azure Monitor
Ubuntu
CentOS
Linux
macOS
Windows
Datadog
Google Cloud Dataflow
Amazon Kinesis
Avro
Azure Functions
Amazon RDS
Azure SQL Database
Kafka Streams
Airflow
Time Analytics
s3-lambda
Google BigQuery
Amazon EMR
SQL
Apache Iceberg
LangChain
Delta Lake
Great Expectations
dbt Cloud
Bash
LangGraph
Microsoft Fabric
Unity Catalog
Factory
Microsoft Purview
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Anusha?
You can contact Anusha and 90k+ other talented remote workers on Himalayas.
Message AnushaGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
