Skip to main content
AS
Open to opportunities

Anusha Shrestha

@anushashrestha

Senior Data Engineer specializing in scalable lakehouse, streaming ETL, and AI-powered data products.

United States
Message

What I'm looking for

I’m looking to build secure, scalable lakehouse + streaming ETL systems, strengthen governance/observability, and apply AI (RAG/LLMs) to create trusted analytics that teams can act on.

I’m a results-driven Senior Data Engineer with 6+ years of experience designing scalable data platforms and ETL/ELT pipelines across AWS, Azure, and GCP. I build cloud-native lakehouse architectures and modern data warehousing solutions using Python, PySpark, Scala, SQL, Apache Spark, Databricks, Delta Lake, and Snowflake—supporting analytics, governance, and self-service reporting.

I also lead end-to-end real-time and near real-time streaming systems with Kafka, Spark Structured Streaming, and cloud messaging services, delivering sub-second analytics and low-latency operational insights. I’ve optimized performance and cost (including reducing processing times by 60% and cloud costs by 40%), automated infrastructure and deployments with Terraform/Docker/Kubernetes and CI/CD, and implemented enterprise data quality, lineage, observability, and security aligned with GDPR, HIPAA, and SOC 2. Lately, I’ve been developing Generative AI and RAG solutions (Azure OpenAI, LangChain, LangGraph, vector databases) to automate document processing, knowledge extraction, and contextual question answering.

Experience

Work history, roles, and key accomplishments

Cardinal Health logoCH
Current

Senior Data Engineer

Designed and optimized enterprise ETL/ELT pipelines using PySpark, Databricks, Delta Lake, and Snowflake to process healthcare datasets for analytics and regulatory reporting. Built real-time streaming and Lakehouse architectures, implemented data governance, and developed RAG solutions using Azure OpenAI and LangChain/LangGraph to support healthcare document processing.

Wells Fargo logoWF

Data Engineer

Jan 2023 - Present (3 years 6 months)

Developed scalable PySpark/Scala/Spark SQL applications and enterprise ETL/ELT pipelines to process large financial and customer datasets, improving processing performance by 35%. Built Medallion Lakehouse pipelines, near real-time streaming solutions, optimized analytics workloads, and supported self-service reporting for 300+ business users.

Education

Degrees, certifications, and relevant coursework

The University of Texas at Arlington logoTA

The University of Texas at Arlington

Bachelor of Science in Business Analytics, Business Analytics

Earned a Bachelor of Science in Business Analytics at the University of Texas at Arlington.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan