Prashis Shrestha
@prashisshrestha
Senior data engineer building near real-time healthcare data pipelines on Spark and Kafka.
What I'm looking for
I build ETL and data ingestion pipelines for healthcare and multi-cloud analytics, from PySpark Data Ingestion frameworks to Hadoop-based batch processing. At Change Healthcare, I optimized large-scale Medicare, Medicaid, and commercial datasets to support advanced analytics.
I’ve led near real-time ingestion with Spark Streaming, wiring Spark to Flink, Kafka, AWS Kinesis, and APIs for low-latency processing and real-time analytics. I also deploy Spark applications with Docker and Kubernetes, and automate infrastructure and workflows using Terraform, AWS Step Functions, and Apache Airflow DAGs.
I migrate and modernize data platforms across environments—moving on-premises SQL Server, Oracle, and MongoDB into Azure Data Lake with Azure Data Factory, SQL Azure, SSIS, and PowerShell, and running Spark on AWS EMR with S3 and Redshift. I’ve standardized transformations with DBT, including schema design work using star and snowflake models.
Earlier, at PNC Bank and Goldman Sachs, I engineered Spark/Hadoop ETL with medallion (Bronze/Silver/Gold) layers, configured Kafka for streaming pipelines, and delivered BI reporting in Tableau and Power BI. I’ve also supported NLP-driven work with Generative AI and LLMs for healthcare text tasks, and supervised 3+ engineers to keep timelines and quality on track.
Experience
Work history, roles, and key accomplishments
Senior Data Engineer
Change Healthcare
Jun 2022 - Present (4 years 2 months)
Designed and optimized ETL pipelines using PySpark for large-scale healthcare datasets. Built real-time data ingestion pipelines with Spark Streaming, Kafka, and AWS Kinesis.
Data Engineer
PNC Bank
Jan 2020 - May 2022 (2 years 4 months)
Developed high-performance Spark code using PySpark and Scala, optimizing data transformations within the Hadoop ecosystem. Engineered scalable ETL pipelines for multi-cloud data integration.
Data Engineer
Goldman Sachs
Mar 2017 - Dec 2019 (2 years 9 months)
Designed and implemented ETL workflows using AWS Glue and GCP Dataflow for multi-cloud integration. Built and managed data pipelines to transform and load datasets into AWS S3, Snowflake, and RDS.
Education
Degrees, certifications, and relevant coursework
Texas Tech University
Bachelor of Business Administration, Information Technology
Bachelor's degree in Business Administration Information Technology from Texas Tech University.
Tech stack
Software and tools used professionally
Matillion
Azure HDInsight
Azure Synapse
Apache Spark
AWS Glue
Talend
Google Cloud Storage
AWS Step Functions
GitHub
GitLab
Kubernetes
Jenkins
NumPy
Pandas
PySpark
dbt
Sqoop
MySQL
PostgreSQL
MongoDB
Cassandra
Hadoop
HBase
Sybase
Yarn
Databricks
Terraform
Azure DevOps
Jira
Java
JSON
PowerShell
MATLAB
XML
Kafka
Linux
macOS
Windows
Serverless
Airflow
Time Analytics
s3-lambda
SQL
Hugging Face
LangChain
Bash
Factory
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Prashis?
You can contact Prashis and 90k+ other talented remote workers on Himalayas.
Message PrashisGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
