Skip to main content
Shreya RautSR
Open to opportunities

Shreya Raut

@shreyaraut

I build scalable, reliable data pipelines for analytics using Python, PySpark, dbt, Airflow, and AWS.

India
Message

What I'm looking for

I'm looking to build scalable, reliable data pipelines and analytics platforms, using Python, PySpark, dbt, Airflow, and AWS to improve data quality, performance, and delivery.

I build scalable, reliable data pipelines for analytics, most recently as a Senior Data Engineer at Simform after data engineering roles at Mindera and Cognizant.

At Mindera, I designed data-quality monitoring with dbt, Apache Airflow, and Tableau, and processed 50M+ Google Analytics records daily across BigQuery, Amazon S3, EMR, and AWS Glue. I also developed dbt ELT pipelines on Trino/Starburst for 10TB+ datasets and built real-time ingestion with Amazon Kinesis and Firehose, reducing latency from hours to minutes.

I automate dependable data platforms with Terraform, Docker, Jenkins, GitHub Actions, AWS Batch, ECS/Fargate, and Lambda. My CI/CD work reduced ETL release time by about 60%, while Hive and Apache Iceberg query optimizations reduced heavy query runtimes by 35–40%.

Earlier at Cognizant, I developed PySpark-based AWS Glue jobs, ETL scripts, MapReduce programs, and integrations across APIs, SFTP, SharePoint, SQL databases, PostgreSQL, and Oracle. I bring a practical foundation in data lakes, warehouses, distributed processing, and AWS-based data integration.

Experience

Work history, roles, and key accomplishments

Mindera - India logoMI

Data Engineer

Sep 2022 - Mar 2026 (3 years 6 months)

* Designed and automated data quality monitoring framework by integrating dbt with Apache Airflow, and visualizing KPIs in Tableau dashboards to validate data accuracy, freshness, and completeness ensuring business stakeholders had reliable, trustworthy metrics for decision-making.
* Processed cross-cloud Google Analytics data by exporting from BigQuery into Amazon S3, running HDFS-based distribut

Cognizant logoCO

Associate

Apr 2022 - Sep 2022 (5 months)

* Developing Glue ETL Job with Pyspark, for Data integration, specially handling complex unstructured data in a structured manner .
* Experience in debugging and scheduling AWS glue jobs Development experience using PySpark and following AWS technologies: Lambda, API Gateway, Glue, EMR, ECS, Kinesis, SQS, SNS, RDS, DynamoDb, Cognito, Redshift, AThena.
* Data Ingestion to one or more AWS Services

Cognizant logoCO

Programmer Analyst

Feb 2020 - Mar 2022 (2 years 1 month)

* Hands-on experience with AWS (EC2, Glue, Athena, Lambda, S3 and CloudWatch).
* Written Map Reduce programs using Pyspark for analyzing Big Data. Familiarity with DynamoDB and Data engineering.
* Skilled in ETL, Spark, and Python, with expertise in BigQuery and distributed computing systems for data storage.
* Developed ETL scripts to enable weekly processing of 2-4 GB data sets into data warehou

Education

Degrees, certifications, and relevant coursework

GT

Guru Nanak Institute of Enginnering and technology

Bachelor of Engineering - BE, Information Technology

2015 - 2019

GT

Guru Nanak Institute of Engineering and Technology

Bachelor of Engineering, Information Technology

2015 - 2019

Pursued a Bachelor of Engineering in Information Technology, gaining a strong foundation in software development and data engineering.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan