Skip to main content
Vanshika GoyalVG
Open to opportunities

Vanshika Goyal

@vanshikagoyal2

Data Engineer with 1.5+ years building scalable Azure cloud data platforms and streaming ETL pipelines.

India
Message

What I'm looking for

I’m looking for a Data Engineering role where I can design scalable Azure data platforms, build reliable streaming and ETL pipelines, optimize Spark performance, and deliver analytics-ready datasets with strong orchestration and monitoring.

I’m a Data Engineer with 1.5+ years of experience designing scalable, cloud-native data platforms on Microsoft Azure. I build distributed ETL and streaming pipelines using PySpark, SQL, Azure Databricks, and Delta Lake, with a strong focus on producing analytics-ready datasets.

In my current role at Samsung Research Institute - Delhi, I’ve designed Bronze/Silver/Gold pipelines following the Medallion Architecture with ADLS Gen2 and Delta Lake. I automate workflow scheduling, monitoring, and dependency management using Apache Airflow, and I optimize Spark workloads through partitioning, caching, and broadcast joins. I also developed reusable Airflow DAGs with retry policies, SLA monitoring, and notifications to support production-grade ETL.

I’m also driven by measurable impact and continuous improvement—like reducing query latency by 50% in my internship and delivering dashboards in Power BI for channel performance and engagement. I hold Microsoft certifications including Azure Data Engineer Associate (DP-203) and Fabric Data Engineer Associate.

Experience

Work history, roles, and key accomplishments

SD
Current

Engineer I

Samsung Research Institute - Delhi

May 2025 - Present (1 year 3 months)

Designed scalable ETL pipelines for the CTS click-to-search feature using PySpark, SQL, and Azure Databricks, implementing Medallion Architecture (Bronze/Silver/Gold) with ADLS Gen2 and Delta Lake. Automated and optimized production workflows with Apache Airflow and Spark performance tuning, and built batch/streaming analytics pipelines for Samsung TV Plus traffic metrics with Power BI reporting.

SD

Associate Intern

Samsung Research Institute - Delhi

Dec 2024 - May 2025 (5 months)

Developed ETL workflows with Azure Databricks, PySpark, and SQL for enterprise analytics, implementing Delta Lake using SCD Type 2 for historical data management. Optimized Spark jobs via partitioning and caching (reducing query latency by 50%) and built analytics-ready datasets on ADLS Gen2 for downstream reporting.

Education

Degrees, certifications, and relevant coursework

AC

ABES Engineering College

Bachelor of Technology (B.Tech), Computer Science

2021 - 2025

Grade: CGPA: 8.2

Pursuing a B.Tech in Computer Science with a CGPA of 8.2.

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan