Harish Rajoori
@harishrajoori
Staff Data Engineer specializing in 10 Years Architecting Cloud-Native Pipelines, lakehouse platforms and data governance
What I'm looking for
Data platform leader and architect with 10+ years of experience designing, scaling, and operating production lakehouse architectures, semantic layers, and high-throughput distributed systems. Currently leading core data platform initiatives at Warner Bros. Discovery, modernizing advertising reporting into a multi-source AWS Iceberg lakehouse served via a Cube REST API to a native React product interface.
Proven background decoupling analytical reporting from operational OLTP databases, orchestrating multi-source ingestion (PostgreSQL via DMS, Snowflake ad streams, greenfield ECS API extractors, Kafka), and establishing pre-aggregated semantic layers for sub-second analytical querying. Demonstrated impact includes slashing distributed batch execution times by 80%+, designing financial reconciliation engines, and scaling telemetry pipelines processing 10M+ daily events.
Core Skills & Specialties
Storage & Lakehouse: Apache Iceberg V2, Medallion Architecture (Bronze/Silver/Gold), Snowflake, Databricks, AWS Athena, AWS Redshift
Processing & Orchestration: Apache Spark (PySpark optimization), AWS Glue, AWS EMR, Apache Airflow (MWAA), Concurrency Mutex Locks
Semantic & Serving Layer: Cube.js (REST/SQL APIs, Pre-aggregations), dbt modeling, Declarative RLS/CLS, API/BFF Data Contracts
Ingestion & Streaming: AWS DMS (CDC & Full-Load), REST API Microservices, Amazon MSK / Apache Kafka, S3 Cross-Account Replication
Cloud & Platform Infrastructure: AWS (ECS Fargate, Lambda, Secrets Manager, VPC Peering), Terraform / Terragrunt, Docker, Kubernetes (EKS), GitHub Actions CI/CD
Location & Work Preferences
Location: Hyderabad, India
Work Preference: Remote (Global / Asynchronous / Distributed Teams)
Experience
Work history, roles, and key accomplishments
Lead/Staff Data Engineer at WBD (Ad Tech) driving the modernization of reporting platforms into an AWS Iceberg medallion lakehouse. Own end-to-end architecture across PySpark, DMS, Snowflake, and MWAA with concurrency mutex locks. Build governed Cube REST semantic layers and data contracts powering sub-second UI reporting, while enforcing multi-tenant security, Lake Formation TBAC, and CI/CD.
Led migration of a fragmented SQL experimentation system to a parallelized AWS EMR and Starburst ecosystem, reducing pipeline runtime. Implemented Apache Iceberg for ACID compliance and schema evolution, and modernized infrastructure and CI/CD for production deployments.
Senior Data Engineer
Intersoft DataLabs
Oct 2021 - Mar 2023 (1 year 5 months)
Designed a scalable lakehouse using AWS S3 and Redshift for high-performance serving. Built serverless ETL with Python, AWS Glue, and optimized PySpark jobs, and automated workflows with Airflow on Kubernetes to ensure data integrity and lineage.
Developed a proprietary framework using Java Spring Boot to orchestrate sensitive financial data workflows. Automated validation of financial records to support regulatory compliance and data privacy.
Data Engineer
SureIT Solutions
Oct 2019 - Aug 2021 (1 year 10 months)
Designed REST APIs and ingestion layers for a computer vision platform to support real-time UI/AI communication. Built services to handle batch image prediction requests across GPU servers for low-latency inference.
Migrated clinical trial data to AWS Redshift using serverless Python with AWS Lambda and AWS Glue pipelines. Refactored retail data pipelines using Apache Spark to reduce execution runtime.
Developed automated Oozie workflows and HiveQL scripts for credit risk reporting and financial data modeling. Delivered ETL in an Agile environment.
Education
Degrees, certifications, and relevant coursework
J.N.T. University
Bachelor of Technology, Electrical and Electronics Engineering
Bachelor of Technology in Electrical and Electronics Engineering from J.N.T. University, Hyderabad (2013).
Tech stack
Software and tools used professionally
Snowflake
Apache Spark
Apache Hive
Google Cloud Storage
Jupyter
MySQL WorkBench
dbt
Microsoft SQL Server
Django
Spring Boot
Google Analytics
Databricks
Terraform
Python
Java
Kafka
Apache NiFi
Django REST framework
Amazon RDS for PostgreSQL
Google Cloud SQL
Docker
Airflow
Google BigQuery
Amazon Web Services (AWS)
SQL
Google Cloud Dataproc
Apache Iceberg
Trino
Availability
Location
Authorized to work in
Salary expectations
Social media
Interested in hiring Harish?
You can contact Harish and 90k+ other talented remote workers on Himalayas.
Message HarishGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
