Gustavo Souza
@gustavosouza
Senior Data Platform Engineer with 8+ years building large-scale data systems. I specialize in streaming pipelines and ML infrastructure.
What I'm looking for
I’m a Senior Data Platform Engineer with 8+ years building large-scale data systems from first principles to production. At iFood, I worked on the company-wide ML Feature Store—serving 17+ squads and powering 200+ concurrent Spark jobs processing petabytes of data daily, with online serving latency under 20ms p99. I also led platform improvements that reduced costs, including a 30%+ cut in Redis infrastructure expenses through better resource allocation and pipeline efficiency.
At McGraw Hill, I own the Databricks platform end-to-end and deliver cloud-native, streaming-focused infrastructure that scales reliably while reducing operational burden for engineering teams. I built Cluster Autopilot to optimize job configurations via the Databricks Jobs API (saving over 30%), and created a Spark Performance Agent that ranks worst jobs and uses an LLM to generate code-level recommendations delivered to owners via Slack. Earlier, I re-architected Airflow on AWS ECS Fargate to eliminate weekly crashes and cut infrastructure costs by 60%+, and built CDC and streaming pipelines with Kafka/Kinesis, Spark Structured Streaming, and modern lakehouse tooling.
Experience
Work history, roles, and key accomplishments
Own the Databricks platform across the organization, including workspace administration, permissions, access controls, and governance. Built automated cluster optimization and a Spark performance agent to reduce Databricks costs and provide code-level recommendations for poorly performing jobs.
Worked on the company-wide ML Feature Store, serving 17+ squads and powering 200+ concurrent Spark jobs processing petabytes of data daily. Designed streaming and batch feature computation tiers and built an online serving layer with sub-20ms p99 latency.
Data Engineer
MadeiraMadeira
Nov 2019 - Dec 2020 (1 year 1 month)
Re-architected Airflow to run as containerized tasks on AWS ECS Fargate, cutting infrastructure costs and eliminating weekly crashes. Built CDC streaming pipelines to replicate SQL and MongoDB changes into a data lake in near real-time.
Junior Data Engineer
MadeiraMadeira
Dec 2018 - Nov 2019 (11 months)
Maintained and created Scala and Python ETL pipelines orchestrated by Apache Airflow. Worked on data modeling and exploration with Apache Spark and helped build the company’s first Data Lake while monitoring systems with Grafana.
Research Intern
Positivo University
Mar 2018 - Dec 2018 (9 months)
Conducted research on automated recognition of facial pain expressions in Wistar rats using computer vision and deep learning. Built supporting desktop tooling in C++ with Qt.
Education
Degrees, certifications, and relevant coursework
Positivo University
Computer Engineer, Computer Engineering
2021 -
Computer Engineering program at Positivo University in Curitiba starting in April 2021.
Availability
Location
Authorized to work in
Portfolio
gus-souza-portfolio.netlify.appSocial media
Job categories
Interested in hiring Gustavo?
You can contact Gustavo and 90k+ other talented remote workers on Himalayas.
Message GustavoGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
