I want to own data infrastructure end to end. Focused on modern data stack work (dbt, Snowflake, Databricks) and data infrastructure for AI products — especially RAG and retrieval systems. Comfortable with messy legacy systems and undocumented business rules; I do my best work building, not just reviewing.
Leonardo Santos
@leonardosantosdev
Senior Data Engineer | dbt, Snowflake, Databricks, PySpark | RAG & LLM Data Pipelines
What I'm looking for
Senior Data Engineer with 5 years building data platforms across financial risk, manufacturing, and US insurance regulatory reporting.
I've owned credit, fraud, and collections pipelines processing billions of records, delivered a year-long on-prem to GCP migration, and modernized legacy Informatica workflows into dbt and Snowflake for regulatory reporting at a Fortune 500 carrier.
Recent work includes building a production RAG pipeline (chunking, embeddings in PostgreSQL/pgvector, and retrieval feeding an LLM agent with tool-calling constraints) alongside the multi-tenant platform it ran on.
Core stack: PySpark, dbt, Snowflake, Databricks, Airflow, Kafka, PostgreSQL, AWS/GCP.
Experience
Work history, roles, and key accomplishments
Supporting modernization of 192 regulatory data calls from Informatica PowerCenter to dbt/Snowflake. Technical reviewer for an external migration vendor, auditing dbt models and surfacing reconciliation gaps. Reverse-engineer dbt models into source-to-target mappings for teams without SQL access. Technical point of contact for dbt/GenAI; led the interview for a dbt Tech Lead hire.
Software Engineer
Space Sales
Jul 2025 - Dec 2025 (5 months)
Delivered a multi-tenant CRM SaaS from architecture to production as the project's only engineer: data model, backend, frontend, and deployment. Built a production RAG pipeline with chunking, embedding generation, and per-message retrieval feeding an LLM agent, with vectors in PostgreSQL/pgvector. Designed an agent orchestration layer with configurable playbooks and tool-calling constraints.
Automated 10+ manual SAP extraction workflows into scheduled Databricks pipelines for a global automotive manufacturer's shop floor and supply chain operations. Reverse-engineered undocumented Power BI/Power Query logic into PySpark and Spark SQL. Built medallion architecture (bronze/silver/gold) on Delta Lake over SAP ERP data. Migrated legacy SQL Server to Databricks.
Data Engineer
2RP Net
Nov 2021 - Aug 2024 (2 years 9 months)
Sole data engineer for a major Brazilian retail credit operation's financial risk squad, owning credit, fraud, and collections pipelines processing billions of records. Built NiFi/Kafka/Airflow/PySpark workflows, eliminated recurring production incidents, and led a year-long OCI-to-GCP migration with hash-based validation. Integrated anti-fraud APIs; promoted from trainee to sole data engineer.
IT Coordinator
Prefeitura Municipal Muzambinho
Jan 2018 - Dec 2020 (2 years 11 months)
Sole IT resource for 8 public health units. Deployed the e-SUS EHR system, trained 50+ clinical staff, and managed server infrastructure and weekly backups.
Education
Degrees, certifications, and relevant coursework
Instituto Federal de Educação, Ciência e Tecnologia do Sul de Minas Gerais
Bachelor of Science, Computer Science
2015 - 2018
Grade: 3.5
Activities and societies: IT Support at Muzambinho City Hall (2017–2018) while studying
Availability
Location
Authorized to work in
Salary expectations
Job categories
Interested in hiring Leonardo?
You can contact Leonardo and 90k+ other talented remote workers on Himalayas.
Message LeonardoGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
