Thiago Nogueira
@thiagonogueira
I build reliable data, retrieval, and AI platforms at production scale.
What I'm looking for
At Nubank, I own the data and measurement layer for engineering productivity and AI adoption across 11,000+ employees and 10+ AI development tools. I build the governed pipelines, metrics, dashboards, and OKR tracking that give leadership a single source of truth for AI’s impact.
I’ve solved difficult data-integration problems by building a record-linkage model that raised coverage from 64.7% to 96.4%, eliminated manual mapping, and improved business-unit accuracy to 86.8%. I also redesigned productivity metrics using millions of activity records and am leading the migration of AI-usage data into a tested, monitored Scala pipeline with lineage.
Previously at Nubank, I built and scaled the engineering-productivity data platform, largely as the team’s sole data engineer. My work cut recovery time from about an hour to 13 minutes, doubled refresh frequency, reduced dashboard load times from 27 to 7 seconds, and brought Jira desynchronization down from 6.53% to 0.36%.
I’ve also built analytics and growth data systems at QuintoAndar, taught Statistics, Data Visualization, and Advanced SQL, and authored a full 170-hour data-analysis curriculum for Instituto Joga Junto. I’m pursuing graduate studies in Information Retrieval and Vector Databases, focused on the indexing, retrieval, and RAG layers behind production AI systems.
Experience
Work history, roles, and key accomplishments
Own the data and measurement layer for engineering productivity and AI adoption across Nubank's engineering organization.
- Own Nubank's company-wide AI-adoption measurement — the metrics, dashboards and OKR tracking usage of 10+ AI development tools across 11,000+ employees in multiple regions — giving leadership a single source of truth for AI's impact on engineering.
- Built a record-linkage m
Built and scaled the data platform behind engineering-productivity measurement at Nubank — for most of the period as the team's sole data engineer — with a focus on reliability, observability and cost efficiency.
- Designed and built an internal orchestration platform to integrate and synchronize data from disconnected internal systems (e.g., Jira), eliminating recurring pipeline incidents, cutti
Instituto Joga Junto is a nonprofit providing free technology training to people in situations of social vulnerability.
- Taught the Statistics, Data Visualization, and Advanced SQL modules of Entre-DADO, a 170-hour data analysis program (90 students per cohort, across online and in-person tracks).
- Led live sessions and designed hands-on exercises to build applied analytical skills, from descri
- Authored the full curriculum for Entre-DADO, a 170-hour data analysis program — writing every course module and defining its pedagogical structure and topic progression.
- Produced the complete base content later turned into all teaching materials, covering the program end to end from data fundamentals through advanced SQL.
- Cut 56 minutes off final data delivery time by refactoring legacy data-warehouse datasets into star-schema dimensional models.
- Built PySpark + Airflow DAGs to ingest product and API data, and engineered a pipeline unifying multiple customer-contact channels (SMS, email, chat, calls) into a single model.
- Implemented data-quality tests across pipelines, improving reliability of downstream mode
- Modeled data in LookML to build Explores and Views in Looker.
- Built strategic KPIs on customer contact and analyzed internal product-usage metrics.
- Designed A/B tests and experiments using Amplitude.
- Stood up the data function for a business unit, modeling and delivering data aligned to business needs.
- Built metric visualizations and dashboards in Looker and Metabase; developed and maintained Looker data models.
- Modeled datamarts in Redshift and SparkSQL; maintained analytical layers (enrichments and star-schema).
Undergraduate Research Assistant · Data Science / Bioinformatics
Jan 2018 - Apr 2018 (3 months)
Aquifer Microbiomes project, funded by Instituto Serrapilheira (a major private science funder in Brazil).
- Built Python crawlers to collect metagenomics data from international public repositories via APIs.
- Ingested JSON data and converted it into structured datasets; handled cleaning and curation.
- Delivered the final curated datasets as CSV to the lab's research team.
Apprenticeship in the technology department of one of Bahia's largest broadcasters, supporting IT and telecom routines.
Education
Degrees, certifications, and relevant coursework
Universidade Federal da Bahia
Bacharelado, Física
Universidade Federal da Bahia
Mestrado em Ciências da Computação, Information Retrieval and Vector Databases
2025 - 2027
Instituto de Ciências Matemáticas e de Computação (ICMC) - USP
Pós-graduação Lato Sensu - Especialização, Inteligência Artificial e Big Data
2025 - 2026
Universidade Federal da Bahia
Bacharelado Interdisciplinar, Ciência e Tecnologia
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Thiago?
You can contact Thiago and 90k+ other talented remote workers on Himalayas.
Message ThiagoGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
