
Lucas Perez
@lucasperez
I build interpretable machine learning systems for time-series retrieval, imputation, and language-model research.
What I'm looking for
At Petrobras, I build representation-learning systems for multivariate well-log retrieval, helping geologists find analogous subsurface patterns for drilling decisions. I develop Transformers, LSTMs, CNNs, and 1D region-proposal models for sensor-stream analysis and geological-event localization.
I'm pursuing an M.Sc. in Computer Science at UFMG while researching Fisher Information-based analyses of pruning, continual learning, sparse autoencoders, and mechanistic interpretability. My work has contributed to publications in Computers & Geosciences, ACM Hypertext, and the ICML 2026 Workshop on Weight-Space Symmetries.
Experience
Work history, roles, and key accomplishments
• Built representation-learning models for similarity search and retrieval of multivariate well-log time series, enabling geologists to find analogous subsurface patterns to support drilling decisions.
• Implemented and tested Transformers, LSTMs, and CNNs to transform raw sensor streams into reusable embeddings for retrieval and downstream analysis.
• Developed a Faster R-CNN-inspired 1D region-p
• Trained and evaluated sparse-autoencoder variants, including Matryoshka and Temporal SAEs, to investigate hierarchical structure in language-model representations.
• Analyzed learned features and model behavior to compare how different SAE architectures recover meaningful hierarchical organization for mechanistic interpretability.
• Oversaw weekly activities for 100+ students across Statistical Foundations of Data Science and Numerical Calculus courses at UFMG.
• Mentored students, developed exercises, graded assignments, answered coursework questions, and supported the preparation of class materials.
• Investigated Fisher Information approximations for continual learning, comparing diagonal and block-diagonal formulations and their effects on model stability and knowledge retention.
• Built and maintained evaluation pipelines to benchmark catastrophic forgetting and retention across continual-learning scenarios.
• Developed an LLM-based pipeline to turn unstructured medical prescriptions into structured fields, combining model outputs with validation and post-processing to reduce manual review.
• Supported an internal OCR system by helping improve data quality and analyzing common failure cases, increasing reliability in production-like inputs.
• Built data-cleaning and preprocessing pipelines for multivariate well-log time series, preparing raw sensor data for similarity-search and retrieval models used by geologists.
• Developed missing-data handling workflows, including imputation modeling using XGBoost, LSTMs and Transformers, to mitigate sensor gaps and ensure consistent inputs for retrieval/embedding training.
• Co-developed a benc
• Graded 500+ second-round OBMEP exams, evaluating mathematical reasoning and written problem-solving.
• Designed and delivered an applied Data Science curriculum covering Python, Pandas, Scikit-learn, and PyTorch.
• Mentored a laboratory-outcome prediction project from problem definition through model evaluation and iteration, supporting faster operational decisions and reduced turnaround time.
• Designed and delivered an applied Data Science curriculum covering Python, Pandas, Scikit-learn, and PyTorch.
• Guided the development of medium-term energy-demand forecasting models to support energy-contracting decisions for industrial operations, with core results published at ABM.
• Built a BeautifulSoup data-collection pipeline covering 1M+ videos and nearly 10K channels from alternative video platforms, including metadata cleaning, validation, and dataset organization.
• Analyzed how YouTube moderation events relate to channel popularity using the collected dataset and YouTube API, contributing to research published at ACM Hypertext 2022.
Education
Degrees, certifications, and relevant coursework
Universidade Federal de Minas Gerais
M.Sc., Computer Science
2025 - 2027
Co-advisors: Renato Assunção and Fabricio Murai
• Developed Fisher Information-based methods to analyze how pruning reshapes statistical dependencies between neural-network weights, building reproducible PyTorch experiments that led to work published at the ICML 2026 Workshop on Weight-Space Symmetries.
Universidade Federal de Minas Gerais
B.Sc., Computational Mathematics
2019 - 2024
DeepLearning.AI
Convolutional Neural Networks
Issued Jul 2024
Departamento de Ciência da Computação - UFMG
Hands-on Deep Learning – 3ª. Edição – Verão 2020
Issued Mar 2020
Colégio Santo Agostinho
2007 - 2018
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Salary expectations
Social media
Skills
Interested in hiring Lucas?
You can contact Lucas and 90k+ other talented remote workers on Himalayas.
Message LucasGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
