At OSI Digital, I build and maintain 100+ production crawlers for the OLX project, extracting listing details, category counts, and structured datasets at millions-of-records scale each day.
I engineer mobile API crawling pipelines by intercepting and replicating application traffic, reverse-engineer hidden REST and GraphQL APIs, and keep high-volume extraction running through Cloudflare challenges, CAPTCHA handling, proxy rotation, and resilient scraping middleware.
I also design GDPR-compliant ETL pipelines processing 5M+ records daily into AWS S3, using PySpark for historical data correction and delivering audit-ready datasets for analytics and compliance teams.
