
Siddartha Reddy Jammula
@siddarthareddyjammul
I build production data platforms and GenAI applications that make complex enterprise knowledge faster to use.
What I'm looking for
At the World Bank, I build PySpark data pipelines processing 2 TB monthly and a retrieval application over 2M+ internal documents that reduced analyst research time by 45%. I also designed a Neo4j graph layer that improved multi-hop question accuracy by 40%.
Previously, I led lakehouse, streaming, governance, and CI/CD work at JP Morgan, Nationwide, and Centene, delivering reliable analytics platforms for financial services, insurance, and HIPAA-sensitive healthcare data. I enjoy turning complex stakeholder requirements into governed, production-ready data and AI systems.
Experience
Work history, roles, and key accomplishments
Data Scientist II
World Bank
Apr 2024 - Present (2 years 5 months)
Build PySpark pipelines on Azure Databricks and Azure Data Factory processing 2 TB monthly into a Delta Lake medallion architecture. Architected a production retrieval application on Azure OpenAI and Azure AI Search over 2M+ internal documents, reducing analyst research time by 45%.
Designed and maintained production ETL/ELT pipelines in Azure Data Factory and Databricks ingesting from 50+ source systems, processing over 1 TB daily at 99.9% on-time delivery. Architected a Delta Lake lakehouse on ADLS Gen2 using a bronze/silver/gold medallion design.
Consolidated policy and claims data from 12 source systems into an Azure SQL data warehouse, building the ingestion pipelines and dimensional models behind cross line-of-business reporting. Migrated on-premises data workloads to ADLS Gen2 and Azure SQL, reducing infrastructure costs by 25%.
Built Python and SQL ETL processes for healthcare claims data, eliminating 10+ hours of manual processing per week. Tuned long-running SQL for claims reconciliation, reducing several overnight jobs to under 2 hours.
Analyzed research datasets in Python and SQL and built visualizations for university research teams. Cleaned and standardized survey data and automated recurring analyses previously performed in spreadsheets.
Wrote and optimized complex SQL with multi-table joins, window functions, and stored procedures over tables with millions of rows, cutting report generation time by 50%. Monitored scheduled ETL loads, troubleshot job failures, and reprocessed data to meet daily delivery SLAs across 3 enterprise client accounts.
Education
Degrees, certifications, and relevant coursework
University of Alabama at Birmingham
Master of Science, Data Science
2019 - 2020
Master of Science in Data Science from the University of Alabama at Birmingham from 2019 to 2020.
Tech stack
Software and tools used professionally
Availability
Location
Authorized to work in
Job categories
Skills
Interested in hiring Siddartha Reddy?
You can contact Siddartha Reddy and 90k+ other talented remote workers on Himalayas.
Message Siddartha ReddyGet matched with your dream remote job
Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!
