At Modal, I build serverless GPU inference and AI model deployment infrastructure across AWS, Google Cloud, and OCI. I develop distributed Python pipelines for large-scale inference, embeddings, and fault-tolerant AI workloads.
Previously at Weaviate, I built RAG and generative-search workflows using document ingestion, embeddings, FAISS, Pinecone, hybrid search, and HNSW indexing. I also designed cloud-native Python, GraphQL, Docker, and Kubernetes services for vector retrieval at scale.
I've built reliable backend systems for Agility Robotics, Lazada Philippines, and GCash, covering robot telemetry, marketplace inventory, order processing, and payment workflows. My work focuses on practical scalability, latency, reliability, and correctness under production traffic.

