At Biuro Informacji Kredytowej, I build low-latency inference pipelines for private open-source models in an air-gapped banking network. I also designed regression tests for structured entity extraction and model output stability.
On OneSpark, I served a 135 GB model on NVIDIA DGX Spark under SGLang. I used HashK compression to reduce the embedding table from 51.2 GB to 12.8 GB, with no degradation on complex reasoning benchmarks.
At Chabre Data & AI Advisory, I advised technology clients on moving from commercial LLM APIs to self-hosted inference, and benchmarked runtimes to reduce cloud compute costs by up to 65%. I also published research on model refusal removal and its effects on decision disposition.

