I’ve built and evaluated LLM-based systems through independent technical work, comparing model behaviour across reasoning, coding, research, conversation, and tool use. I designed prompts and practical rubrics to assess instruction following, context retention, factual reliability, and recurring failure modes.
I built persistent memory and context-assembly workflows for multi-session interactions, including chronological recall and failure recovery. I also engineered multi-step agent workflows with provider abstraction, model routing, structured handoffs, and explicit task state.
For real-time voice interaction, I implemented streaming speech systems with local STT, TTS coordination, wake-word inference, and interruption handling. I’ve developed and tested asynchronous systems using TypeScript, Swift, and Python, and evaluated audio turn detection, latency, buffering, and reliability.

