I work mainly with Python, LLM APIs and RAG.
I’m interested in retrieval, evaluation, structured outputs and failure analysis — basically the parts that become important once “it gave me a good answer once” stops being an acceptable test.
Most of what I’m building at the moment comes back to the same question: how do you know the AI system is actually working?
Currently building
A technical RAG system over industrial PC datasheets — with metadata-aware retrieval,
structured query analysis, LangSmith evaluation and grounded answer generation.
What I’m working on next
Generation evaluation | reranking | hybrid search | improving corpus-wide retrieval
Background
Software development background in Swift/iOS, with later training in data science, analytics and applied AI.
