AI engineer at Bind, a Finnish AI startup. Building agent tooling, writing about what works.
- Escalating Evals: 72% Fewer LLM Judge Calls
- TDD in the age of agents
- Agent-as-Judge: Harder Evals, More Computation
- Jev-as-Judge: A Confidence Signal for Evals
- Principles of Loop Engineering
More on samikuikka.com.
Building apo โ an opinionated framework for testing agent systems end-to-end.



