I automate application delivery and build the infrastructure and operational systems required to run applications reliably.
My focus is DevOps, Platform engineering, Cloud infrastructure, and Site reliability engineering. I am particularly interested in the problems at the boundary between applications and infrastructure; ie where deployments, dependencies, capacity, latency, resource pressure, and operational decisions interact.
- Kubernetes & Platforms — designing standardized environments that make applications easier to deploy, operate, and scale.
- Application Delivery & Packaging — building automated workflows that validate, test, build, package, and containerize applications into reproducible deployable artifacts.
- Infrastructure & Cloud — building reproducible infrastructure and automated environments that provide a reliable foundation for workloads.
- Reliability Engineering — testing and analyzing failure modes, dependencies, latency, saturation, capacity limits, failure propagation, and recovery.
- Observability — building metrics, monitoring, and operational signals that make system behavior visible and explainable.
- Automation — replacing repetitive and error-prone operational work with Python, Bash, infrastructure automation, and CI/CD workflows.
My goal is to turn systems that are manual, fragile, or difficult to understand into systems that are repeatable, observable, resilient, and easier to operate.
A healthy platform should assumes that people will make mistakes. A single mistake should not become a serious production incident; it should be caught by validation, isolated in a controlled environment, limited in blast radius, quickly visible, or easy to reverse. A system can be healthy one moment and approaching saturation the next. A dependency can be reachable while still being unable to complete useful work. Retries intended to improve reliability can increase pressure on an already struggling service. A deployment can succeed technically while introducing problems that only become visible under real traffic. So my understanding of a healthy system is in its reliability, not status. That is why I am interested in understanding the relationship between:
Validate → Deploy → Load → Latency → Queueing → Saturation → Errors → Failure Propagation → Recovery
When something goes wrong, I want to know: What failed, why it failed, how the failure propagated, what signals revealed it, and how the system can recover.
AWS · Microsoft Azure · Terraform · Linux
Docker · Kubernetes · Helm
Prometheus · Grafana · Alertmanager · Metrics · Monitoring · Alerting · SLIs/SLOs · Incident Analysis
Python · Bash · Git · GitHub Actions · Argo CD · GitOps · CI/CD
Application Delivery · Cloud Infrastructure · Platform Engineering · Distributed Systems · Reliability Engineering · Failure Analysis · Capacity & Saturation · Operational Automation
I'm interested in engineering problems involving cloud infrastructure, application delivery, platform engineering, reliability, distributed systems, observability, and automation.
Email: aniokemark@gmail.com
LinkedIn: linkedin