Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
27 changes: 27 additions & 0 deletions .github/workflows/section-divider-check.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
name: No section dividers

on:
push:
branches: [main]
pull_request:
workflow_dispatch:

jobs:
check:
name: scan markdown for section dividers
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
with:
fetch-depth: 1

- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.13"

- name: Run section-divider check on all tracked posts
run: scripts/check_section_dividers.py
9 changes: 9 additions & 0 deletions .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,15 @@ repos:
# so the script's "if argv: scan those" path takes over.
# No additional args needed.

- id: no-section-dividers
name: no horizontal-rule section dividers in posts
entry: scripts/check_section_dividers.py
language: python
types: [markdown]
files: ^content/
# Pre-commit passes only the staged files, so the script scans
# those; the script itself skips vendored/generated content docs.

- id: paradigm-shift-warn
name: warn on antithesis / paradigm-shift sentences
entry: scripts/check_paradigm_shift.py
Expand Down
6 changes: 0 additions & 6 deletions content/208-day-kernel-bug-lessons.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,8 +25,6 @@ return ns;

Once the computed value exceeded `0xffffffffffffffff`, it wrapped around, leading to undefined behavior in the scheduler and an unrecoverable state requiring a manual reboot.

---

## Why This Matters to DevSecOps

This bug is more than a curiosity. It's a classic case study in:
Expand All @@ -37,17 +35,13 @@ This bug is more than a curiosity. It's a classic case study in:

Affected systems included RHEL 5.0 through 5.5 and early RHEL 6 versions running kernels below `2.6.32-220.4.*`. Some Debian-based distributions were likely impacted, though documentation was less complete.

---

## Takeaways for Modern Systems

- **Live patching tools** like Ksplice, KernelCare, and kpatch can reduce reboot pressure
- **Observability stacks** should alert on uptime thresholds and kernel messages (`dmesg`, `uptime`, scheduler warnings)
- **Compliance frameworks** often require timely OS patching, and this bug illustrates why
- **CI/CD pipelines for OS-level components** should test for edge cases, including time-based and overflow scenarios

---

Even today, this incident reminds us that uptime isn't always a badge of honor. In some cases, it's a quiet countdown to failure. For another time-math failure at cloud scale, see [The Chaos of the Leap Second (2012)]({filename}leap-second-chaos-2012.md).

*Originally inspired by a 2012 analysis of the `sched_clock()` bug affecting Linux systems with prolonged uptime.*
Expand Down
18 changes: 0 additions & 18 deletions content/hardening-k8s.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,6 @@ securityContext:

This enforces that containers don't run as UID 0, reducing the blast radius of any compromise.

---

### 2. Drop All Linux Capabilities

```yaml
Expand All @@ -41,8 +39,6 @@ securityContext:

Drop all capabilities by default, then add only what you need.

---

### 3. Disable Privilege Escalation

```yaml
Expand All @@ -52,8 +48,6 @@ securityContext:

This prevents processes inside the container from gaining additional privileges, even if compromised.

---

### 4. Use Read-Only Filesystem

```yaml
Expand All @@ -63,8 +57,6 @@ securityContext:

This blocks attackers from writing malicious files or installing tools inside the container.

---

### 5. Avoid Host Access

```yaml
Expand All @@ -75,8 +67,6 @@ hostIPC: false

Avoid `hostPath` volumes unless absolutely required. These settings ensure your workloads remain isolated from the host.

---

### 6. Use Trusted Images and Scan Them

Use minimal base images (Alpine, Distroless) and trusted registries. Always scan them:
Expand All @@ -87,8 +77,6 @@ trivy image your-registry/app:tag

This helps catch known CVEs before deployment.

---

### 7. Handle Secrets via Volumes (Not Env Vars)

```yaml
Expand All @@ -107,8 +95,6 @@ containers:

Mounting secrets as volumes avoids accidental exposure via logs or `/proc`.

---

### 8. Restrict Network Traffic with NetworkPolicies

```yaml
Expand All @@ -124,8 +110,6 @@ spec:

Start with a default-deny policy per namespace, then explicitly allow only the traffic your services need. Without NetworkPolicies, any pod can communicate with any other pod in the cluster.

---

### 9. Harden ServiceAccount Usage

```yaml
Expand All @@ -134,8 +118,6 @@ automountServiceAccountToken: false

Disable automatic token mounting for pods that don't need API server access. Create dedicated ServiceAccounts with minimal RBAC bindings rather than relying on the `default` account, which often accumulates unnecessary permissions.

---

## Final Thoughts

Tools matter less than secure defaults. These practices help harden your Kubernetes workloads using the Restricted Pod Security Standard and reduce risks across the board.
Expand Down
14 changes: 0 additions & 14 deletions content/hiring-discovery-layer-broken.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,8 +20,6 @@ Many of those same engineers struggle to get past the first screen.

That gap says something important about the hiring system. Hiring has a discovery problem. The industry has signal. The signal is poorly indexed, poorly routed, and poorly evaluated.

---

## Observability everywhere except the hiring funnel

Modern engineering organizations spend enormous effort building observability into production systems. They trace requests, measure latency, monitor failure modes, and build dashboards so they can understand what is happening under pressure.
Expand All @@ -32,8 +30,6 @@ A resume is a lossy compression format for engineering judgment. It takes years

That format can find familiar labels. It struggles to find judgment.

---

## Senior signal lives in context

Senior engineering skill usually shows up in context:
Expand All @@ -52,8 +48,6 @@ This creates a predictable failure mode. Candidates optimize resumes for search.

The system then calls the result a talent shortage.

---

## Some shortages are real. This one is mostly routing.

Some shortages are real. Security, infrastructure, identity, Kubernetes, cloud architecture, incident response, and compliance automation all require people who have made real decisions under constraint. Those people are hard to find.
Expand All @@ -70,8 +64,6 @@ If your hiring process cannot distinguish operational judgment from keyword opti

A process that screens mostly for tool names will undervalue people who understand failure modes. A funnel that treats context as noise will discard the signal senior roles depend on.

---

## The fix starts with calibration

Hiring managers need to give recruiters sharper signal than a list of tools. A good intake should include examples of real problems the person will solve, failure modes the team cares about, and evidence that would prove competence.
Expand All @@ -86,8 +78,6 @@ For example:

Those questions produce better screens than a wall of product names.

---

## Better artifacts than resumes

Candidates also need better artifacts.
Expand All @@ -100,8 +90,6 @@ Most hiring systems ignore them because they are not structured for automated pa

The shape works both ways. My own writeup of a program that produced measurable outcomes plus the political stall I caused myself, [Adoption Is a Security Control]({filename}paved-road-adoption-as-control.md), is the kind of artifact I would rather be evaluated on than a bullet list. The complement to this post is [The Trust Decay]({filename}trust-decay-adversarial-hiring.md), which frames the same pipeline from the candidate side.

---

## What hiring teams should actually ask for

Teams should look for direct proof.
Expand All @@ -115,8 +103,6 @@ Teams should look for direct proof.

These questions surface judgment. Keyword filters surface vocabulary. The two are not the same signal.

---

## The routing failure has a name

The hiring market feels broken because discovery is broken. The people exist. The work exists. The evidence exists. The pipeline loses too much signal before it reaches the people qualified to evaluate it.
Expand Down
12 changes: 0 additions & 12 deletions content/iam-blast-radius-architecture-problem.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,8 +21,6 @@ At that point, the policy review is often just cleanup. Useful cleanup, but stil

The real IAM question is not only, "Does this policy look least privileged?" The real question is, "What happens when this identity is compromised?"

---

## Policies as documents vs identities as architecture

A lot of IAM programs stay shallow because they focus on permissions as isolated documents instead of treating identities as part of the system architecture. A policy can avoid wildcards and still create an unacceptable failure domain. A role can look narrow in isolation and still become dangerous because of who can assume it, what it can mutate, or what other paths it unlocks.
Expand All @@ -36,8 +34,6 @@ Least privilege is not just fewer actions. Least privilege is smaller failure do

This is why reviewing IAM as JSON hygiene misses the point.

---

## Better questions than "does this policy contain `*`?"

The dangerous question is rarely, "Does this policy contain `Action: *`?"
Expand All @@ -58,8 +54,6 @@ Those questions move IAM reviews out of the policy document and back into the ar

I turned those questions into a browser tool: [IAM Blast Radius analyzer](/tools/iam-blast-radius/) checks a pasted IAM policy for escalation paths, role-assumption reach, and data-access exposure locally, grading each finding by certainty - potential blast radius, not effective permissions.

---

## Attackers experience the environment as a graph

Attackers do not experience your environment as a set of clean policy JSONs. They experience it as a graph. Trust relationships, deployment pipelines, secrets, buckets, queues, roles, runners, clusters, and logs all become edges in that graph.
Expand All @@ -74,8 +68,6 @@ A permission that looks harmless in one context can become serious in another:

IAM design has to account for that graph.

---

## Where compliance checklists fall short

Compliance checklists can tell you whether a policy contains obvious broad permissions. They can tell you whether MFA is required. They can tell you whether a role was unused for 90 days.
Expand All @@ -86,8 +78,6 @@ They do not always tell you whether a compromised CI runner can mutate productio

The checklist verifies the shape of the policy. It does not verify the shape of the blast radius.

---

## Good IAM architecture starts with containment

Containment is a design choice made before the policy review begins:
Expand All @@ -104,8 +94,6 @@ This is slower than rubber-stamping a policy PR. It is also the work that actual

The safer default is easier to enforce when it is baked into the tooling. A small library like [iam-safe-defaults]({filename}iam-safe-defaults-fail-loud.md) can move the argument from "should this role have a permissions boundary" to "why is this role opting out of one," which is where you want the conversation to live.

---

## Trust made visible

Beyond access control, IAM is one of the main ways cloud architecture expresses trust. Every role, policy, permission boundary, service account, and deployment credential says something about what the system believes can safely happen.
Expand Down
10 changes: 0 additions & 10 deletions content/leap-second-chaos-2012.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,6 @@ On June 30, 2012, a **leap second** was inserted into atomic time via NTP to kee

What followed were 500 errors, high latency, and CPU usage spikes that crippled backend services.

---

## Why Did It Break?

Though seemingly minor, the leap second broke systems in subtle and severe ways:
Expand All @@ -28,8 +26,6 @@ Though seemingly minor, the leap second broke systems in subtle and severe ways:
- **Cloud Weakness Exposure:** An Amazon EC2 outage the day before had already left infrastructure strained. With fewer available instances, systems were more vulnerable when the leap second hit.
- **Limited Real-World Testing:** Simulating leap seconds under actual load, in full-stack distributed systems, proved nearly impossible. Pre-patch validations missed edge behavior.

---

## Real-World Impact

This bug hit nearly every high-scale Java-based system:
Expand All @@ -40,8 +36,6 @@ This bug hit nearly every high-scale Java-based system:

In many cases, the kernel *didn't fail*. The chaos came from how services processed time at runtime.

---

## Mitigation & Takeaways

### Immediate Fixes in 2012
Expand Down Expand Up @@ -70,8 +64,6 @@ sudo /etc/init.d/ntp start
- **Design for Temporal Anomalies**: Distributed systems should assume wall-clock time can regress, freeze, or desync, and gracefully degrade when it does.
- **Simulated Testing Isn't Enough**: Always combine synthetic load with chaos testing under unusual real-world conditions (e.g., leap seconds, DNS failures, NTP skew).

---

## Epilogue

The 2012 leap second chaos wasn't caused by incompetence. Many teams patched, prepared, and tested. But the leap second hit during degraded cloud capacity, exposed fragile JVM behavior, and stressed assumptions in time-sensitive code.
Expand All @@ -80,8 +72,6 @@ A single second exposed fault lines in the foundations of the modern internet.

In 2022, the CGPM (General Conference on Weights and Measures) voted to abolish leap seconds by 2035, largely driven by incidents like this one. Until then, the mitigations above remain essential for any system that touches wall-clock time.

---

**What other "just time" failures have caught you off guard in production? Let's share war stories.**

Related reading: the [208.5-day kernel bug]({filename}208-day-kernel-bug-lessons.md) - another case where a piece of time math, left unpatched, was a silent countdown to an outage.
Expand Down
12 changes: 0 additions & 12 deletions content/oom-killer-process-prioritization.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,8 +15,6 @@ In resource-constrained environments (especially virtual private servers, CI age

The OOM Killer uses heuristics (like memory usage and the `oom_score_adj` value) to select processes it deems less essential. But you don't have to leave that critical decision entirely to the kernel's default logic.

---

## The Incident

Years ago, I had to recover a VPS via remote console. A quick dive into `/var/log/messages` showed that the OOM Killer had struck, terminating critical services. The culprit? A perfect storm:
Expand All @@ -27,8 +25,6 @@ Years ago, I had to recover a VPS via remote console. A quick dive into `/var/lo

This combination overwhelmed system memory. Without process priority tuning, the OOM Killer started targeting processes based on its heuristics, which felt indiscriminate from an operational view as it even took down `sshd`.

---

## The Mitigation Strategy

You can significantly influence OOM Killer decisions using the `/proc/<pid>/oom_score_adj` setting for a process. This value ranges from -1000 to +1000. The kernel uses this score, combined with memory usage, to decide kill priority; a lower score makes the process less likely to be chosen relative to others.
Expand Down Expand Up @@ -98,8 +94,6 @@ done < "$CONFIG"
echo "OOM score adjustment complete."
```

---

## Running the Script

You can run this periodically via cron or on boot with systemd. For example:
Expand All @@ -126,8 +120,6 @@ sudo systemctl daemon-reload
sudo systemctl enable --now oom-adjuster.service
```

---

## Security Considerations

From a [DevSecOps perspective](devsecops-guide.html), OOM prioritization is a security hardening technique as much as an uptime one:
Expand All @@ -139,16 +131,12 @@ From a [DevSecOps perspective](devsecops-guide.html), OOM prioritization is a se

Misconfigured systems where critical daemons (like `iptables`, `auditd`, `sshd`, or VPN tunnels) are killed first expose themselves to avoidable downtime and security gaps.

---

## Modern Use Cases

- **Kubernetes nodes**: Influence OOM behavior via Quality of Service (QoS) classes (set by defining resource requests/limits in pod specs), or apply node-level tuning using methods like the script above for critical node components (e.g., kubelet, container runtime). For the pod-side of the same problem, see [Hardening Kubernetes Deployments]({filename}hardening-k8s.md).
- **CI/CD runners**: Protect build agents or essential runner services from being killed during resource-intensive test suites or concurrent builds.
- **Shared hosting / VPS**: Prioritize core services (web server, database, SSH) over potentially less critical user processes or background tasks.

---

## Conclusion

The OOM Killer is an essential part of the Linux kernel, but leaving process termination order purely to default heuristics can be risky in production. By strategically assigning `oom_score_adj` values based on business continuity and security priorities, you can significantly reduce recovery time and harden your systems against memory pressure scenarios.
Expand Down
Loading
Loading