diff --git a/.gitignore b/.gitignore
index cdf5437..8eea27e 100644
--- a/.gitignore
+++ b/.gitignore
@@ -9,6 +9,7 @@ __pycache__/
.Python
env/
venv/
+.venv/
ENV/
build/
develop-eggs/
@@ -37,3 +38,4 @@ wheels/
# Temporary files
*.log
*.tmp
+maintainers-theme-evidence/
diff --git a/docs/about.md b/docs/about.md
index c6d07d7..77112e7 100644
--- a/docs/about.md
+++ b/docs/about.md
@@ -1,25 +1,5 @@
# About
-Personal site. ORNL is my employer; nothing here is an official ORNL product or statement.
+I'm Joshua N. Grant, a geospatial systems architect at Oak Ridge National Laboratory, where I work on spatial data infrastructure, geospatial data warehouses, raster processing, and streaming systems. I came into software through plant-science research and data analysis, then worked in data engineering at ORNL and Bold Penguin before returning to ORNL in my current role. Outside work I build data tools, generative graphics, music systems, and small games — including PARQONAUT, NUMBRANE, and the other projects on [Projects](projects/index.md). This is my personal site and does not represent ORNL.
-## Work
-
-Geospatial systems architect at **Oak Ridge National Laboratory** — spatial data warehouses, raster processing, streaming geospatial pipelines, and the Postgres/Kafka/object-storage stack around them.
-
-**Bold Penguin** (2022–2023): data engineering, Prefect 2 migration, CI/CD, services on Kubernetes.
-
-**Oak Ridge National Laboratory** (2019–2022): ETL on Kubernetes and Airflow, containerized services, COVID tracking data systems.
-
-## Outside work
-
-[Projects](projects/index.md): PARQONAUT, dotfiles, NUMBRANE, games, MIDI, and the rest.
-
-## Education
-
-MS and BS, plant sciences, University of Tennessee. Software came later, through data analysis and pipelines.
-
-## Contact
-
-[jngrant@live.com](mailto:jngrant@live.com) · [GitHub](https://github.com/sempervent) · [LinkedIn](https://linkedin.com/in/joshuanagrant) · [Blog](https://notjustadatum.blogspot.com)
-
-[Contact & Collaboration](getting-started.md) for phone and mailing address.
+[Contact](getting-started.md)
diff --git a/docs/adr/0001-site-architecture.md b/docs/adr/0001-site-architecture.md
deleted file mode 100644
index 51a5bea..0000000
--- a/docs/adr/0001-site-architecture.md
+++ /dev/null
@@ -1,129 +0,0 @@
-# ADR-0001: Site Architecture — MkDocs + Material + Repo-as-Docs
-
-**Status**: Accepted
-
-**Date**: 2026-02-28
-
-**Deciders**: Joshua N. Grant (site owner)
-
-**Technical Story**: Build a fast, searchable, maintainable documentation site that supports both professional best practices and playful experimentation — without a CMS, a backend, or a build team.
-
----
-
-## Context
-
-This site is a personal technical portfolio and documentation hub. It has two distinct audiences with different needs:
-
-1. **Engineers evaluating patterns** — want stable reference material, architectural rationale, and production-ready examples they can adapt. They browse by topic; they don't read linearly.
-2. **Curious generalists** — stumble in via search or a link, want something interesting to read or build. They respond to personality and novelty.
-
-The site needed to satisfy both without becoming a bloated CMS or a hand-coded React app. Constraints:
-
-- Single author (no content team, no editorial workflow)
-- Content lives in Git (version control, diff-ability, PR review for future collaborators)
-- Must deploy to a free static host (GitHub Pages)
-- Must be fast to build, cheap to maintain, and not require frontend engineering to update
-- Must support full-text search without a search backend
-- Must scale to hundreds of pages without navigation collapse
-
----
-
-## Decision
-
-**MkDocs** with the **Material for MkDocs** theme, deployed to **GitHub Pages** from this repository, with content organized under `docs/` as plain Markdown files.
-
-### Content domains
-
-Three top-level content domains with distinct contracts:
-
-| Domain | Purpose | Structure contract |
-|---|---|---|
-| **Best Practices** | Stable reference: patterns, governance, architecture | Conceptual; no assumed "do this right now" task |
-| **Tutorials** | Task-oriented: step-by-step, copy-paste runnable | Prereqs → Steps → Verify → Troubleshoot |
-| **Just for Fun** | Experimental: creative, playful, technically rigorous | No format constraint; must be reproducible |
-
-### Navigation philosophy
-
-- Tabs for top-level domains (`navigation.tabs`)
-- Section grouping within tabs (`navigation.sections`)
-- Breadcrumbs on every page (`navigation.path`)
-- Long subsections (PostgreSQL, Docker) use nested subgroups — not flat 25-item lists
-- Every major section has an `index.md` landing page
-
-### Discovery mechanisms
-
-- **Tags**: small, controlled vocabulary (≤ 15 tags total); rendered on a `/tags` index page
-- **Search**: lunr-based full-text with symbol-aware separator (`[\s\-\_\.]+`); `search.suggest` and `search.share` enabled
-- **What's New**: manually curated `whats-new.md`; updated when new content lands; no automation
-- **See Also**: admonition blocks at the bottom of tutorial and best-practice pages; 2–3 links each; contextually adjacent only
-
-### Deployment
-
-GitHub Actions builds the site on push to `main` via `mkdocs gh-deploy`. No server. No CDN to configure. GitHub Pages serves static HTML.
-
----
-
-## Alternatives Considered
-
-### Hugo
-
-Fast at scale; single binary. Rejected: Go template syntax is a maintenance burden for a Markdown-first author. Material-equivalent theme quality requires significant front-end setup. No meaningful advantage at this site's scale.
-
-### Docusaurus (React)
-
-Strong MDX support; excellent for product docs with interactive components. Rejected: Node.js build pipeline; React authoring expected for advanced layouts; overkill for a site with no interactive components. Bundle size and build complexity add nothing here.
-
-### Sphinx (reStructuredText)
-
-Excellent for API reference documentation; first-class Python ecosystem. Rejected: reST markup is higher friction than Markdown for non-API content; theme ecosystem is weaker; search quality is inferior to Material's lunr integration.
-
-### GitBook / Notion / Confluence
-
-Managed SaaS tools with rich editors and collaboration features. Rejected: content is not Git-native; export lock-in; no control over URL structure; pricing; branding constraints. This site's content is code — it belongs in a code repository.
-
-### Single giant README or GitHub Wiki
-
-Minimal tooling; works for small projects. Rejected: no search, no nav structure, no code block copy buttons, no dark mode, no tagging, no per-section index pages. Collapses at scale.
-
----
-
-## Consequences
-
-### Positive
-
-- Zero server infrastructure; zero hosting cost
-- Content is fully version-controlled; blame, diff, PR review all work
-- MkDocs builds in < 15 seconds even at 300+ pages
-- Material provides search, dark mode, code copy, admonitions, and tabs out of the box — no custom front-end work
-- URL structure is stable (file path = URL); easy to cross-link and share
-- `navigation.path` breadcrumbs solve "where am I?" without custom JS
-
-### Negative / Trade-offs
-
-- Nav hierarchy lives in `mkdocs.yml` YAML; verbose at scale and error-prone to maintain manually
-- No dynamic content (no comments, no user accounts, no live search suggestions beyond lunr)
-- Material plugin upgrades occasionally break config; `requirements.txt` must be pinned and maintained
-- "What's New" is manual — if the author doesn't update it, it goes stale
-- Tags are manual frontmatter — no auto-tagging
-
-### Follow-ups / Future Work
-
-- [ ] Automate `whats-new.md` generation from `git log` via a pre-commit hook or CI step
-- [ ] Add `mkdocs-redirects` entries when pages are moved (plugin already installed)
-- [ ] Evaluate `mkdocs-awesome-pages-plugin` to reduce YAML verbosity in nav
-- [ ] Add a `projects/` section with structured project cards
-- [ ] Revisit tag vocabulary as content grows; target ≤ 20 tags total
-
----
-
-## Operational Notes (guide future decisions)
-
-**URL preservation**: Never move a page without adding a redirect entry in `mkdocs.yml` via the `redirects` plugin. The plugin is already installed (`mkdocs-redirects`). A broken link from an external site is permanent damage.
-
-**Nav sanity**: Any subsection with more than 12 items should be split into named subgroups. The PostgreSQL section (26 items) is the canonical example — it was split into Core & Design, Performance & Operations, Advanced Features, and Integration & Deployment.
-
-**Cross-linking policy**: Every tutorial gets a `!!! tip "See also"` admonition with 2–3 links to contextually adjacent pages. Best-practice pages link to at least one tutorial that implements the pattern. Links must be relative and verified at build time.
-
-**Tag policy**: Tags are kept to a small, stable vocabulary. Before adding a new tag, check whether an existing tag covers it. Tags should reflect technology domains (`geospatial`, `postgresql`, `docker`), not content types or difficulty levels.
-
-**Content domain assignment**: When a page is ambiguous between Best Practices and Tutorials, assign it to the section that matches the primary reader intent. A page that readers open to *understand* goes in Best Practices. A page that readers open to *do something right now* goes in Tutorials. Cross-link between them.
diff --git a/docs/adr/0002-why-mkdocs-material.md b/docs/adr/0002-why-mkdocs-material.md
deleted file mode 100644
index e7917c5..0000000
--- a/docs/adr/0002-why-mkdocs-material.md
+++ /dev/null
@@ -1,83 +0,0 @@
-# ADR-0002: Use MkDocs + Material Theme for Documentation Site
-
-**Status**: Accepted
-
-**Date**: 2024-01-15
-
-**Deciders**: Joshua N. Grant
-
-**Tags**: documentation, tooling, site
-
----
-
-## Context
-
-This site needed a documentation framework that could handle several hundred markdown files, support a rich nav hierarchy, provide fast full-text search, and deploy cleanly to GitHub Pages with zero server infrastructure.
-
-The author writes documentation in Markdown and wanted to keep authoring in Markdown — not learn a template DSL or manage a database. The site should build in CI, not require a local build step for every page edit, and look professional without heavy front-end engineering.
-
-## Decision
-
-Use **MkDocs** (the static site generator) with the **Material for MkDocs** theme.
-
-## Options Considered
-
-### Option 1: MkDocs + Material (chosen)
-
-**Pros:**
-- Pure Markdown authoring; no shortcodes or template syntax needed
-- Material theme provides tabs, search, code copy, dark mode, admonitions out of the box
-- GitHub Pages deployment via `gh-pages` branch is first-class and well-documented
-- `navigation.tabs`, `navigation.path`, `search.suggest`, `toc.integrate` cover all UX needs
-- Plugin ecosystem: `git-revision-date-localized`, `tags`, `section-index`, `redirects`
-- Active development; Material 9.x is stable and widely used
-
-**Cons:**
-- Python dependency (minor; managed via `requirements.txt`)
-- No JavaScript server-side rendering; purely static
-- Large nav trees can produce long YAML in `mkdocs.yml`
-
-### Option 2: Hugo
-
-**Pros:**
-- Faster builds at very large scale (thousands of pages)
-- Single binary, no Python required
-
-**Cons:**
-- Go template syntax is complex for non-Go authors
-- Fewer out-of-the-box documentation UX patterns
-- Material-equivalent theme quality requires more setup work
-
-### Option 3: Docusaurus (React)
-
-**Pros:**
-- First-class MDX support (Markdown + JSX)
-- Strong ecosystem for product documentation
-
-**Cons:**
-- Node.js build pipeline; heavier dependency surface
-- React component authoring expected for advanced layouts
-- Overkill for a personal documentation site with no interactive components
-
-## Rationale
-
-MkDocs + Material is the most productive choice for a single-author, Markdown-first documentation site. Material provides the full UX feature set this site needs — tabs, search, dark mode, code copy, admonitions — with zero custom front-end work. Hugo is faster at extreme scale but adds template complexity that doesn't pay off here. Docusaurus is designed for product docs with interactive React components, which this site does not need.
-
-## Consequences
-
-### Positive
-- Zero front-end engineering required to maintain the site
-- Full-text search works out of the box with Material's lunr integration
-- GitHub Actions CI builds and deploys in < 2 minutes
-
-### Negative
-- Nav hierarchy is expressed in `mkdocs.yml` YAML, which becomes verbose at large scale
-- Material requires specific plugin versions; upgrades occasionally break config
-
-### Neutral / Trade-offs
-- Python version pinning matters for reproducible builds; `requirements.txt` must be maintained
-
-## Related Documents
-
-- [Creating MkDocs GitHub Site](../tutorials/quick-start/creating-mkdocs-github-site.md)
-- [ADR and Technical Decision Governance](../best-practices/architecture-design/adr-decision-governance.md)
diff --git a/docs/adr/0003-why-best-practices-vs-tutorials.md b/docs/adr/0003-why-best-practices-vs-tutorials.md
deleted file mode 100644
index 62c0591..0000000
--- a/docs/adr/0003-why-best-practices-vs-tutorials.md
+++ /dev/null
@@ -1,87 +0,0 @@
-# ADR-0003: Separate "Best Practices" from "Tutorials" as Distinct Top-Level Sections
-
-**Status**: Accepted
-
-**Date**: 2024-02-01
-
-**Deciders**: Joshua N. Grant
-
-**Tags**: information-architecture, documentation
-
----
-
-## Context
-
-As the site grew beyond 50 pages, two distinct types of content emerged with different authoring patterns and reader goals:
-
-1. **Conceptual/opinionated guides** — "How to think about X", pattern references, governance frameworks, architectural principles. Readers consult these when making decisions. They rarely run commands from these pages.
-
-2. **Step-by-step implementations** — "How to do X right now", with runnable commands, copy-paste configs, and clear success criteria. Readers follow these linearly to accomplish a specific task.
-
-Mixing both types under a single nav section forced readers to guess which kind of content they were getting before clicking.
-
-## Decision
-
-Maintain **two top-level sections**:
-
-- **Best Practices** — patterns, architectures, governance, opinionated guides. No assumed prerequisite of "do this exact thing right now."
-- **Tutorials** — step-by-step, task-oriented, runnable. Has clear Prerequisites, Steps, Verify, and Troubleshoot structure.
-
-Cross-link between the two wherever a tutorial implements a best practice.
-
-## Options Considered
-
-### Option 1: Single "Documentation" section (flat)
-
-**Pros:**
-- Simpler nav structure
-- No category ambiguity
-
-**Cons:**
-- Readers can't tell if a page is "read to understand" vs "follow to implement"
-- Harder to maintain consistent page structure when types are mixed
-- Harder to recommend "start here" paths for different reader goals
-
-### Option 2: Two sections: Best Practices + Tutorials (chosen)
-
-**Pros:**
-- Mental model matches reader intent: "I want to understand" vs "I want to implement"
-- Consistent page templates per section (best practice has Context/Pattern/Trade-offs; tutorial has Prereqs/Steps/Verify)
-- Cross-links between sections surface the relationship between concept and implementation
-
-**Cons:**
-- Some content genuinely fits both categories (e.g., a "best practices" page with runnable examples)
-- Duplication risk if a best-practice and tutorial cover the same topic without linking to each other
-
-### Option 3: Tag-based (single section, discover by tag)
-
-**Pros:**
-- No category assignment needed
-- Tags can span multiple dimensions
-
-**Cons:**
-- Requires robust tagging from day one (which doesn't exist at authoring time)
-- Harder to build a coherent "recommended reading order" without section structure
-
-## Rationale
-
-The two-section model matches the dominant reader intents on this site. Engineers who know what they want to build open Tutorials. Engineers who are evaluating approaches open Best Practices. Keeping them separate preserves the ability to enforce different page templates per section and recommend distinct "Start Here" paths per audience.
-
-## Consequences
-
-### Positive
-- Clear authoring expectations per section
-- Separate "Start Here" recommendations per section (technical decision-makers vs implementors)
-- Cross-links between sections create a richer content graph
-
-### Negative
-- Some pages are genuinely ambiguous in category (handled by placing in the section that best matches the primary reader intent, then cross-linking)
-- Nav depth increases because each section has its own subsections
-
-### Neutral / Trade-offs
-- The "Just for Fun" section under Tutorials is an intentional exception: it's neither purely conceptual nor purely task-driven, but it belongs in Tutorials because readers follow it to build something
-
-## Related Documents
-
-- [ADR-0002: Why MkDocs + Material](0002-why-mkdocs-material.md)
-- [Technical Documentation](../documentation.md) — site structure and navigation
diff --git a/docs/adr/0004-why-just-for-fun-exists.md b/docs/adr/0004-why-just-for-fun-exists.md
deleted file mode 100644
index 3c8cd0f..0000000
--- a/docs/adr/0004-why-just-for-fun-exists.md
+++ /dev/null
@@ -1,94 +0,0 @@
-# ADR-0004: Include a "Just for Fun" Section in a Technical Documentation Site
-
-**Status**: Accepted
-
-**Date**: 2024-03-10
-
-**Deciders**: Joshua N. Grant
-
-**Tags**: documentation, information-architecture
-
----
-
-## Context
-
-A personal technical documentation site that presents only enterprise patterns and production best practices risks becoming indistinguishable from thousands of other engineering blogs. It signals "I can execute corporate patterns" but not "I explore ideas at the edges of what's technically possible."
-
-The site author builds projects that don't fit neatly into "Best Practices" or "Tutorials" — things like:
-
-- A Raspberry Pi sample library server with USB MIDI live audition
-- Recursive cathedral generators from L-system grammars
-- Metrics sonification via SuperCollider + OSC
-- WebGL generative art from PostGIS raster data
-
-These projects are technically rigorous, reproducible, and often more instructive about system design than a standard tutorial — but their primary motivation is curiosity, not production deployment.
-
-## Decision
-
-Create a **"Just for Fun" subsection** under Tutorials. Include creative, experimental, and exploratory projects that are:
-
-- Technically reproducible (not just concept posts)
-- Non-trivial in implementation (demonstrate real engineering)
-- Built for intrinsic interest, not business requirements
-
-Do not artificially restrict the section to specific technology areas. Let it grow organically.
-
-## Options Considered
-
-### Option 1: Exclude creative projects entirely
-
-**Pros:**
-- Site maintains a purely professional tone
-- No reader confusion about what the site is "for"
-
-**Cons:**
-- Loses the most distinctive content on the site
-- The best projects — the ones that demonstrate creative problem-solving — get no home
-- Portfolio signals only "follows conventional patterns," not "thinks originally"
-
-### Option 2: Separate blog/personal site for fun projects
-
-**Pros:**
-- Clean separation of concerns
-- Professional docs site stays focused
-
-**Cons:**
-- Two sites to maintain; cross-traffic is lost
-- The technical depth of these projects is documentation-worthy, not blog-post-worthy
-- Discoverability drops: readers who arrive for best practices never discover the creative work
-
-### Option 3: "Just for Fun" section within the documentation site (chosen)
-
-**Pros:**
-- Readers who arrive for serious content discover unexpected creative depth
-- The same documentation framework serves both content types
-- Cross-links between fun projects and related best practices reinforce conceptual connections (e.g., Pi sample server → Docker best practices)
-- Signals that the author is a whole engineer, not just a pattern-follower
-
-**Cons:**
-- Some readers may find the tone inconsistent with the rest of the site
-- Section needs its own voice: playful but technically precise, not corporate
-
-## Rationale
-
-The "Just for Fun" section is a differentiator, not a distraction. Every page in the section demonstrates real engineering — recursive L-systems, WebAudio API internals, MIDI protocol handling, SQLite query optimization. The playful framing makes these approachable; the technical depth makes them credible. Excluding this content would make the site less interesting and less representative of how the author actually works.
-
-## Consequences
-
-### Positive
-- Site has a unique character that distinguishes it from generic engineering blogs
-- Creative projects get the same documentation quality as production content
-- Cross-links from fun projects into best practices and tutorials create unexpected discovery paths
-
-### Negative
-- Section requires active curation to maintain quality bar (reproducible, non-trivial)
-- Voice calibration is harder: playful but not cringe; experimental but not hand-wavy
-
-### Neutral / Trade-offs
-- "Managing People in Software Development" is in this section despite being non-technical — it fits the spirit of "unexpected but useful" content that defines the section
-
-## Related Documents
-
-- [Just for Fun Overview](../tutorials/just-for-fun/index.md)
-- [ADR-0003: Best Practices vs Tutorials](0003-why-best-practices-vs-tutorials.md)
-- [Technical Documentation](../documentation.md) — site structure and navigation
diff --git a/docs/adr/0005-esp32-section-architecture.md b/docs/adr/0005-esp32-section-architecture.md
deleted file mode 100644
index 73429f3..0000000
--- a/docs/adr/0005-esp32-section-architecture.md
+++ /dev/null
@@ -1,96 +0,0 @@
-# ADR 0005: Establish ESP32 & Embedded Systems as a First-Class Documentation Section
-
-- **Status**: Accepted
-- **Date**: 2026-02-26
-- **Deciders**: mate (owner)
-- **Technical Story**: ESP32-based projects have accumulated across the site without a coherent home. Embedded systems combine hardware, firmware, electrical safety, and security in ways that require a dedicated, structured documentation pillar — not a footnote under "just for fun."
-
----
-
-## Context
-
-The site already has stable pillars for server-side concerns (PostgreSQL, Docker, Python, Rust, Go) and creative projects (Just for Fun). But embedded systems occupy a gap:
-
-- Hardware and software decisions are deeply coupled — you cannot explain firmware without explaining the circuit.
-- Electrical safety is non-negotiable and must be part of the documentation, not an afterthought.
-- ESP32 projects span best practices (architecture, power, security) *and* tutorials (capstone builds), requiring both documentation modes.
-- Without a dedicated section, embedded content risks scattering across "Just for Fun" (too casual), "Docker" (wrong stack), or nowhere at all.
-
-Specific drivers:
-1. Growing number of ESP32 personal projects (sensor nodes, room controllers, art frames).
-2. Need to document safety-first patterns that apply across all embedded projects, not just one tutorial.
-3. Desire to distinguish *durable practices* (deep sleep design, ISR safety, power budgets) from *playful experiments* (e-ink art frames).
-4. Embedded docs must cross-link hardware ↔ firmware ↔ security — a flat structure cannot represent these dependencies well.
-
----
-
-## Decision
-
-**Create `Best Practices → 🔌 Embedded Systems & ESP32`** as a first-class section at the same level as Python, Docker, and PostgreSQL, with the following structure:
-
-| Content type | Location |
-|---|---|
-| Durable patterns (architecture, power, safety, security, sensors, display) | `docs/best-practices/esp32/` |
-| Applied project tutorials | `docs/tutorials/embedded/` |
-| Architectural rationale (this ADR) | `docs/adr/` |
-
-**Safety as mandatory content**: every best practice document in this section must address safety implications — electrical, power, or security — as a primary concern, not a sidebar.
-
-**Cross-linking policy**: every tutorial references the relevant best practice pages via `!!! tip "See also"` admonitions. Every best practice page links back to at least one applied tutorial.
-
-**Naming convention**: `kebab-case.md`, `esp32-` prefix for ESP32-specific docs, `embedded/` for the tutorial subdirectory.
-
----
-
-## Alternatives Considered
-
-### Keep ESP32 under Just for Fun only
-
-Rejected. "Just for Fun" signals low stakes and experimentation. Electrical safety documentation, power budget analysis, and OTA security do not belong under that framing. A reader looking for grounded hardware guidance would not find it.
-
-### Generic "Hardware" catch-all section
-
-Rejected. "Hardware" is too broad. The site has no Raspberry Pi HAT tutorials, no PCB design content, no STM32 content — a generic section would be an empty promise. ESP32 is the concrete, bounded scope that exists today.
-
-### One giant "ESP32 reference" page
-
-Rejected. A monolithic page cannot represent the relationship between topics (e.g., "power gating depends on both the circuit design page and the firmware architecture page"). Separate, cross-linked pages enable navigation by reader intent.
-
-### Separate repo
-
-Rejected. The value of co-locating embedded content with Docker, Go, and security content is cross-pollination: a reader learning about OTA updates should easily reach the general secrets management guide. Splitting the repo breaks that graph.
-
----
-
-## Consequences
-
-### Positive
-
-- Clear separation between durable patterns (best practices) and applied projects (tutorials) — consistent with [ADR-0003](0003-why-best-practices-vs-tutorials.md)
-- Safety documentation is guaranteed visible and linked from every project page
-- Embedded becomes a stable, growing pillar alongside server-side content
-- Cross-links between embedded tutorials and Docker/Security/Kubernetes content are natural and short
-- New contributors have a clear template to follow for future embedded topics
-
-### Negative / Trade-offs
-
-- Increases nav surface area by ~10 entries; sidebar grows slightly longer
-- Structural complexity: readers must understand that "programming architecture" is in best-practices, while "RF controller" is in tutorials
-- Requires maintaining cross-links as new content is added (manageable with the See Also convention)
-
-### Follow-ups
-
-- [ ] Add LoRa best practices page (SX1276/SX1278, spreading factor, duty cycle limits)
-- [ ] Add power electronics section (MOSFET switches, LiPo charging circuits, solar input)
-- [ ] Add printable ESP32 safety checklist (HTML/PDF export from Markdown)
-- [ ] Add ESP32-S3 and ESP32-C3 notes (USB native, RISC-V, BLE 5.0 differences)
-- [ ] Consider an `embedded/` best-practices subdirectory if non-ESP32 MCUs are added (STM32, RP2040)
-
----
-
-## Notes
-
-- Follow the [URL preservation policy from ADR-0001](0001-site-architecture.md): if any embedded page is ever moved, add a redirect.
-- Nav rule: if the embedded section exceeds 12 items, group into sub-sections (e.g., "Core", "Communication", "Power").
-- Tag policy: use `embedded` and `esp32` tags consistently; add `security` or `performance` where appropriate.
-- This ADR does not cover non-ESP32 microcontrollers. A future ADR should decide whether to expand the section or create a sibling section.
diff --git a/docs/adr/0006-embedded-expansion-lora-power.md b/docs/adr/0006-embedded-expansion-lora-power.md
deleted file mode 100644
index 86f5784..0000000
--- a/docs/adr/0006-embedded-expansion-lora-power.md
+++ /dev/null
@@ -1,79 +0,0 @@
-# ADR 0006: Expand Embedded Section — LoRa, Power Electronics, and Variant Notes
-
-- **Status**: Accepted
-- **Date**: 2026-02-26
-- **Deciders**: mate (owner)
-- **Supersedes**: N/A — extends [ADR-0005](0005-esp32-section-architecture.md)
-
----
-
-## Context
-
-[ADR-0005](0005-esp32-section-architecture.md) established the `best-practices/esp32/` section and committed to follow-ups: LoRa best practices, power electronics, a printable safety checklist, and ESP32-S3/C3 notes.
-
-Implementation revealed a structural decision: **Power Electronics** does not fit cleanly under `esp32/`. The content (MOSFET switching, LiPo charging circuits, solar, buck vs LDO) applies to any embedded microcontroller — not only ESP32. It belongs at a higher level of abstraction.
-
-This ADR formalizes:
-1. The creation of `best-practices/embedded/` as a cross-cutting hardware folder
-2. The placement of LoRa, printable checklist, and S3/C3 notes under `esp32/`
-3. The nav grouping for `🔋 Power Electronics & Embedded Hardware`
-
----
-
-## Decision
-
-**Create `docs/best-practices/embedded/`** as a sibling to `best-practices/esp32/`, `best-practices/docker-infrastructure/`, etc.
-
-| Content | Location | Rationale |
-|---|---|---|
-| Power Electronics (MOSFETs, LiPo, solar, buck/LDO) | `best-practices/embedded/` | Platform-agnostic hardware concerns |
-| LoRa best practices (SX127x) | `best-practices/esp32/` | ESP32-centric driver and firmware guidance |
-| Printable safety checklist | `best-practices/esp32/` | ESP32-specific checklist; referenceable from hardware doc |
-| ESP32-S3/C3 notes | `best-practices/esp32/` | Variant comparison for ESP32 ecosystem |
-
-**Nav grouping**: `🔋 Power Electronics & Embedded Hardware` appears in mkdocs.yml alongside `🔌 Embedded Systems & ESP32` under Best Practices.
-
-**Cross-linking policy**: the `embedded/` folder links back to `esp32/` for firmware context; `esp32/` pages link to `embedded/` for hardware context.
-
----
-
-## Alternatives Considered
-
-### Put power electronics under `esp32/`
-
-Rejected. A page titled "Power Electronics for ESP32" is still useful to Raspberry Pi users, STM32 users, and Arduino users. Placing it under `esp32/` would mislead the IA and make it harder to find for non-ESP32 readers. The content's value is platform-agnostic.
-
-### Create a `hardware/` folder instead of `embedded/`
-
-Rejected. "Hardware" is a broader signal than intended. "Embedded" is more accurate — it connotes electronics that host a microcontroller, which is precisely the scope.
-
-### Single `embedded/` folder replacing `esp32/`
-
-Rejected. The `esp32/` folder contains firmware-level guidance (FreeRTOS, ISRs, deep sleep API) that is ESP32-specific. Merging would dilute the precision of that content.
-
-### No new folder — just add power electronics to `esp32/`
-
-Rejected — see first alternative.
-
----
-
-## Consequences
-
-### Positive
-
-- Power electronics guidance is reachable by readers working on any embedded platform
-- Clear separation: `esp32/` = ESP32 firmware + peripherals; `embedded/` = universal hardware concerns
-- Sets a precedent for future MCU-agnostic content (RP2040, STM32, AVR) to have a home without needing to dilute the ESP32 section
-- Printable safety checklist is directly reachable from the nav and from hardware build references
-
-### Negative / Trade-offs
-
-- Nav surface area increases by one section and two entries
-- A reader unfamiliar with the split may look in the wrong section first — mitigated by cross-links
-
-### Follow-ups
-
-- [ ] Add RP2040 / Raspberry Pi Pico notes to `embedded/` if Pico projects are documented
-- [ ] Add STM32 power notes if STM32 tutorials are added
-- [ ] Consider `embedded/antenna-and-rf-layout.md` for shared RF layout guidance across LoRa, WiFi, BLE
-- [ ] Evaluate whether the printable checklist would benefit from MkDocs print CSS customization
diff --git a/docs/adr/0007-deep-dives-section.md b/docs/adr/0007-deep-dives-section.md
deleted file mode 100644
index c326512..0000000
--- a/docs/adr/0007-deep-dives-section.md
+++ /dev/null
@@ -1,95 +0,0 @@
-# ADR 0007: Introduce "Deep Dives" Section for Analytical Long-Form Content
-
-- **Status**: Accepted
-- **Date**: 2026-02-26
-- **Deciders**: mate (owner)
-- **Related**: [ADR-0003](0003-why-best-practices-vs-tutorials.md) (Best Practices vs Tutorials distinction)
-
----
-
-## Context
-
-The site has two content modes: prescriptive (Best Practices) and procedural (Tutorials). Both are optimized for practical output — a reader leaves with a configuration, a working build, or a mental model of "what to do."
-
-Some subjects resist this framing. They are not best solved by prescription or procedure. They require comparison, historical context, regulatory analysis, threat modeling, and architectural argumentation. Examples already emerging in the content:
-
-- LoRaWAN vs raw LoRa: a decision shaped by spectrum law, security assumptions, infrastructure cost, and fleet size
-- MQTT vs CoAP vs HTTP for IoT: protocol choice determined by transport semantics, not code examples
-- Scratch vs distroless: not a tutorial topic; a philosophical and operational stance
-
-Attempting to contain these analyses in a Best Practice page produces either a superficial treatment or an overwhelming document that violates the section's prescriptive purpose. Containing them in a tutorial adds irrelevant instructional scaffolding to what is fundamentally an argument.
-
-The existing section structure does not provide a natural home for analytical, argument-driven content.
-
----
-
-## Decision
-
-Create `docs/deep-dives/` as a top-level content section with its own nav tab: **Deep Dives**.
-
-Content placed here must satisfy all of the following:
-
-- **Comparative**: examines two or more approaches, protocols, architectures, or tools in genuine tension
-- **Analytical**: presents reasoning and evidence, not instructions
-- **Argument-driven**: arrives at a defensible position or decision framework
-- **Less code-heavy**: code snippets are illustrative, never the primary content
-- **Intentional**: a deep dive is written because the topic *demands* depth — not as a way to expand the site
-
-Content that is prescriptive belongs in Best Practices. Content that is procedural belongs in Tutorials. Content that is analytical and comparative belongs in Deep Dives.
-
-**Placement**: top-level nav between Best Practices and Tutorials. Deep Dives is neither a reference nor a guide — it is a section a reader visits to think, not to do.
-
----
-
-## Alternatives Considered
-
-### Overload Best Practices with long-form comparisons
-
-Rejected. Best Practices are prescriptive by contract (see ADR-0003). A 4000-word analysis of LoRaWAN vs raw LoRa is not a best practice — it is a decision support document. Mixing the two dilutes the purpose of the Best Practices section and makes it harder to navigate.
-
-### Write excessively long tutorials
-
-Rejected. Tutorials are procedural. Adding comparative analysis to a tutorial inflates it and frustrates readers who came to build something. A reader who wants to understand the LoRaWAN security model is not the same reader who wants firmware wiring instructions.
-
-### External blog (separate from the site)
-
-Rejected. A separate blog creates friction: different URL, different deployment, different search index, different navigation context. The goal is to keep high-signal content co-located with the technical reference material it references, enabling cross-linking and unified search.
-
-### One overloaded "analysis" page per topic area
-
-Rejected. A single "Embedded Systems Analysis" page containing every comparative discussion would be incoherent and unsearchable. Individual, focused deep dives are more discoverable and maintainable.
-
----
-
-## Consequences
-
-### Positive
-
-- Provides a legitimate outlet for analytical depth without corrupting the purpose of existing sections
-- Raises the intellectual surface area of the site — a reader can move from "how to configure a LoRa transmitter" (Best Practices) to "why LoRaWAN might be wrong for your deployment" (Deep Dives) within the same site
-- Encourages higher-quality writing: the analytical mode demands more careful argumentation than a checklist
-- Cross-linking between Deep Dives and Best Practices / Tutorials creates a richer content graph
-
-### Negative / Trade-offs
-
-- Increases nav surface area by one top-level tab
-- Requires editorial discipline: Deep Dives must remain intentional and rare — if every topic gets a deep dive, the section loses its signal
-- Analytical writing takes longer to produce and review than prescriptive writing
-
----
-
-## Operational Rules
-
-1. A Deep Dive is created only when the topic genuinely resists prescription or procedure
-2. Each Deep Dive must cross-link to at least one Best Practice or Tutorial page
-3. Deep Dives do not include step-by-step wiring instructions, setup commands, or "how to install X"
-4. Style: analytical, comparative, journalistic — not tutorial, not reference
-5. If a Deep Dive eventually produces a clear prescription, extract that prescription into a Best Practice page and link back
-
----
-
-## Follow-ups
-
-- [ ] Add a style guide for Deep Dives (tone, structure, code usage policy) as a meta-page in the section
-- [ ] Candidate topics: MQTT vs CoAP vs HTTP for IoT, Scratch vs distroless container philosophy, ADR-driven architecture decision culture, LoRaWAN vs Zigbee vs Thread for smart home
-- [ ] Evaluate whether Deep Dives warrant their own RSS feed or tag filtering
diff --git a/docs/adr/0008-expand-deep-dives-section.md b/docs/adr/0008-expand-deep-dives-section.md
deleted file mode 100644
index 2d11f29..0000000
--- a/docs/adr/0008-expand-deep-dives-section.md
+++ /dev/null
@@ -1,100 +0,0 @@
----
-tags:
- - adr
- - documentation
----
-
-# ADR 0008: Expand Deep Dives as a First-Class Analytical Section
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Author**: Joshua N. Grant
-
----
-
-## Context
-
-The site's Deep Dives section was established in ADR 0007 with a single entry (LoRaWAN vs Raw LoRa) as a proof of concept for long-form analytical content. That entry demonstrated the format's value: structured comparative analysis, trade-off matrices, historical context, and decision frameworks — distinct from both tutorials and best-practice checklists.
-
-Several recurring content needs within this site's scope are poorly served by existing content types:
-
-- **Format and storage decisions** involve trade-offs that span years of ecosystem evolution and cannot be reduced to a checklist.
-- **System selection decisions** (databases, engines, protocols) require comparative analysis across dimensions that change with organizational scale.
-- **Infrastructure philosophy** (container base images, infrastructure automation) involves trade-offs between correctness, operational complexity, and security that are contextual rather than prescriptive.
-
-These topics warrant a separate content type with its own structural guardrails.
-
----
-
-## Decision
-
-Expand the Deep Dives section from one entry to eight, adding:
-
-1. **Parquet vs CSV vs ORC vs Avro** — columnar storage theory, compression, predicate pushdown, ecosystem analysis
-2. **Lakehouse vs Warehouse vs Database** — historical arc, table format emergence, governance complexity, cost structures
-3. **DuckDB vs PostgreSQL vs Spark** — execution models, data gravity, hybrid architecture
-4. **Geospatial File Format Choices** — raster vs vector, Cloud-Optimized GeoTIFF, GeoParquet emergence
-5. **Container Base Image Philosophy** — glibc vs musl, attack surface analysis, operational debugging cost
-6. **Infrastructure as Code vs GitOps** — control plane models, drift detection, hybrid governance
-7. **MQTT vs HTTP in IoT Systems** — connection models, power consumption, broker centralization, security
-
-These entries are organized under four sub-sections in the navigation:
-
-- Data Formats & Storage
-- Data Systems & Architecture
-- Infrastructure & Automation
-- Embedded & Radio
-
----
-
-## Why Analytical Essays Deserve Structural Separation
-
-Best practices pages answer the question: "given that I am doing X, how should I do it well?" Tutorials answer: "how do I accomplish X step by step?"
-
-Deep dives answer a different question: "should I be doing X at all, and under what conditions does X beat Y?" This requires a different structure — comparative analysis, historical context, trade-off matrices — and a different tone: analytical and opinionated rather than prescriptive and procedural.
-
-Mixing these content types in the same section would dilute both. Best practices readers want concrete guidance; deep dive readers want structured argument. Structural separation preserves the value of each.
-
----
-
-## Guardrails Against Content Sprawl
-
-Deep dives carry an expansive mandate ("analytical exploration") that can attract poorly scoped content. The following criteria apply to all deep dive candidates:
-
-**A deep dive is appropriate when**:
-- The topic involves a genuine architectural or technology selection decision with non-obvious trade-offs
-- The decision consequences extend across months or years of system operation
-- Comparative analysis across at least two options is possible and meaningful
-- The topic cannot be adequately addressed in a best-practice page without becoming a comparative essay
-
-**A deep dive is not appropriate when**:
-- The correct answer is well-established and non-controversial
-- The topic is better served by a tutorial (step-by-step) or best-practice (how-to-well)
-- The topic is too narrow to warrant the deep dive structure (decision framework, trade-off matrix, historical context)
-- The topic duplicates an existing deep dive without adding a meaningfully different perspective
-
----
-
-## Criteria for New Deep Dives
-
-New deep dives should meet all of the following before being added:
-
-1. **Genuine tension**: two or more approaches with legitimate use cases — not a comparison where one option is obviously correct
-2. **Consequential decision**: a choice that affects system architecture, data portability, operational complexity, or long-term maintainability
-3. **Analytical completeness**: the entry must include a historical or contextual arc, at least one comparative matrix, ASCII diagram(s) where clarifying, and a decision framework section
-4. **Appropriate length**: 1,000–3,000 words; shorter entries should be best practices; longer entries should be split into multiple deep dives
-
----
-
-## Consequences
-
-The Deep Dives section becomes a stable, growing analytical corpus alongside tutorials and best practices. Navigation is organized by domain (formats, systems, infrastructure, embedded) rather than by chronological addition. Cross-links between related deep dives and related tutorials/best practices are maintained as part of the authoring checklist.
-
-The risk of content sprawl is managed by the criteria above. The ADR process itself (this document) serves as the governance mechanism for section-level structural changes.
-
----
-
-## Related Documents
-
-- [ADR 0007: Deep Dives Section](0007-deep-dives-section.md) — original decision to create the section
-- [ADR 0003: Best Practices vs Tutorials](0003-why-best-practices-vs-tutorials.md) — the content type distinction that deep dives extend
diff --git a/docs/adr/0009-deep-dives-expansion.md b/docs/adr/0009-deep-dives-expansion.md
deleted file mode 100644
index ee7e99e..0000000
--- a/docs/adr/0009-deep-dives-expansion.md
+++ /dev/null
@@ -1,91 +0,0 @@
----
-tags:
- - adr
- - documentation
----
-
-# ADR 0009: Expand Deep Dives into Core Analytical Layer
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Author**: Joshua N. Grant
-
----
-
-## Context
-
-ADR 0008 formalized the expansion of Deep Dives from one entry (LoRaWAN vs Raw LoRa) to eight entries, establishing structural sub-sections and criteria for new entries. This document addresses a further expansion: four additional deep dives covering data pipeline reliability, observability architecture, embedded platform selection, and workflow orchestration philosophy.
-
-The expansion brings the Deep Dives section to twelve entries across five sub-sections (Data Formats & Storage, Data Systems & Architecture, Operations & Reliability, Infrastructure & Automation, Embedded & Radio). This constitutes a material increase in scope — the section is no longer a collection of isolated analytical essays but an interconnected analytical corpus with deliberate cross-linking and a consistent epistemic standard.
-
----
-
-## Decision
-
-Accept the following four deep dives into the corpus:
-
-1. **Why Most Data Pipelines Fail** — organizational and architectural failure modes, schema drift, ownership ambiguity, the CI/CD gap in data systems, and survivable architectural patterns
-2. **Observability vs Monitoring** — monitoring vs observability semantics, three pillars analysis, cardinality economics, data pipeline observability, cultural factors, tooling landscape
-3. **ESP32 vs Raspberry Pi** — microcontroller determinism vs Linux convenience, power profiles, security surface, ecosystem maturity, and deployment decision matrix
-4. **Prefect vs Airflow** — DAG-first vs Python-native orchestration philosophy, execution model architecture, failure semantics, scaling models, and operational burden
-
-Add a new sub-section to the nav: **Operations & Reliability** (currently containing the Observability vs Monitoring deep dive, with room for additional entries in the reliability and incident management domain).
-
----
-
-## Why These Are Not Blog Posts
-
-Blog posts are timely, personal, and conversational. They may be opinionated without systematic justification, and they do not typically include formal decision frameworks or comparative analyses.
-
-Deep Dives in this corpus are:
-
-- **Structured**: every entry follows a consistent schema (thesis, context, analysis, comparison, decision framework) that enables comparative reading across entries
-- **Durable**: the analysis is written to remain relevant across years, not weeks; where technology changes rapidly, the structural principles rather than specific tools are the primary focus
-- **Referenced**: cross-links to related deep dives, best practices, and tutorials make each entry part of a navigable analytical corpus
-- **Analytical**: positions are taken and justified with evidence and logical argument, not asserted as opinions
-- **Non-promotional**: vendor mentions are comparative and contextual, not endorsements
-
-The test question: would this content belong in an engineering organization's internal knowledge base, cited in architecture review documents? If yes, it is a deep dive. If it reads like a personal take on a recent controversy, it is a blog post.
-
----
-
-## Criteria for Inclusion
-
-A proposed deep dive must satisfy all of the following:
-
-**1. Genuine and consequential tension**: the topic must involve a real architectural decision with non-obvious trade-offs and meaningful consequences for system longevity, team capability, or operational cost. "Which library should I use for X?" is not a deep dive topic. "What does the choice of orchestration model imply for organizational ownership of data pipelines?" is.
-
-**2. Analytical completeness**: the entry must include at minimum: a historical or contextual arc, a comparative analysis of two or more approaches, at least one structured comparison (table, matrix, or diagram), and a decision framework that provides actionable guidance.
-
-**3. Appropriate scope**: a deep dive should address a single coherent question or tension. If the scope requires more than 3,000 words to be analytically complete, it is likely too broad and should be split. If it can be addressed in fewer than 800 words, it is likely a best practice annotation, not a deep dive.
-
-**4. Durability**: the analysis should remain relevant for at least two to three years without significant revision. Technology-specific entries (e.g., "Airflow 2.6 vs 2.7") fail this criterion. Platform-philosophy entries (e.g., "DAG-first vs code-first orchestration models") satisfy it.
-
-**5. Non-overlap with existing content**: a proposed deep dive must not substantially duplicate an existing deep dive. It may reference, contrast, or extend existing deep dives, but the primary analytical contribution must be distinct.
-
----
-
-## Guardrails Against Dilution
-
-The Deep Dives section degrades in analytical value if it accumulates:
-
-- Entries that are tutorials in analytical clothing (detailed implementation instructions with light framing)
-- Entries that are opinion pieces without analytical structure
-- Entries on topics so narrow that they would better serve as best practice annotations
-- Entries on topics so broad that they cannot be addressed rigorously in a single document
-- Entries duplicating content from existing deep dives without a meaningfully distinct analytical contribution
-
-The ADR process serves as the formal check on section-level structural changes (adding new sub-sections, reorganizing the corpus). Individual deep dive entries do not require an ADR if they meet the criteria above. New sub-sections or reorganizations require an ADR update.
-
----
-
-## Consequences
-
-The Deep Dives section becomes a twelve-entry analytical corpus with five sub-sections. Cross-links between related entries are maintained as part of authoring discipline. The "Operations & Reliability" sub-section is established as a home for future entries addressing failure analysis, incident management, and reliability engineering at the conceptual level.
-
----
-
-## Related Documents
-
-- [ADR 0008: Expand Deep Dives as a First-Class Analytical Section](0008-expand-deep-dives-section.md) — previous expansion and criteria formalization
-- [ADR 0007: Deep Dives Section](0007-deep-dives-section.md) — original section creation
diff --git a/docs/adr/0010-deep-dives-governance.md b/docs/adr/0010-deep-dives-governance.md
deleted file mode 100644
index f9ea2e0..0000000
--- a/docs/adr/0010-deep-dives-governance.md
+++ /dev/null
@@ -1,126 +0,0 @@
----
-tags:
- - adr
- - documentation
----
-
-# ADR 0010: Governance Model for Deep Dives
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Author**: Joshua N. Grant
-
----
-
-## Context
-
-The Deep Dives section has grown to seventeen entries across six sub-sections through ADRs 0007–0009. Three successive expansion ADRs have each restated criteria for inclusion with minor variations. This ADR consolidates the governance model into a single authoritative reference, superseding the inclusion criteria statements in ADRs 0008 and 0009.
-
-The growth of the section raises a specific risk: as the entry count increases, the average quality and analytical rigor of entries tends to decline. The first entries in any analytical corpus establish a standard that becomes harder to maintain as the pressure to fill in topics increases. Explicit governance makes the standard visible and enforceable.
-
----
-
-## Why Long-Form Analytical Content Requires Curation
-
-Long-form analytical content is not self-governing. The absence of curation produces a set of tendencies that degrade corpus quality over time:
-
-**Scope inflation**: without defined scope, entries expand to cover everything tangentially related to the topic. An entry that begins as an analysis of distributed tracing trade-offs expands to include step-by-step setup instructions, vendor comparison tables, and configuration recommendations. The analytical character is diluted by operational detail that belongs in best practices or tutorials.
-
-**Tone drift**: without a defined tone standard, analytical entries drift toward tutorial voice (instructional), blog voice (conversational and opinionated without systematic justification), or vendor-advocacy voice (promotional). Each is appropriate in a different content type; none is appropriate for this section.
-
-**Redundancy**: without cross-reference awareness, entries duplicate analysis that already exists in the corpus. The economics of observability overlap with the observability vs monitoring entry; the appropriate use of microservices overlaps with the monolith entry. Redundancy is acceptable if the analytical angle is genuinely distinct; it is wasteful if it is merely repetition.
-
-**Obsolescence**: analytical content on rapidly changing technologies can become misleading without update. An entry on a specific tool's architecture that was accurate in 2024 may be incorrect by 2026. The governance model must include a policy for handling obsolescence.
-
----
-
-## Criteria for Inclusion
-
-A proposed deep dive must satisfy all criteria in each category.
-
-### Topic Criteria
-
-1. **Genuine architectural tension**: the topic involves a decision with consequential trade-offs that cannot be resolved by a simple rule. Decisions where one option is obviously correct for all contexts do not merit deep dive treatment.
-
-2. **Consequential scope**: the decision affects system longevity, team capability, operational cost, or organizational structure across a time horizon of at least one to two years. Short-term implementation decisions belong in tutorials or best practices.
-
-3. **Non-duplication**: the proposed entry must not substantially duplicate an existing entry. A different analytical angle on the same topic is acceptable if the angle is explicitly identified and the new entry cross-references the existing one. Mere restatement is not acceptable.
-
-4. **Durability**: the core analysis must remain relevant for two to three years without requiring substantive revision. Tool-specific version comparisons and feature checklists fail this criterion. Philosophy-level comparisons of execution models, governance approaches, or architectural trade-offs satisfy it.
-
-### Structural Criteria
-
-Every deep dive must include, in order:
-
-1. **Opening Thesis** — a clear statement of the analytical claim the document supports. Not "this document discusses X"; a substantive claim about X that the document then argues for.
-
-2. **Historical Context** — the conditions that produced the current situation, the predecessor approaches, and why the current alternatives exist.
-
-3. **Analytical Core** — the comparison, trade-off analysis, or architectural examination. This section should include at least one structured comparison (table, matrix, or ASCII diagram) and at least two to three paragraphs of dense analytical prose per major claim.
-
-4. **Decision Framework** — actionable guidance organized by stakeholder context (team size, regulatory environment, operational maturity, workload characteristics). Not "it depends" without criteria; specific criteria for each branch of the decision.
-
-5. **Cross-links** — at minimum two links to related content in the site (deep dives, best practices, tutorials, ADRs). Cross-links are not decorative; they should reflect genuine analytical relationships.
-
-### Tone Criteria
-
-All deep dives must:
-
-- Avoid tutorial voice: no step-by-step instructions, no "to do X, first do Y" sequences
-- Avoid marketing voice: no vendor endorsements, no superlative claims without evidence, no promotional language for any tool or platform
-- Avoid blog voice: positions must be argued with evidence and logical structure, not asserted as personal opinions
-- Avoid hedging to the point of uselessness: the Decision Framework must make actual recommendations, not merely list considerations
-
----
-
-## Sub-Section Governance
-
-New sub-sections may be created when at least two deep dives share a coherent analytical domain that is not adequately captured by existing sub-sections. A single entry does not justify a new sub-section. Sub-section creation requires an ADR update.
-
-Current sub-sections and their intended scope:
-
-| Sub-section | Scope |
-|---|---|
-| Data Formats & Storage | Physical storage formats, encoding trade-offs, cloud-native implications |
-| Data Systems & Architecture | Databases, data platforms, pipeline architecture, orchestration |
-| Systems Design & Architecture | Service decomposition, distributed systems, scalability trade-offs |
-| Operations & Reliability | Observability, monitoring, incident management, operational economics |
-| Infrastructure & Automation | Container philosophy, IaC, GitOps, deployment models |
-| Embedded & Radio | Microcontroller platforms, IoT protocols, radio systems |
-
----
-
-## Handling Obsolescence
-
-Entries that contain analysis that has become materially incorrect due to technology changes should be updated rather than deleted. The update should be noted at the top of the document with the date and nature of the update. If the core analytical claim has been invalidated (not merely that specific tools have changed), the entry should be marked as superseded and a new entry created.
-
-Entries that are superseded or significantly updated trigger a review of cross-links from other entries to ensure referenced conclusions remain accurate.
-
----
-
-## Cross-Link Expectations
-
-Cross-links serve two functions: they help readers navigate the analytical corpus, and they expose the analytical relationships between topics. Entries that are analytically related but not cross-linked create isolated analysis that readers cannot discover through navigation.
-
-Minimum cross-link expectations:
-- Every deep dive links to at least two other entries in the corpus (deep dives, best practices, or tutorials)
-- Every new entry is linked from at least one existing entry or from the index
-- The deep-dives/index.md is updated for every new entry
-
-Cross-links are maintained as part of the authoring process, not as a retrospective cleanup task.
-
----
-
-## Consequences
-
-This ADR supersedes the inclusion criteria statements in ADRs 0008 and 0009. Those ADRs remain as historical records of the expansion decisions they document; their criteria sections are no longer authoritative.
-
-Future deep dive entries do not require an ADR. They require adherence to the criteria above. ADR updates are required only for structural changes to the section: new sub-sections, reorganization, or material changes to the governance model itself.
-
----
-
-## Related Documents
-
-- [ADR 0009: Deep Dives Expansion](0009-deep-dives-expansion.md) — previous expansion (criteria superseded by this ADR)
-- [ADR 0008: Expand Deep Dives Section](0008-expand-deep-dives-section.md) — previous expansion (criteria superseded by this ADR)
-- [ADR 0007: Deep Dives Section](0007-deep-dives-section.md) — original section creation decision
diff --git a/docs/adr/0011-deep-dives-expansion-phase-2.md b/docs/adr/0011-deep-dives-expansion-phase-2.md
deleted file mode 100644
index 421e578..0000000
--- a/docs/adr/0011-deep-dives-expansion-phase-2.md
+++ /dev/null
@@ -1,90 +0,0 @@
----
-tags:
- - adr
- - documentation
----
-
-# ADR 0011: Deep Dives Expansion Phase II
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Author**: Joshua N. Grant
-
----
-
-## Context
-
-ADR 0010 established a consolidated governance model for the Deep Dives section and noted that future deep dive entries do not require an ADR — they require adherence to the criteria established in ADR 0010. This ADR documents Phase II of the corpus expansion as a record of scope decisions rather than as a governance update.
-
-Phase II adds five entries across three new topic clusters: real-time systems economics, data lake governance failure, and distributed systems theory. A fourth entry enhances the existing metadata-as-infrastructure document with a data plane/control plane framework and a metadata debt section. The Human Cost of Automation addresses the organizational side of the infrastructure automation deep dives already present in the corpus.
-
-The corpus now contains twenty-two entries across six sub-sections.
-
----
-
-## Entries Added in Phase II
-
-**Data Systems & Architecture**
-
-- **The Hidden Cost of Real-Time Systems** — latency taxonomy (hard/soft/streaming/batch), infrastructure costs of stateful stream processing (checkpointing, backpressure, deduplication), observability explosion in streaming, cost comparison table, and a decision framework from batch through micro-batch to streaming
-- **Why Most Data Lakes Become Data Swamps** — structural decay mechanisms (missing metadata, schema drift, orphaned datasets, version confusion), technical anti-patterns, cultural failures, lakehouse as partial solution, and zone-based architecture for prevention
-- **Metadata as Infrastructure (enhanced)** — addition of data plane vs metadata plane ASCII diagram, metadata debt taxonomy (stale descriptions, broken lineage, inconsistent tags), and updated title framing metadata as a hidden control plane rather than mere documentation
-
-**Systems Design & Architecture**
-
-- **Distributed Systems and the Myth of Infinite Scale** — CAP theorem and PACELC analysis, N² coordination growth diagram, data gravity and cloud egress cost, partial failure taxonomy, rolling upgrade complexity, and a decision framework distinguishing genuine distribution requirements from distribution-as-prestige
-
-**Infrastructure & Automation**
-
-- **The Human Cost of Automation** — skill atrophy mechanisms, complexity transfer from visible toil to invisible fragility, psychological impact of reduced operator agency, automation maturity model (manual through self-healing), and a decision framework with explicit criteria for when automation creates more organizational risk than it removes
-
----
-
-## Criteria for Future Inclusion
-
-ADR 0010 remains the authoritative governance document. For reference, the criteria are:
-
-**Topic**: genuine architectural tension, consequential scope (1–3 year horizon), non-duplication of existing entries, and analytical durability (not version-specific).
-
-**Structure**: Opening Thesis → Historical Context → Analytical Core (with comparison and diagram) → Decision Framework → Cross-links.
-
-**Tone**: analytical (not tutorial, not marketing, not blog voice). Positions must be argued with evidence and logical structure.
-
----
-
-## Avoiding Intellectual Dilution
-
-The risk at twenty-two entries is that the corpus begins to include topics that are adjacent to but not clearly within the analytical scope of the section. The test question from ADR 0010 remains the primary filter: would this content belong in an engineering organization's internal knowledge base, cited in architecture review documents?
-
-Two additional tests apply at scale:
-
-**The novelty test**: does the proposed entry add an analytical perspective not already present in the corpus? An entry on "Why Most Kubernetes Deployments Are Overcomplicated" that makes the same structural argument as "Why Most Microservices Should Be Monoliths" applied to a different technology does not pass the novelty test. The corpus should provide a range of analytical perspectives, not the same perspective applied to many technology choices.
-
-**The cross-link test**: a well-positioned deep dive should naturally link to and from at least two existing entries. A proposed entry with no natural cross-links to existing corpus entries is likely analytically isolated from the corpus's themes — a signal that it is not well-suited to the section.
-
----
-
-## Cross-Linking Discipline
-
-Cross-links in the Deep Dives corpus serve a function beyond navigation: they reveal the analytical graph structure of the corpus. Entries that form a cluster of mutual cross-links address related analytical questions. The current corpus has two main analytical clusters:
-
-**Data systems cluster**: Parquet, Lakehouse, DuckDB, Pipelines, Data Swamps, Metadata, Real-Time, Prefect/Airflow — all address different dimensions of the same set of trade-offs in analytical data architecture.
-
-**Systems and operations cluster**: Microservices, Distributed Systems, Observability, Economics of Observability, Automation, IaC/GitOps, Containers — all address different dimensions of how software systems are built, operated, and governed.
-
-**Embedded cluster**: LoRaWAN, MQTT, ESP32 — domain-specific analytical content for the IoT/embedded systems audience.
-
-Cross-links should predominantly connect entries within clusters. Cross-cluster links are appropriate when the analytical connection is genuine (e.g., the Economics of Observability links to Data Pipeline observability) rather than merely topical.
-
----
-
-## Consequences
-
-The corpus is mature enough to serve as a coherent analytical reference rather than a collection of independent essays. Future entries should be evaluated for their contribution to the corpus's analytical structure — do they deepen an existing cluster, or do they begin a new cluster with at least two related entries? Isolated single entries on topics without natural corpus neighbors should be evaluated skeptically.
-
----
-
-## Related Documents
-
-- [ADR 0010: Governance Model for Deep Dives](0010-deep-dives-governance.md) — authoritative governance (criteria, tone, structure, cross-linking)
-- [ADR 0009: Deep Dives Expansion](0009-deep-dives-expansion.md) — Phase I expansion record
diff --git a/docs/adr/0012-deep-dives-scale-governance.md b/docs/adr/0012-deep-dives-scale-governance.md
deleted file mode 100644
index 47e2ffc..0000000
--- a/docs/adr/0012-deep-dives-scale-governance.md
+++ /dev/null
@@ -1,138 +0,0 @@
----
-tags:
- - adr
- - governance
- - deep-dives
----
-
-# ADR 0012: Governance Model for Large-Scale Deep Dive Section
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Supersedes**: Curation criteria in ADR 0010 and ADR 0011
-**Context**: Deep Dives section has grown to 24 essays across 6 thematic clusters.
-
----
-
-## Context
-
-The Deep Dives section was established in ADR 0007 to provide a space for long-form analytical essays distinct from tutorials and best practices. It has since expanded through four phases of addition (ADR 0008, 0009, 0010, 0011) and now contains 24 essays spanning data formats, data systems, systems architecture, observability, infrastructure automation, embedded systems, and GPU economics.
-
-At this scale, the section faces three governance challenges that require explicit policy:
-
-1. **Curation discipline**: not every analytical question warrants a new deep dive. Without a clear novelty test, the section will accumulate essays that repeat one another, dilute intellectual signal, and create maintenance burden without proportional reader value.
-
-2. **Cross-linking requirements**: at 24 essays, the potential cross-link surface is large. Without policy, cross-linking becomes arbitrary (linking everything to everything) or stale (links added at creation time but not maintained as the corpus grows).
-
-3. **Thematic clustering**: the navigation must remain legible as the section grows. Sub-sections that contain too few or too many essays degrade navigation usability. Cluster boundaries must be governed explicitly.
-
----
-
-## Decision
-
-### Curation Discipline
-
-A new deep dive is warranted only when it satisfies all of the following:
-
-**Novelty test**: the essay addresses a question not adequately covered by any existing deep dive, even by implication. If the answer to "does this already exist?" is "mostly, in another essay," the appropriate action is to add a section to the existing essay, not create a new file.
-
-**Analytical substance test**: the essay must contain genuine comparative analysis — trade-off matrices, historical context, ecosystem maturity commentary, economic implications — not just an informed opinion or a list of considerations. If it cannot sustain 800+ words of structured argument, it is not a deep dive.
-
-**Durability test**: the essay must address a structural question with multi-year relevance. Technology announcements, version-specific comparisons, and tactical "how to choose" guides are not deep dives. A deep dive on Kubernetes overhead remains relevant as Kubernetes itself evolves; an essay on "Kubernetes 1.29 upgrade tips" does not.
-
-**Non-tutorial test**: the essay must not contain step-by-step implementation instructions. If it does, the implementation content belongs in the Tutorials section and the analytical content belongs in Deep Dives or Best Practices.
-
-### Cross-Linking Requirements
-
-Every new deep dive must include at minimum 2 and at maximum 5 internal cross-links to other deep dives, placed at the opening (header note) and/or closing (See also admonition). Cross-links must be directionally meaningful — they should explain *why* the linked essay is relevant, not merely that it exists.
-
-When a new deep dive is added, at minimum 2 existing essays that are conceptually adjacent must be updated to add a reciprocal cross-link. This maintains the web of connections as the corpus grows and prevents new essays from being isolated.
-
-Cross-links to Best Practices and Tutorial pages are appropriate when the deep dive directly extends or contextualizes a practice or tutorial. These links should appear in the See also admonition only.
-
-### Thematic Clustering
-
-The six current thematic clusters are:
-
-| Cluster | Current count | Soft maximum |
-|---|---|---|
-| Data Formats & Storage | 3 | 6 |
-| Data Systems & Architecture | 8 | 12 |
-| Systems Design & Architecture | 4 | 7 |
-| Operations & Reliability | 2 | 5 |
-| Infrastructure & Automation | 4 | 7 |
-| Embedded & Radio | 3 | 6 |
-
-When a cluster approaches its soft maximum, the appropriate response is one of:
-- Split the cluster into two sub-clusters with distinct thematic identity
-- Merge essays that substantially overlap into a revised composite essay
-- Enforce the novelty test more strictly for proposed additions to that cluster
-
-Sub-cluster maximum: 12 essays per thematic cluster in the navigation. Beyond 12, navigation usability degrades to the point where sub-cluster splitting is mandatory.
-
-### Obsolescence Handling
-
-Deep dives become partially obsolete when the ecosystem evolves significantly: a technology ceases to be relevant, a comparison changes due to product acquisition or open-source abandonment, or the fundamental trade-off the essay describes resolves in one direction definitively. The response is:
-
-- **Minor obsolescence** (one section is stale): add a dated note at the beginning of the affected section. Do not rewrite the essay.
-- **Major obsolescence** (the central thesis is invalidated): mark the essay with a deprecation notice at the top, link to any successor essay, and leave the original content in place for historical context.
-- **Supersession** (a new essay covers the same territory with superior analysis): the new essay should explicitly state what it supersedes; the old essay should link to the new one.
-
----
-
-## Thematic Cluster Map (Current State)
-
-At the time of this ADR:
-
-```
-Deep Dives (24 essays)
-│
-├── Data Formats & Storage (3)
-│ Parquet · Geospatial Formats · Polars vs Pandas Geospatial
-│
-├── Data Systems & Architecture (8)
-│ Lakehouse · DuckDB vs PG vs Spark · Data Pipeline Failures ·
-│ Data Lakes → Swamps · Prefect vs Airflow · Metadata as Infra ·
-│ Hidden Cost of Real-Time · End of the Data Warehouse?
-│
-├── Systems Design & Architecture (4)
-│ Monolith vs Microservices · Appropriate Microservices ·
-│ Distributed Scale Myths · Kubernetes Overhead
-│
-├── Operations & Reliability (2)
-│ Observability vs Monitoring · Economics of Observability
-│
-├── Infrastructure & Automation (4)
-│ Container Base Images · IaC vs GitOps ·
-│ Human Cost of Automation · Economics of GPU Infrastructure
-│
-└── Embedded & Radio (3)
- LoRaWAN vs Raw LoRa · MQTT vs HTTP · ESP32 vs Raspberry Pi
-```
-
----
-
-## Rationale
-
-The governance model prioritizes intellectual coherence over volume. A corpus of 24 high-quality essays is more valuable than a corpus of 50 essays of inconsistent depth. The soft cluster maximums create a forcing function: adding a new essay to a near-full cluster requires either demonstrating the novelty test is met or splitting the cluster, both of which impose a productive discipline on scope.
-
-The cross-linking requirement prevents the section from fragmenting into isolated essays. The reciprocal link requirement is the mechanism that keeps the web of connections current as the corpus grows.
-
----
-
-## Alternatives Considered
-
-**No explicit governance**: allow the section to grow organically, with ADRs created only for major structural changes. Rejected because at 24 essays, without explicit policy, the section will accumulate redundant essays and navigation will degrade.
-
-**Hard essay limits**: set an absolute maximum (e.g., 30 essays). Rejected because the correct limit is topical coverage, not count. A section with 25 essays each addressing a distinct structural question is better than 20 essays with significant overlap.
-
-**Separate site section for each thematic cluster**: treat Data Systems deep dives as a separate section from Systems Design deep dives. Rejected because the cross-cluster linkage is a feature — data system decisions interact with systems design decisions — and the unified Deep Dives section preserves navigational context for that interaction.
-
----
-
-## Consequences
-
-- Future deep dives must pass the novelty, substance, durability, and non-tutorial tests before creation.
-- New essays must include 2–5 internal cross-links and trigger reciprocal cross-link updates in at least 2 existing essays.
-- Navigation sub-cluster sizes must be monitored; splitting is mandatory when a cluster exceeds 12 entries.
-- Obsolescence handling policy is now explicit and must be applied when significant ecosystem changes render existing essays partially stale.
diff --git a/docs/adr/0013-deep-dives-curation-model.md b/docs/adr/0013-deep-dives-curation-model.md
deleted file mode 100644
index 39a3751..0000000
--- a/docs/adr/0013-deep-dives-curation-model.md
+++ /dev/null
@@ -1,112 +0,0 @@
----
-tags:
- - adr
- - governance
- - deep-dives
----
-
-# ADR 0013: Deep Dive Curation and Thematic Clustering Model
-
-**Status**: Accepted
-**Date**: 2026-02-26
-**Supersedes**: Curation criteria in ADR 0012 (governance thresholds now updated)
-**Context**: Deep Dives section has grown to 29 essays across 6 thematic clusters.
-
----
-
-## Context
-
-ADR 0012 established a governance model for the Deep Dives section when it contained 24 essays. This ADR refines that model as the corpus has grown to 29 essays — crossing the ADR 0012 soft cluster maximum for two clusters (Data Formats & Storage: 5 entries; Data Systems & Architecture: 10 entries) — and as experience with the cross-linking and curation requirements has revealed patterns that require more explicit guidance.
-
-The primary governance challenge at this scale is **conceptual drift**: essays that individually satisfy the novelty and substance tests but that, in aggregate, produce a corpus whose thematic coherence is harder for readers to perceive. A corpus of 29 essays organized into 6 clusters is legible only if the cluster boundaries are sharp enough that readers can use them to navigate rather than reading the entire index.
-
-The secondary challenge is **cross-link maintenance**: at 29 essays, the potential cross-link surface is large enough that maintaining all cross-links manually as essays are added requires disciplined tracking. The current practice of adding cross-links only at creation time and updating 2–5 existing pages produces a cross-link graph that grows slower than the corpus.
-
----
-
-## Decision
-
-### Curation: Thematic Fit as Prerequisite
-
-Every proposed deep dive must be assignable to exactly one existing thematic cluster before creation. If an essay spans multiple clusters with equal weight, it either belongs in the cluster where its *primary argument* lives, or it reveals a missing cluster boundary. Essays that cannot be assigned to a cluster without forcing are evidence of scope problem rather than thematic novelty.
-
-The curation question is: **"what cluster would a reader look in to find this essay?"** If the answer requires describing a cluster that does not exist, the question becomes: "is there sufficient topical density to justify a new cluster, or is the essay an edge case?"
-
-### Cluster Definitions (Updated)
-
-The six clusters are now explicitly defined by their organizing analytical question:
-
-| Cluster | Organizing question | Current count |
-|---|---|---|
-| Data Formats & Storage | What physical and format-level constraints govern how data is stored and accessed? | 5 |
-| Data Systems & Architecture | What architectural and organizational forces shape analytical data systems? | 10 |
-| Systems Design & Architecture | What structural limits constrain the design of distributed and decomposed software systems? | 4 |
-| Operations & Reliability | What operational and economic forces govern system observability, monitoring, and reliability? | 2 |
-| Infrastructure & Automation | What organizational and technical trade-offs govern infrastructure acquisition, automation, and compute economics? | 5 |
-| Embedded & Radio | What physical and ecosystem constraints govern embedded system and radio protocol selection? | 3 |
-
-**Data Systems & Architecture** has reached the soft maximum of 10. The next addition to this cluster requires either demonstrating that the proposed essay has a sufficiently distinct primary argument from the 10 existing essays, or splitting the cluster into two sub-clusters. Candidate split: "Data Storage & Table Formats" (Lakehouse, Warehouse, Parquet, Metadata, Data Lakes, Data Pipelines) vs "Data Computation & Orchestration" (DuckDB/PG/Spark, Prefect/Airflow, Real-Time, ML Systems).
-
-### Cross-Linking: Minimum Requirements
-
-Every new essay must:
-
-1. Include a header cross-link (`*See also: ...*`) to 1–2 essays in *other* clusters, explaining why the reader should continue there.
-2. Include a closing `!!! tip "See also"` admonition with 3–5 cross-links with directionally meaningful descriptions.
-3. Trigger reciprocal updates in at least 2 existing essays that are the most thematically adjacent.
-
-The purpose of cross-cluster header links is navigational: a reader who arrived at a Data Systems essay from the index should have a visible path to the Infrastructure cluster if the essay assumes infrastructure concepts. The purpose of closing cross-links is analytical: they tell the reader what questions to pursue after this essay.
-
-### Avoiding Conceptual Drift
-
-The symptom of conceptual drift is when a reader looking at the index cannot determine, from the title and one-sentence description, what distinguishes one essay from three adjacent essays. The prevention mechanism is the **distinction test**: before creating a new essay, identify the 2–3 most topically similar existing essays and articulate, in one sentence, how the new essay's central argument differs from each. If the distinction cannot be articulated clearly, the essay is not ready for creation.
-
----
-
-## Current Corpus Map (29 essays)
-
-```
-Deep Dives (29 essays)
-│
-├── Data Formats & Storage (5)
-│ Parquet · Geospatial Formats · Polars vs Pandas Geospatial ·
-│ Physics of Storage Systems · Operational Geometry of Spatial Systems
-│
-├── Data Systems & Architecture (10) ← approaching split threshold
-│ Lakehouse · DuckDB vs PG vs Spark · Data Pipeline Failures ·
-│ Data Lakes → Swamps · Prefect vs Airflow · Metadata as Infra ·
-│ Hidden Cost of Real-Time · End of the Data Warehouse? ·
-│ Hidden Cost of Metadata Debt · Why ML Systems Fail in Production
-│
-├── Systems Design & Architecture (4)
-│ Monolith vs Microservices · Appropriate Microservices ·
-│ Distributed Scale Myths · Kubernetes Overhead
-│
-├── Operations & Reliability (2)
-│ Observability vs Monitoring · Economics of Observability
-│
-├── Infrastructure & Automation (5)
-│ Container Base Images · IaC vs GitOps ·
-│ Human Cost of Automation · Economics of GPU Infrastructure ·
-│ Myth of Serverless Simplicity
-│
-└── Embedded & Radio (3)
- LoRaWAN vs Raw LoRa · MQTT vs HTTP · ESP32 vs Raspberry Pi
-```
-
----
-
-## Rationale
-
-The governance model deliberately imposes friction on new additions. The cost of this friction is slower corpus growth. The benefit is a corpus whose individual essays are more distinct from one another, whose cluster boundaries are more navigable, and whose cross-link graph is maintained rather than stale. The section is more valuable at 30 high-distinction essays than at 50 essays with significant conceptual overlap.
-
-The cluster split decision for Data Systems & Architecture is deferred to the next addition; the split should be made based on where the 11th essay falls, not in anticipation of a hypothetical essay.
-
----
-
-## Consequences
-
-- Proposed essays must satisfy the distinction test (articulated difference from 2–3 most similar existing essays) before creation.
-- Data Systems & Architecture approaching 10 essays requires monitoring; the 11th addition triggers a cluster split decision.
-- Cross-linking minimums (header link + 3–5 closing links + 2 reciprocal updates) are formalized as creation requirements.
-- ADR 0012 curation criteria remain valid; this ADR refines them for a larger corpus and adds the distinction test and cluster definition model.
diff --git a/docs/adr/0014-elevate-site-to-systems-doctrine.md b/docs/adr/0014-elevate-site-to-systems-doctrine.md
deleted file mode 100644
index a63c966..0000000
--- a/docs/adr/0014-elevate-site-to-systems-doctrine.md
+++ /dev/null
@@ -1,119 +0,0 @@
-# ADR 0014: Elevate Site from Documentation Collection to Cohesive Systems Doctrine
-
-- Status: Accepted
-- Date: 2026-02-26
-- Deciders: Joshua Grant
-- Technical Story: Transition the site from a growing collection of tutorials and analytical essays into a coherent, navigable systems-level doctrine.
-
----
-
-## Context
-
-The site has evolved beyond tutorials and best practices into a substantial body of analytical deep dives covering:
-
-- Distributed systems
-- Storage theory
-- Observability
-- Data architecture
-- Embedded systems
-- Infrastructure economics
-- Geospatial systems
-
-While the intellectual depth has increased, structural cohesion has not kept pace.
-
-The risks are:
-
-- Content silos
-- Cognitive overload for new readers
-- Redundant arguments across deep dives
-- Lack of narrative entry points
-- Weak conceptual mapping between sections
-
-The site is at an inflection point:
-It can either become a loosely connected archive, or a structured intellectual framework.
-
----
-
-## Decision
-
-We will introduce structural layers that transform the site into a cohesive systems doctrine:
-
-1. **Architectural Entry Points**
- - A “Start Here” decision tree page
- - Curated Reading Tracks
-
-2. **Thematic Cohesion**
- - Cross-cutting theme declarations
- - A Decision Framework Index
- - An Anti-Patterns Index
-
-3. **Intellectual Framing**
- - A “Philosophy of the Site” page
- - A Systems Thinking Glossary
- - A Canon / Recommended Reading order
-
-4. **Cross-Section Synthesis**
- - Meta-essays connecting deep dives
- - Explicit cross-links between related doctrines
-
-5. **Intentional Structure**
- - Minimal redundancy
- - Clear semantic relationships
- - Designed, not accumulated navigation
-
----
-
-## Alternatives Considered
-
-### 1. Continue Adding Content Without Structural Reform
-Rejected.
-This risks fragmentation and reader fatigue.
-
-### 2. Create an External Blog for Philosophy
-Rejected.
-The philosophy is intrinsic to the technical content and should remain integrated.
-
-### 3. Major Structural Rewrite
-Rejected.
-We prefer incremental, compositional refinement.
-
----
-
-## Consequences
-
-### Positive
-
-- Increased intellectual cohesion
-- Clear onboarding path for new readers
-- Stronger narrative continuity
-- Reduced conceptual redundancy
-- Elevated authority and intentionality
-
-### Negative
-
-- Additional structural overhead
-- Increased editorial discipline required
-- Slightly more complex navigation
-
----
-
-## Follow-ups
-
-- Implement "Start Here" decision tree
-- Add Reading Tracks page
-- Add Decision Framework Index
-- Add Anti-Patterns index
-- Add Philosophy page
-- Add Systems Glossary
-- Introduce light thematic labeling
-
----
-
-## Governance
-
-Future deep dives must:
-
-- Declare themes
-- Cross-link to at least two existing essays
-- Avoid duplication of core arguments
-- End with a structured decision framework
\ No newline at end of file
diff --git a/docs/adr/0015-diagrams-mermaid.md b/docs/adr/0015-diagrams-mermaid.md
deleted file mode 100644
index 1ec0ef1..0000000
--- a/docs/adr/0015-diagrams-mermaid.md
+++ /dev/null
@@ -1,93 +0,0 @@
----
-tags:
- - adr
- - diagrams
- - documentation
----
-
-# ADR 0015: Standardize Site Diagrams on Mermaid (with SVG as Optional Exception)
-
-**Status**: Accepted
-**Date**: 2026-02-27
-**Deciders**: Joshua N. Grant
-**Technical Story**: Replace ad-hoc ASCII diagrams with a consistent, maintainable diagram standard that is version-friendly and renders cleanly across the site.
-
----
-
-## Context
-
-The site increasingly relies on conceptual diagrams to explain architecture, trade-offs, and system boundaries. Many diagrams are currently written as ASCII. While ASCII is portable and fast to write, it has several drawbacks:
-
-- Limited expressive power for complex relationships
-- Inconsistent visual vocabulary across pages
-- Reduced readability on mobile and narrow screens
-- Difficult to reuse or standardize diagram patterns
-- Perceived "sketchy" quality even when content is rigorous
-
-SVG diagrams can look excellent but tend to introduce tooling overhead:
-
-- Requires authoring tools or generation pipelines
-- Harder to edit and review as text diffs
-- Risk of inconsistent styling unless a strict convention exists
-
-Mermaid offers a middle ground:
-
-- Text-based diagrams (diff-friendly, reviewable)
-- Easy to embed in Markdown
-- Supports common diagram types (flowcharts, sequence diagrams, state diagrams)
-- Low friction for authors
-- Rendered by MkDocs Material through `pymdownx.superfences` — already configured
-
----
-
-## Decision
-
-We will standardize on **Mermaid** for the majority of diagrams across the site.
-
-- Mermaid becomes the default for architecture, flow, state, and sequence diagrams.
-- SVG remains an optional exception for:
- - High-resolution posters and figures
- - Geospatial illustrations and maps
- - Cases where Mermaid is insufficient or visually unclear
-
-We will:
-
-1. Confirm Mermaid rendering is enabled in `mkdocs.yml` (it is — via `pymdownx.superfences` custom fence).
-2. Add a diagram style guide at `docs/diagrams/style-guide.md`.
-3. Convert a starter set of high-value ASCII diagrams to Mermaid (see Follow-ups).
-4. Gradually convert remaining ASCII diagrams during normal content evolution.
-
----
-
-## Alternatives Considered
-
-**Continue using ASCII diagrams**: Rejected. Lacks consistency and visual clarity at scale.
-
-**Adopt SVG as the default**: Rejected. Higher authoring friction and tooling requirements.
-
-**Use both equally without a default**: Rejected. Without a default standard, style fragments and diagrams drift.
-
----
-
-## Consequences
-
-**Positive**:
-- Higher clarity and visual consistency across the site
-- Better mobile readability
-- Text-based diffs for diagrams in pull requests
-- Reduced diagram authoring friction compared to SVG
-- Stronger perceived polish without sacrificing maintainability
-
-**Negative / Trade-offs**:
-- Mermaid features vary by diagram type and renderer version
-- Some complex diagrams may require careful layout constraints
-- Small learning curve for contributors
-
----
-
-## Follow-ups
-
-- [x] Confirm `pymdownx.superfences` mermaid custom fence is in `mkdocs.yml`
-- [x] Add `docs/diagrams/style-guide.md` with snippet library
-- [x] Convert core doctrine diagrams: Prefect vs Airflow, ESP32 vs Pi, Data Pipelines, Observability vs Monitoring
-- [ ] Convert remaining ASCII diagrams gradually during content evolution
diff --git a/docs/adr/_templates/0001-template-standard.md b/docs/adr/_templates/0001-template-standard.md
deleted file mode 100644
index ab1e9df..0000000
--- a/docs/adr/_templates/0001-template-standard.md
+++ /dev/null
@@ -1,81 +0,0 @@
-# ADR-XXXX: [Short Descriptive Title]
-
-**Status**: Proposed | Accepted | Superseded | Rejected | Deprecated
-
-**Date**: YYYY-MM-DD
-
-**Deciders**: [Names or roles, e.g., "Architecture Working Group", "Tech Leads"]
-
-**Tags**: [Optional: component, domain, e.g., "observability", "security", "data"]
-
----
-
-## Context
-
-[2-4 paragraphs describing the situation, problem, or requirement that led to this decision. Include:
-- What problem are we solving?
-- What constraints exist?
-- What are the business/technical drivers?]
-
-## Decision
-
-[1-2 paragraphs stating the decision clearly and concisely. Be specific about what is being decided.]
-
-We will [specific action/choice].
-
-## Options Considered
-
-### Option 1: [Name of Option]
-
-**Description**: [What this option entails]
-
-**Pros**:
-- [Benefit 1]
-- [Benefit 2]
-
-**Cons**:
-- [Drawback 1]
-- [Drawback 2]
-
-### Option 2: [Name of Option]
-
-[Same structure as Option 1]
-
-### Option 3: [Name of Option]
-
-[Same structure as Option 1]
-
-## Rationale
-
-[2-3 paragraphs explaining why the chosen option was selected. Reference specific pros/cons, constraints, or requirements that drove the decision.]
-
-## Consequences
-
-### Positive
-
-- [Expected positive outcome 1]
-- [Expected positive outcome 2]
-
-### Negative
-
-- [Expected negative outcome 1]
-- [Expected negative outcome 2]
-
-### Neutral / Trade-offs
-
-- [Trade-off or neutral consequence]
-
-## Related Documents
-
-- [Link to design doc, RFC, or implementation guide]
-- [Link to best-practices document]
-- [Link to runbook or operational guide]
-
-## Implementation Notes
-
-[Optional: Brief notes on how this decision will be implemented, key milestones, or dependencies]
-
-## References
-
-[Optional: Links to external resources, research, or benchmarks that informed the decision]
-
diff --git a/docs/adr/_templates/0002-template-lightweight.md b/docs/adr/_templates/0002-template-lightweight.md
deleted file mode 100644
index 96f58d5..0000000
--- a/docs/adr/_templates/0002-template-lightweight.md
+++ /dev/null
@@ -1,28 +0,0 @@
-# ADR-XXXX: [Short Descriptive Title]
-
-**Status**: Proposed | Accepted | Superseded | Rejected | Deprecated
-
-**Date**: YYYY-MM-DD
-
-**Deciders**: [Names or roles]
-
----
-
-## Context
-
-[1-2 paragraphs: What problem or situation led to this decision?]
-
-## Decision
-
-[1 paragraph: What is being decided?]
-
-We will [specific action/choice].
-
-## Impact / Notes
-
-[1-2 paragraphs: What are the key consequences, trade-offs, or implementation considerations?]
-
-## Related Documents
-
-- [Link to relevant docs]
-
diff --git a/docs/adr/index.md b/docs/adr/index.md
deleted file mode 100644
index 4c6cbc0..0000000
--- a/docs/adr/index.md
+++ /dev/null
@@ -1,30 +0,0 @@
-# Architecture Decision Records
-
-Architecture Decision Records (ADRs) document significant decisions made while designing and building this site — what was chosen, what was rejected, and why.
-
-See [ADR and Technical Decision Governance](../best-practices/architecture-design/adr-decision-governance.md) for the methodology behind these records.
-
-## ADR Index
-
-| Number | Title | Status | Date |
-|--------|-------|--------|------|
-| [ADR-0001](0001-site-architecture.md) | Site Architecture — MkDocs + Material + Repo-as-Docs | Accepted | 2026-02-28 |
-| [ADR-0002](0002-why-mkdocs-material.md) | Use MkDocs + Material Theme | Accepted | 2024-01-15 |
-| [ADR-0003](0003-why-best-practices-vs-tutorials.md) | Separate Best Practices from Tutorials | Accepted | 2024-02-01 |
-| [ADR-0004](0004-why-just-for-fun-exists.md) | Include a Just for Fun Section | Accepted | 2024-03-10 |
-| [ADR-0005](0005-esp32-section-architecture.md) | Establish ESP32 & Embedded Systems as a First-Class Section | Accepted | 2026-02-26 |
-| [ADR-0006](0006-embedded-expansion-lora-power.md) | Expand Embedded Section — LoRa, Power Electronics, Variant Notes | Accepted | 2026-02-26 |
-| [ADR-0007](0007-deep-dives-section.md) | Introduce "Deep Dives" Section for Analytical Long-Form Content | Accepted | 2026-02-26 |
-| [ADR-0008](0008-expand-deep-dives-section.md) | Expand Deep Dives as a First-Class Analytical Section | Accepted | 2026-02-26 |
-| [ADR-0009](0009-deep-dives-expansion.md) | Expand Deep Dives into Core Analytical Layer | Accepted | 2026-02-26 |
-| [ADR-0010](0010-deep-dives-governance.md) | Governance Model for Deep Dives | Accepted | 2026-02-26 |
-| [ADR-0011](0011-deep-dives-expansion-phase-2.md) | Deep Dives Expansion Phase II | Accepted | 2026-02-26 |
-| [ADR-0012](0012-deep-dives-scale-governance.md) | Governance Model for Large-Scale Deep Dive Section | Accepted | 2026-02-26 |
-| [ADR-0013](0013-deep-dives-curation-model.md) | Deep Dive Curation and Thematic Clustering Model | Accepted | 2026-02-26 |
-| [ADR-0014](0014-elevate-site-to-systems-doctrine.md) | Elevate Site from Documentation Collection to Cohesive Systems Doctrine | Accepted | 2026-02-26 |
-| [ADR-0015](0015-diagrams-mermaid.md) | Standardize Site Diagrams on Mermaid | Accepted | 2026-02-27 |
-
-## Templates
-
-- **[Standard ADR Template](_templates/0001-template-standard.md)** — For most architectural decisions
-- **[Lightweight ADR Template](_templates/0002-template-lightweight.md)** — For smaller, focused decisions
diff --git a/docs/tutorials/best-practices-integration/api-first-polyglot-protobuf.md b/docs/archive/tutorials/best-practices-integration/api-first-polyglot-protobuf.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/api-first-polyglot-protobuf.md
rename to docs/archive/tutorials/best-practices-integration/api-first-polyglot-protobuf.md
diff --git a/docs/tutorials/best-practices-integration/cost-optimized-geokg-capacity-planning.md b/docs/archive/tutorials/best-practices-integration/cost-optimized-geokg-capacity-planning.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/cost-optimized-geokg-capacity-planning.md
rename to docs/archive/tutorials/best-practices-integration/cost-optimized-geokg-capacity-planning.md
diff --git a/docs/tutorials/best-practices-integration/cost-optimized-ml-pipeline-capacity-retention.md b/docs/archive/tutorials/best-practices-integration/cost-optimized-ml-pipeline-capacity-retention.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/cost-optimized-ml-pipeline-capacity-retention.md
rename to docs/archive/tutorials/best-practices-integration/cost-optimized-ml-pipeline-capacity-retention.md
diff --git a/docs/tutorials/best-practices-integration/data-pipeline-quality-governance-observability.md b/docs/archive/tutorials/best-practices-integration/data-pipeline-quality-governance-observability.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/data-pipeline-quality-governance-observability.md
rename to docs/archive/tutorials/best-practices-integration/data-pipeline-quality-governance-observability.md
diff --git a/docs/tutorials/best-practices-integration/developer-experience-taxonomy-repository.md b/docs/archive/tutorials/best-practices-integration/developer-experience-taxonomy-repository.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/developer-experience-taxonomy-repository.md
rename to docs/archive/tutorials/best-practices-integration/developer-experience-taxonomy-repository.md
diff --git a/docs/tutorials/best-practices-integration/documentation-driven-adr-governance.md b/docs/archive/tutorials/best-practices-integration/documentation-driven-adr-governance.md
similarity index 99%
rename from docs/tutorials/best-practices-integration/documentation-driven-adr-governance.md
rename to docs/archive/tutorials/best-practices-integration/documentation-driven-adr-governance.md
index 378b88d..99897b4 100644
--- a/docs/tutorials/best-practices-integration/documentation-driven-adr-governance.md
+++ b/docs/archive/tutorials/best-practices-integration/documentation-driven-adr-governance.md
@@ -4,7 +4,7 @@
This tutorial combines:
- **[Documentation Best Practices](../../best-practices/architecture-design/documentation.md)** - Writing and organizing useful documentation
-- **[ADR and Technical Decision Governance](../../best-practices/architecture-design/adr-decision-governance.md)** - Architecture Decision Records
+- **[Architecture & Design documentation](../../best-practices/architecture-design/documentation.md)** - includes ADR patterns as a general practice
- **[Reference Architecture Diagrams](../../best-practices/architecture-design/reference-architecture-diagrams.md)** - System documentation
- **[Cognitive Load Management and Developer Experience](../../best-practices/architecture-design/cognitive-load-developer-experience.md)** - Developer experience optimization
diff --git a/docs/tutorials/best-practices-integration/event-driven-geokg-observability.md b/docs/archive/tutorials/best-practices-integration/event-driven-geokg-observability.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/event-driven-geokg-observability.md
rename to docs/archive/tutorials/best-practices-integration/event-driven-geokg-observability.md
diff --git a/docs/tutorials/best-practices-integration/event-driven-microservices-observability-data-governance.md b/docs/archive/tutorials/best-practices-integration/event-driven-microservices-observability-data-governance.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/event-driven-microservices-observability-data-governance.md
rename to docs/archive/tutorials/best-practices-integration/event-driven-microservices-observability-data-governance.md
diff --git a/docs/tutorials/best-practices-integration/geospatial-data-mesh-cost-capacity.md b/docs/archive/tutorials/best-practices-integration/geospatial-data-mesh-cost-capacity.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/geospatial-data-mesh-cost-capacity.md
rename to docs/archive/tutorials/best-practices-integration/geospatial-data-mesh-cost-capacity.md
diff --git a/docs/tutorials/best-practices-integration/high-performance-caching-topology.md b/docs/archive/tutorials/best-practices-integration/high-performance-caching-topology.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/high-performance-caching-topology.md
rename to docs/archive/tutorials/best-practices-integration/high-performance-caching-topology.md
diff --git a/docs/tutorials/best-practices-integration/high-performance-geokg-async-caching.md b/docs/archive/tutorials/best-practices-integration/high-performance-geokg-async-caching.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/high-performance-geokg-async-caching.md
rename to docs/archive/tutorials/best-practices-integration/high-performance-geokg-async-caching.md
diff --git a/docs/archive/tutorials/best-practices-integration/index.md b/docs/archive/tutorials/best-practices-integration/index.md
new file mode 100644
index 0000000..2d0f459
--- /dev/null
+++ b/docs/archive/tutorials/best-practices-integration/index.md
@@ -0,0 +1,10 @@
+# Archived: Best Practices Integration tutorials
+
+These walkthroughs combined many topics into single “full stack” narratives. They read like generated course material and are **not maintained** as part of the public tutorial set.
+
+Use instead:
+
+- [Tutorials index](../../../tutorials/index.md)
+- [Best Practices index](../../../best-practices/index.md)
+
+Individual archived files remain in this directory for reference during cleanup.
diff --git a/docs/tutorials/best-practices-integration/multi-environment-deployment-secrets-config-governance.md b/docs/archive/tutorials/best-practices-integration/multi-environment-deployment-secrets-config-governance.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/multi-environment-deployment-secrets-config-governance.md
rename to docs/archive/tutorials/best-practices-integration/multi-environment-deployment-secrets-config-governance.md
diff --git a/docs/tutorials/best-practices-integration/multi-region-dr-temporal-blast-radius.md b/docs/archive/tutorials/best-practices-integration/multi-region-dr-temporal-blast-radius.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/multi-region-dr-temporal-blast-radius.md
rename to docs/archive/tutorials/best-practices-integration/multi-region-dr-temporal-blast-radius.md
diff --git a/docs/tutorials/best-practices-integration/multi-region-geokg-temporal-dr.md b/docs/archive/tutorials/best-practices-integration/multi-region-geokg-temporal-dr.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/multi-region-geokg-temporal-dr.md
rename to docs/archive/tutorials/best-practices-integration/multi-region-geokg-temporal-dr.md
diff --git a/docs/tutorials/best-practices-integration/polyglot-streaming-saga-cqrs-multicloud.md b/docs/archive/tutorials/best-practices-integration/polyglot-streaming-saga-cqrs-multicloud.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/polyglot-streaming-saga-cqrs-multicloud.md
rename to docs/archive/tutorials/best-practices-integration/polyglot-streaming-saga-cqrs-multicloud.md
diff --git a/docs/tutorials/best-practices-integration/progressive-delivery-release-fitness.md b/docs/archive/tutorials/best-practices-integration/progressive-delivery-release-fitness.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/progressive-delivery-release-fitness.md
rename to docs/archive/tutorials/best-practices-integration/progressive-delivery-release-fitness.md
diff --git a/docs/tutorials/best-practices-integration/quality-governed-geokg-retention-lifecycle.md b/docs/archive/tutorials/best-practices-integration/quality-governed-geokg-retention-lifecycle.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/quality-governed-geokg-retention-lifecycle.md
rename to docs/archive/tutorials/best-practices-integration/quality-governed-geokg-retention-lifecycle.md
diff --git a/docs/tutorials/best-practices-integration/resilient-microservices-chaos-engineering.md b/docs/archive/tutorials/best-practices-integration/resilient-microservices-chaos-engineering.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/resilient-microservices-chaos-engineering.md
rename to docs/archive/tutorials/best-practices-integration/resilient-microservices-chaos-engineering.md
diff --git a/docs/tutorials/best-practices-integration/secure-polyglot-identity-encryption.md b/docs/archive/tutorials/best-practices-integration/secure-polyglot-identity-encryption.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/secure-polyglot-identity-encryption.md
rename to docs/archive/tutorials/best-practices-integration/secure-polyglot-identity-encryption.md
diff --git a/docs/tutorials/best-practices-integration/semantic-knowledge-graph-rdf-lineage.md b/docs/archive/tutorials/best-practices-integration/semantic-knowledge-graph-rdf-lineage.md
similarity index 100%
rename from docs/tutorials/best-practices-integration/semantic-knowledge-graph-rdf-lineage.md
rename to docs/archive/tutorials/best-practices-integration/semantic-knowledge-graph-rdf-lineage.md
diff --git a/docs/assets/overrides/main.html b/docs/assets/overrides/main.html
index 524d44b..8bbc215 100644
--- a/docs/assets/overrides/main.html
+++ b/docs/assets/overrides/main.html
@@ -1,26 +1,7 @@
{% extends "base.html" %}
{% block extrahead %}
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-
-{% endblock %}
-
-{% block content %}
- {{ super() }}
-
-
-
{% endblock %}
diff --git a/docs/assets/overrides/partials/footer.html b/docs/assets/overrides/partials/footer.html
deleted file mode 100644
index 99eacac..0000000
--- a/docs/assets/overrides/partials/footer.html
+++ /dev/null
@@ -1,57 +0,0 @@
-
-
diff --git a/docs/assets/overrides/partials/header.html b/docs/assets/overrides/partials/header.html
deleted file mode 100644
index efd931f..0000000
--- a/docs/assets/overrides/partials/header.html
+++ /dev/null
@@ -1,85 +0,0 @@
-
-
diff --git a/docs/assets/styles/custom.css b/docs/assets/styles/custom.css
index 050c0f2..5c8692b 100644
--- a/docs/assets/styles/custom.css
+++ b/docs/assets/styles/custom.css
@@ -252,17 +252,30 @@
margin: 1.5em 0;
}
-.project-card {
- background: var(--md-default-bg-color);
- border: 1px solid var(--md-default-fg-color--lightest);
+[data-md-color-scheme="default"] {
+ --project-card-bg: hsl(0, 0%, 100%);
+ --project-card-fg: hsla(0, 0%, 0%, 0.87);
+ --project-card-border: hsla(0, 0%, 0%, 0.12);
+}
+
+[data-md-color-scheme="slate"] {
+ --project-card-bg: hsl(232, 7%, 16%);
+ --project-card-fg: hsla(0, 0%, 100%, 0.87);
+ --project-card-border: hsla(0, 0%, 100%, 0.12);
+}
+
+.md-typeset .project-card {
+ background: var(--project-card-bg);
+ color: var(--project-card-fg);
+ border: 1px solid var(--project-card-border);
border-radius: 8px;
- padding: 1.25em 1.35em;
- box-shadow: 0 1px 3px rgba(0, 0, 0, 0.05);
+ padding: 1em 1.15em;
+ box-shadow: none;
}
.project-card h3 {
margin-top: 0;
- font-size: 1.15em;
+ font-size: 1.1em;
color: var(--md-primary-fg-color);
}
@@ -274,9 +287,9 @@
.project-card__image {
display: block;
width: 100%;
- max-width: 140px;
+ max-width: 120px;
height: auto;
- margin: 0 auto 0.75em;
+ margin: 0 auto 0.5em;
border-radius: 6px;
}
diff --git a/docs/best-practices/architecture-design/adr-decision-governance.md b/docs/best-practices/architecture-design/adr-decision-governance.md
deleted file mode 100644
index 1510a67..0000000
--- a/docs/best-practices/architecture-design/adr-decision-governance.md
+++ /dev/null
@@ -1,1599 +0,0 @@
-# Best Practices for Architecture Decision Records (ADRs) and Technical Decision Governance
-
-**Objective**: Establish a consistent, discoverable practice for recording and governing architecture and technical decisions. When you need to understand why a system is built a certain way, when you want to make decisions explicit and reviewable, when you need to prevent architectural amnesia—this guide provides the framework.
-
-## Goals & Non-Goals
-
-### Goals
-
-- **Create a consistent practice** for recording architecture and technical decisions
-- **Make decisions discoverable, reviewable, and auditable** over time
-- **Integrate ADRs into existing Git-based workflows** (PRs, issues, reviews)
-- **Reduce tribal knowledge** and forgotten decisions
-- **Enable better onboarding** by documenting "why" alongside "how"
-
-### Non-Goals
-
-- **Not replacing detailed design docs or RFCs**: ADRs capture decisions; design docs capture implementation details
-- **Not mandating a single ADR tooling stack**: Focus on patterns and structure that work with any tooling
-- **Not requiring ADRs for every change**: Only significant architectural decisions
-- **Not a replacement for code review**: ADRs complement, don't replace, code review
-
-## Why ADRs Matter (Especially in This Kind of Repo)
-
-### The Problem: Architectural Amnesia
-
-Your repository already contains extensive best-practices documentation:
-- Docker and containerization patterns
-- PostgreSQL optimization and security
-- NGINX configuration and routing
-- Observability stacks (Grafana, Prometheus, Loki)
-- SBOM and CVE mitigation workflows
-- Geospatial data engineering
-- ML/AI deployment patterns
-
-**These documents answer "how"**—they describe proven patterns and implementation details.
-
-**What's missing is "why"**—the context, trade-offs, and rationale behind each decision.
-
-### Problems Without ADRs
-
-1. **Decisions live in ephemeral channels**: Chat, email, or someone's memory
-2. **Repeated questions**: New contributors ask "Why Postgres here and not X?", "Why Loki over ELK?", "Why PgAudit instead of native logging?"
-3. **Architectural thrash**: Teams make conflicting changes because original decisions are forgotten
-4. **Onboarding friction**: New engineers spend weeks understanding "why things are this way"
-5. **Compliance gaps**: Auditors ask "Why this security posture?" and there's no documented answer
-
-### The Value of ADRs
-
-ADRs provide:
-- **Historical context**: Why a decision was made at a specific point in time
-- **Trade-off transparency**: What alternatives were considered and why they were rejected
-- **Decision audit trail**: Who decided, when, and what the expected consequences were
-- **Living documentation**: Decisions can be superseded, but history is preserved
-
-### Example: The PgAudit Decision
-
-**Without ADR**: A new engineer sees PgAudit configured and wonders:
-- "Why not use native PostgreSQL logging?"
-- "Why CSV logs instead of JSON?"
-- "Why PgCron for rotation instead of logrotate?"
-
-**With ADR**: ADR-0005 documents:
-- Context: Compliance requirements for audit trails
-- Options: Native logging, PgAudit, third-party tools
-- Decision: PgAudit + CSV logs + PgCron
-- Rationale: Structured parsing, automated rotation, database-native scheduling
-- Consequences: Performance overhead, maintenance complexity, compliance coverage
-
-## ADR Basics: What, When, and How Much Detail
-
-### What Is an ADR?
-
-An **Architecture Decision Record (ADR)** is a short, versioned document that captures:
-- **A single significant decision**
-- **The context** that led to the decision
-- **Options considered** and why they were rejected
-- **The decision** itself
-- **Expected consequences** (positive and negative)
-
-### When to Write an ADR
-
-Write an ADR when you make a decision that:
-- **Introduces a new major technology** (e.g., switching from ELK to Loki)
-- **Changes security posture** (e.g., adopting SBOM/Trivy policy)
-- **Affects data model or storage architecture** (e.g., choosing Parquet over JSONB)
-- **Defines a standard pattern** (e.g., GitFlow workflow, NGINX routing strategy)
-- **Will be questioned later** ("Why did we choose X instead of Y?")
-
-### When NOT to Write an ADR
-
-Skip ADRs for:
-- **Trivial refactors**: Code cleanup, variable renaming
-- **Local implementation details**: Covered by code comments
-- **Temporary workarounds**: Documented in issue trackers
-- **Obvious choices**: When there's only one reasonable option
-
-### Decision Checklist: "Should This Be an ADR?"
-
-Ask yourself:
-1. **Will a future engineer ask "why?"** about this decision?
-2. **Does this affect multiple systems or teams?**
-3. **Are there multiple reasonable alternatives?**
-4. **Will this decision be referenced in design docs or runbooks?**
-5. **Does this change our security, compliance, or operational posture?**
-6. **Is this a pattern we'll reuse across projects?**
-7. **Would reversing this decision be expensive or disruptive?**
-
-If you answer "yes" to 3+ questions, write an ADR.
-
-## ADR Repository Structure & Naming
-
-### Recommended Structure
-
-**For Documentation Repositories** (like `sempervent.github.io`):
-
-```
-/docs
- /adr
- /_templates
- 0001-template-standard.md
- 0002-template-lightweight.md
- 0001-use-postgres-for-audit-log.md
- 0002-standardize-on-openmaptiles-for-basemaps.md
- 0003-sbom-trivy-policy-for-docker-images.md
- 0004-observability-stack-standardization.md
- 0005-pgaudit-pgcron-audit-pipeline.md
- index.md
-```
-
-**For Code Repositories**:
-
-```
-/adr
- 0001-title.md
- 0002-title.md
- index.md
-```
-
-### Naming Convention
-
-**Format**: `NNNN-kebab-case-title.md`
-
-- **Sequential numeric prefix**: `0001-`, `0002-`, etc. (padded to 4 digits)
-- **Kebab-case title**: Lowercase, hyphens for spaces
-- **No gaps**: If ADR-0005 is deleted, don't reuse the number; keep sequence continuous
-
-**Examples**:
-- `0001-use-postgres-for-audit-log.md`
-- `0002-standardize-on-openmaptiles-for-basemaps.md`
-- `0003-sbom-trivy-policy-for-docker-images.md`
-
-### Status Labels
-
-Each ADR must have a status:
-
-- **`Proposed`**: Draft, under review, not yet accepted
-- **`Accepted`**: Decision is active and should be followed
-- **`Superseded`**: Replaced by a newer ADR (include link to superseding ADR)
-- **`Rejected`**: Decision was not adopted (include reason)
-- **`Deprecated`**: Decision is no longer relevant but kept for historical context
-
-### Linking Related ADRs
-
-Use explicit links in ADR documents:
-
-```markdown
-## Related ADRs
-
-- **Supersedes**: [ADR-0003: Legacy Observability Stack](../adr/0003-legacy-observability-stack.md)
-- **Superseded by**: [ADR-0010: Unified Observability Platform](../adr/0010-unified-observability-platform.md)
-- **Related**: [ADR-0005: Database Auditing Strategy](../adr/0005-database-auditing-strategy.md)
-```
-
-## ADR Templates
-
-### Standard ADR Template
-
-Use this template for most architectural decisions:
-
-```markdown
-# ADR-XXXX: [Short Descriptive Title]
-
-**Status**: Proposed | Accepted | Superseded | Rejected | Deprecated
-
-**Date**: YYYY-MM-DD
-
-**Deciders**: [Names or roles, e.g., "Architecture Working Group", "Tech Leads"]
-
-**Tags**: [Optional: component, domain, e.g., "observability", "security", "data"]
-
----
-
-## Context
-
-[2-4 paragraphs describing the situation, problem, or requirement that led to this decision. Include:
-- What problem are we solving?
-- What constraints exist?
-- What are the business/technical drivers?]
-
-## Decision
-
-[1-2 paragraphs stating the decision clearly and concisely. Be specific about what is being decided.]
-
-We will [specific action/choice].
-
-## Options Considered
-
-### Option 1: [Name of Option]
-
-**Description**: [What this option entails]
-
-**Pros**:
-- [Benefit 1]
-- [Benefit 2]
-
-**Cons**:
-- [Drawback 1]
-- [Drawback 2]
-
-### Option 2: [Name of Option]
-
-[Same structure as Option 1]
-
-### Option 3: [Name of Option]
-
-[Same structure as Option 1]
-
-## Rationale
-
-[2-3 paragraphs explaining why the chosen option was selected. Reference specific pros/cons, constraints, or requirements that drove the decision.]
-
-## Consequences
-
-### Positive
-
-- [Expected positive outcome 1]
-- [Expected positive outcome 2]
-
-### Negative
-
-- [Expected negative outcome 1]
-- [Expected negative outcome 2]
-
-### Neutral / Trade-offs
-
-- [Trade-off or neutral consequence]
-
-## Related Documents
-
-- [Link to design doc, RFC, or implementation guide]
-- [Link to best-practices document]
-- [Link to runbook or operational guide]
-
-## Implementation Notes
-
-[Optional: Brief notes on how this decision will be implemented, key milestones, or dependencies]
-
-## References
-
-[Optional: Links to external resources, research, or benchmarks that informed the decision]
-```
-
-### Lightweight ADR Template
-
-Use this template for smaller, focused decisions:
-
-```markdown
-# ADR-XXXX: [Short Descriptive Title]
-
-**Status**: Proposed | Accepted | Superseded | Rejected | Deprecated
-
-**Date**: YYYY-MM-DD
-
-**Deciders**: [Names or roles]
-
----
-
-## Context
-
-[1-2 paragraphs: What problem or situation led to this decision?]
-
-## Decision
-
-[1 paragraph: What is being decided?]
-
-We will [specific action/choice].
-
-## Impact / Notes
-
-[1-2 paragraphs: What are the key consequences, trade-offs, or implementation considerations?]
-
-## Related Documents
-
-- [Link to relevant docs]
-```
-
-## Decision Scope & Granularity
-
-### Best Practices
-
-**Each ADR should cover one main decision**. Avoid bundling unrelated choices into a single ADR.
-
-### Scope Guidelines
-
-| Scope | Example | Good/Bad |
-|-------|---------|----------|
-| **Too Big** | "Observability, Logging, Monitoring, and All Data Pipelines v3" | ❌ Too broad, multiple unrelated decisions |
-| **Just Right** | "Standardize on Grafana + Prometheus + Loki for Observability" | ✅ Single coherent decision |
-| **Just Right** | "Use PgAudit + CSV Logs + PgCron for Audit Logging Pipeline" | ✅ Related components, single purpose |
-| **Just Right** | "Adopt GitFlow with Protected Main/Develop and Release Branches" | ✅ Single workflow decision |
-| **Just Right** | "Generate Dark OpenMapTiles for US at Z≤12 as Default Basemap" | ✅ Single technology/pattern decision |
-| **Too Small** | "Use `asyncpg` instead of `psycopg2` in service X" | ❌ Implementation detail, not architectural |
-
-### Good ADR Examples (From This Ecosystem)
-
-1. **"Standardize on Grafana + Prometheus + Loki for Observability"**
- - Single decision: Which observability stack?
- - Clear scope: Metrics, logs, dashboards
- - Related but distinct: Could have separate ADRs for alerting, tracing
-
-2. **"Use PgAudit + CSV Logs + PgCron for Audit Logging Pipeline"**
- - Single decision: How to implement database auditing?
- - Related components: All part of one pipeline
- - Clear boundary: Database auditing, not application logging
-
-3. **"Adopt GitFlow with Protected Main/Develop and Release Branches"**
- - Single decision: Which branching model?
- - Clear scope: Version control workflow
- - Related but distinct: Could have separate ADR for CI/CD integration
-
-4. **"Generate Dark OpenMapTiles for US at Z≤12 as Default Basemap"**
- - Single decision: Which basemap technology and configuration?
- - Clear scope: Geospatial visualization
- - Related but distinct: Could have separate ADR for tile serving infrastructure
-
-### When to Split or Combine
-
-**Split if**:
-- ADR covers multiple unrelated systems (e.g., "Database and Message Queue Strategy")
-- Different teams own different parts
-- Decisions can be made independently
-
-**Combine if**:
-- Components are tightly coupled (e.g., PgAudit + CSV logs + PgCron)
-- One decision depends on another
-- They form a single coherent pattern
-
-## Integrating ADRs with Git & Pull Requests
-
-### Workflow: Creating a New ADR
-
-**Step 1: Create Branch**
-
-```bash
-# Create feature branch
-git checkout -b adr/0007-sbom-trivy-policy
-
-# Or use descriptive name
-git checkout -b adr/0007-docker-image-security-policy
-```
-
-**Step 2: Create ADR File**
-
-```bash
-# Create ADR file
-touch docs/adr/0007-sbom-trivy-policy-for-docker-images.md
-
-# Copy template
-cp docs/adr/_templates/0001-template-standard.md docs/adr/0007-sbom-trivy-policy-for-docker-images.md
-```
-
-**Step 3: Draft ADR**
-
-- Set status to `Proposed`
-- Fill in context, options, decision, rationale
-- Link to related documents
-
-**Step 4: Update ADR Index**
-
-```markdown
-# Architecture Decision Records
-
-| Number | Title | Status | Date |
-|--------|-------|--------|------|
-| 0007 | SBOM and Trivy Policy for Docker Images | Proposed | 2024-01-15 |
-```
-
-### Workflow: Review Process
-
-**Step 1: Open Pull Request**
-
-**PR Title**: `ADR-0007: SBOM and Trivy Policy for Docker Images`
-
-**PR Description Template**:
-
-```markdown
-## ADR Proposal
-
-This PR proposes [brief summary of decision].
-
-**Status**: Proposed
-
-**Deciders**: @tech-lead-1, @security-lead, @ops-lead
-
-**Related Issues**: #123, #456
-
-## Review Checklist
-
-- [ ] Context and problem statement are clear
-- [ ] All reasonable options have been considered
-- [ ] Rationale is well-documented
-- [ ] Consequences (positive and negative) are identified
-- [ ] Related documents are linked
-- [ ] ADR index is updated
-
-## Questions for Reviewers
-
-- [Question 1]
-- [Question 2]
-```
-
-**Step 2: Request Review**
-
-Request review from:
-- **Tech leads / architects**: Validate technical soundness
-- **Stakeholders**: Security, ops, data teams affected by decision
-- **Domain experts**: People with relevant expertise
-
-**Step 3: Address Feedback**
-
-- Update ADR based on review comments
-- Document any changes in PR comments
-- Keep status as `Proposed` until consensus
-
-**Step 4: Acceptance**
-
-Once approved:
-- Change status to `Accepted`
-- Merge PR
-- Optionally tag repository with release that includes decision
-
-### Commit Message Convention
-
-**Format**: `docs(adr): add ADR-XXXX: [Title]`
-
-**Examples**:
-
-```bash
-git commit -m "docs(adr): add ADR-0007: SBOM and Trivy Policy for Docker Images
-
-- Proposes mandatory SBOM generation for all Docker images
-- Requires Trivy scan with CRITICAL/HIGH severity gates
-- Establishes automated CVE mitigation workflow
-
-Status: Proposed"
-```
-
-### Workflow: Superseding an ADR
-
-**Step 1: Create New ADR**
-
-```bash
-git checkout -b adr/0010-unified-observability-platform
-touch docs/adr/0010-unified-observability-platform.md
-```
-
-**Step 2: Reference Old ADR**
-
-In new ADR:
-
-```markdown
-## Related ADRs
-
-- **Supersedes**: [ADR-0004: Observability Stack Standardization](../adr/0004-observability-stack-standardization.md)
-```
-
-**Step 3: Update Old ADR**
-
-In old ADR:
-
-```markdown
-**Status**: Superseded
-
-## Related ADRs
-
-- **Superseded by**: [ADR-0010: Unified Observability Platform](../adr/0010-unified-observability-platform.md)
-```
-
-**Step 4: Merge Both Changes**
-
-- New ADR: Status `Accepted`
-- Old ADR: Status `Superseded`
-
-## Linking ADRs to Best-Practices & Runbooks
-
-### Best-Practices Documents Should Reference ADRs
-
-**Pattern**: Add ADR references at the top or bottom of best-practices documents.
-
-**Example: NGINX Best Practices**
-
-```markdown
-# NGINX Best Practices: Patterns, Hardening, and Multi-API Replication
-
-**This document is guided by**: [ADR-0003: NGINX as Edge Reverse Proxy for Internal APIs](../adr/0003-nginx-edge-reverse-proxy.md)
-
-[Rest of document...]
-
----
-
-## Related Documents
-
-- [ADR-0003: NGINX as Edge Reverse Proxy](../adr/0003-nginx-edge-reverse-proxy.md)
-- [ADR-0008: Multi-API Gateway Pattern](../adr/0008-multi-api-gateway-pattern.md)
-```
-
-**Example: Observability Best Practices**
-
-```markdown
-# Best Practices for Monitoring & Observability
-
-**This approach is defined by**: [ADR-0004: Observability Stack Standardization](../adr/0004-observability-stack-standardization.md)
-
-[Rest of document...]
-```
-
-### ADRs Should Link to Implementation Guides
-
-In ADR's "Related Documents" section:
-
-```markdown
-## Related Documents
-
-- **[Best Practices: Grafana, Prometheus, Loki](../operations-monitoring/grafana-prometheus-loki-observability.md)** - Implementation guide
-- **[Tutorial: Setting Up Prometheus Scraping](../tutorials/operations-monitoring/prometheus-setup.md)** - Step-by-step setup
-- **[Runbook: Observability Stack Incident Response](../runbooks/observability-incident-response.md)** - Operational procedures
-```
-
-### Handling Divergence
-
-If best-practices diverge from ADR:
-
-```markdown
-## Note on ADR Alignment
-
-This document follows [ADR-0004: Observability Stack Standardization](../adr/0004-observability-stack-standardization.md) with one exception:
-
-**Divergence**: We use Promtail instead of Fluentd for log ingestion.
-
-**Reason**: Promtail provides better integration with Loki and reduces operational complexity in our Kubernetes environment.
-
-**Impact**: This is an implementation detail that doesn't change the core decision in ADR-0004.
-```
-
-## Decision Lifecycle & Governance
-
-### Governance Model
-
-**Who Can Propose ADRs?**
-
-- **Any engineer** can propose an ADR
-- No gatekeeping on proposal; all proposals are welcome
-
-**Who Can Approve ADRs?**
-
-- **Architecture Working Group** (or equivalent): For system-wide decisions
-- **Tech Leads Council**: For cross-team decisions
-- **Domain Experts**: For domain-specific decisions (e.g., security team for security ADRs)
-- **Consensus-based**: For smaller teams, require approval from 2-3 relevant stakeholders
-
-**Decision Authority Matrix**:
-
-| Decision Scope | Approver | Example |
-|----------------|----------|---------|
-| **System-wide** | Architecture Working Group | Observability stack, database strategy |
-| **Security** | Security Team + Tech Leads | SBOM policy, audit logging |
-| **Domain-specific** | Domain Experts + Tech Leads | Geospatial basemap, ML deployment |
-| **Team-specific** | Team Tech Lead | Service-specific patterns |
-
-### Decision Lifecycle
-
-```mermaid
-graph LR
- A[Proposed] --> B{In Review}
- B -->|Approved| C[Accepted]
- B -->|Rejected| D[Rejected]
- C -->|Superseded| E[Superseded]
- C -->|No longer relevant| F[Deprecated]
-
- style A fill:#fff4cc
- style B fill:#cfe2f3
- style C fill:#d9ead3
- style D fill:#f4cccc
- style E fill:#e1d5e7
- style F fill:#fce5cd
-```
-
-**States**:
-- **Proposed**: Draft, under review
-- **In Review**: Actively being discussed (optional explicit state)
-- **Accepted**: Decision is active
-- **Rejected**: Decision was not adopted
-- **Superseded**: Replaced by newer ADR
-- **Deprecated**: No longer relevant but kept for history
-
-### Review Cadence
-
-**Quarterly ADR Review Meeting**:
-
-**Agenda**:
-1. **Review stale ADRs**: Identify ADRs that may need updating
-2. **Mark as Deprecated**: ADRs for systems that no longer exist
-3. **Identify Superseded ADRs**: Find ADRs that have been implicitly replaced
-4. **Update statuses**: Ensure all ADRs have correct status
-5. **Archive old ADRs**: Move deprecated ADRs to archive (optional)
-
-**Process**:
-- Tech leads review ADR index
-- Identify candidates for status changes
-- Discuss in meeting
-- Update ADRs via PRs
-
-### Handling Disagreements
-
-**Use ADRs to make disagreements explicit**:
-
-1. **Capture competing options**: Document all serious alternatives
-2. **Record rationale**: Explain why each option was considered
-3. **Document trade-offs**: Be explicit about what's being traded off
-4. **Record the compromise**: The decision is the documented compromise
-
-**Example**:
-
-```markdown
-## Options Considered
-
-### Option 1: ELK Stack (Elasticsearch, Logstash, Kibana)
-
-**Pros**: Mature, feature-rich, excellent search capabilities
-
-**Cons**: Resource-intensive, complex to operate, licensing costs at scale
-
-### Option 2: Grafana + Prometheus + Loki
-
-**Pros**: Lightweight, cost-effective, excellent Grafana integration
-
-**Cons**: Less mature than ELK, different query language (LogQL vs Lucene)
-
-## Rationale
-
-After evaluation, we chose Option 2 (Grafana + Prometheus + Loki) because:
-- Cost constraints favor open-source, self-hosted solutions
-- Existing Prometheus investment reduces operational overhead
-- LogQL, while different, is sufficient for our use cases
-- Team familiarity with Grafana reduces training burden
-
-**Trade-off**: We sacrifice some advanced search capabilities for operational simplicity and cost efficiency.
-```
-
-## Concrete Examples
-
-### Example 1: Standardize on PgAudit + CSV Logs + PgCron for Database Auditing
-
-```markdown
-# ADR-0005: Use PgAudit + CSV Logs + PgCron for Database Auditing
-
-**Status**: Accepted
-
-**Date**: 2024-01-15
-
-**Deciders**: Database Team, Security Team, Architecture Working Group
-
-**Tags**: database, security, compliance
-
----
-
-## Context
-
-We need comprehensive audit logging for PostgreSQL databases to meet compliance requirements (SOC2, GDPR). The audit trail must:
-- Log all DDL, DML, and function calls
-- Be queryable and analyzable
-- Rotate automatically to manage disk space
-- Integrate with our existing observability stack
-
-Current PostgreSQL native logging provides basic query logging but lacks structured audit events and automated rotation capabilities.
-
-## Decision
-
-We will use **PgAudit** for structured audit logging, **CSV logs** for parsing and ingestion, and **PgCron** for automated log rotation and data refresh.
-
-## Options Considered
-
-### Option 1: Native PostgreSQL Logging
-
-**Description**: Use PostgreSQL's built-in `log_statement` and `log_min_duration_statement` settings.
-
-**Pros**:
-- No additional extensions required
-- Simple configuration
-- Low overhead
-
-**Cons**:
-- No structured audit events (just query text)
-- Difficult to parse and analyze
-- No automated rotation (requires external tools)
-- Limited filtering capabilities
-
-### Option 2: PgAudit + External Log Rotation
-
-**Description**: Use PgAudit for audit logging, but rely on OS-level logrotate for rotation.
-
-**Pros**:
-- Structured audit events
-- Standard OS tooling
-- No database extension for rotation
-
-**Cons**:
-- Requires coordination between database and OS
-- Less flexible for database-specific retention policies
-- Harder to integrate with database-driven refresh jobs
-
-### Option 3: PgAudit + CSV Logs + PgCron (Chosen)
-
-**Description**: Use PgAudit for audit events, CSV logs for structured parsing, and PgCron for automated rotation and data refresh.
-
-**Pros**:
-- Structured audit events via PgAudit
-- CSV format enables easy parsing and ingestion
-- Database-native scheduling via PgCron
-- Can build analytical views on audit data
-- Automated refresh keeps audit tables current
-
-**Cons**:
-- Requires two extensions (pgaudit, pg_cron)
-- Performance overhead from audit logging (5-15%)
-- More complex setup and maintenance
-
-## Rationale
-
-We chose Option 3 because:
-- **Compliance requirements** demand structured, queryable audit trails
-- **CSV logs** enable us to build analytical views (top users, top tables, activity logs)
-- **PgCron** provides database-native scheduling, reducing operational complexity
-- **Integrated approach** allows us to refresh audit tables automatically and build dashboards
-
-The performance overhead is acceptable given compliance requirements, and we can tune PgAudit settings (e.g., `pgaudit.log = 'ddl,write'` instead of `'all'`) to reduce impact.
-
-## Consequences
-
-### Positive
-
-- Comprehensive audit trail for compliance
-- Queryable audit data via database views
-- Automated rotation and refresh reduce operational burden
-- Can build dashboards and alerts on audit data
-
-### Negative
-
-- Performance overhead (5-15% depending on logging level)
-- Additional maintenance for PgAudit and PgCron extensions
-- Storage requirements for audit logs and parsed data
-
-### Neutral / Trade-offs
-
-- More complex setup, but better long-term maintainability
-- Database-native scheduling vs OS-level tooling
-
-## Related Documents
-
-- **[Best Practices: PostgreSQL Security](../postgres/postgres-security-best-practices.md)**
-- **[Tutorial: Auditing PostgreSQL with PgAudit and PgCron](../../tutorials/database-data-engineering/postgres-pgaudit-pgcron-auditing.md)**
-
-## Implementation Notes
-
-- Enable PgAudit and PgCron in `postgresql.conf`
-- Configure CSV logging with appropriate rotation settings
-- Create audit schema with staging and parsed log tables
-- Build refresh function and analytical views
-- Schedule PgCron jobs for rotation and refresh
-```
-
-### Example 2: Adopt Grafana + Prometheus + Loki as Default Observability Stack
-
-```markdown
-# ADR-0004: Standardize on Grafana + Prometheus + Loki for Observability
-
-**Status**: Accepted
-
-**Date**: 2024-01-10
-
-**Deciders**: Architecture Working Group, Operations Team
-
-**Tags**: observability, monitoring, logging
-
----
-
-## Context
-
-We need a unified observability stack for metrics, logs, and dashboards across all systems (infrastructure, applications, data pipelines). Current state:
-- Metrics scattered across multiple systems
-- Logs in various formats and locations
-- No unified dashboarding solution
-- High operational overhead from multiple tools
-
-Requirements:
-- Cost-effective (prefer open-source, self-hosted)
-- Scalable to hundreds of services
-- Easy integration with existing systems
-- Strong visualization capabilities
-
-## Decision
-
-We will standardize on **Grafana** for dashboards, **Prometheus** for metrics, and **Loki** for logs. All new systems must integrate with this stack.
-
-## Options Considered
-
-### Option 1: ELK Stack (Elasticsearch, Logstash, Kibana)
-
-**Description**: Use Elasticsearch for logs, Logstash for processing, Kibana for dashboards. Add Prometheus for metrics.
-
-**Pros**:
-- Mature, feature-rich
-- Excellent search capabilities (Lucene)
-- Strong ecosystem
-
-**Cons**:
-- Resource-intensive (Elasticsearch requires significant RAM/CPU)
-- Complex to operate at scale
-- Licensing costs for advanced features
-- Different stack for logs vs metrics
-
-### Option 2: Datadog / New Relic / Splunk (SaaS)
-
-**Description**: Use commercial SaaS observability platform.
-
-**Pros**:
-- Fully managed, no operational overhead
-- Excellent UI and features
-- Strong support
-
-**Cons**:
-- High cost at scale (per-host or per-GB pricing)
-- Vendor lock-in
-- Data residency concerns for sensitive data
-- Limited customization
-
-### Option 3: Grafana + Prometheus + Loki (Chosen)
-
-**Description**: Use Grafana for visualization, Prometheus for metrics, Loki for logs.
-
-**Pros**:
-- Cost-effective (open-source, self-hosted)
-- Unified stack (metrics and logs in Grafana)
-- PromQL and LogQL are powerful query languages
-- Excellent Grafana integration
-- Can reuse existing Prometheus investment
-
-**Cons**:
-- Less mature than ELK for log search
-- LogQL different from Lucene (learning curve)
-- Requires operational expertise
-
-## Rationale
-
-We chose Option 3 because:
-- **Cost constraints** favor open-source, self-hosted solutions
-- **Unified stack** reduces operational complexity (one dashboard tool)
-- **Existing Prometheus investment** means we can leverage current setup
-- **LogQL**, while different from Lucene, is sufficient for our use cases
-- **Team familiarity** with Grafana reduces training burden
-
-The trade-off of less advanced log search is acceptable given cost and operational simplicity benefits.
-
-## Consequences
-
-### Positive
-
-- Unified observability stack reduces operational overhead
-- Cost-effective (no per-host or per-GB SaaS fees)
-- Strong Grafana integration for metrics and logs
-- Can scale horizontally
-
-### Negative
-
-- Requires operational expertise to run Prometheus and Loki
-- LogQL learning curve for teams used to Lucene
-- Less mature log search than ELK
-
-### Neutral / Trade-offs
-
-- Self-hosted vs managed: More control, more operational burden
-- LogQL vs Lucene: Different but sufficient for our needs
-
-## Related Documents
-
-- **[Best Practices: Monitoring & Observability](../operations-monitoring/grafana-prometheus-loki-observability.md)**
-- **[Tutorial: Setting Up Prometheus Scraping](../../tutorials/operations-monitoring/prometheus-setup.md)**
-
-## Implementation Notes
-
-- Deploy Prometheus for metrics collection
-- Deploy Loki for log aggregation
-- Deploy Grafana with Prometheus and Loki datasources
-- Migrate existing dashboards to Grafana
-- Train teams on PromQL and LogQL
-```
-
-### Example 3: Use OpenMapTiles Dark Basemap (US z≤12) for All Geospatial Dashboards
-
-```markdown
-# ADR-0002: Standardize on OpenMapTiles Dark Basemap for Geospatial Dashboards
-
-**Status**: Accepted
-
-**Date**: 2024-01-05
-
-**Deciders**: Data Engineering Team, Frontend Team
-
-**Tags**: geospatial, visualization, frontend
-
----
-
-## Context
-
-We need a consistent basemap for all geospatial dashboards (Grafana, custom web apps, internal tools). Requirements:
-- Dark theme (reduces eye strain, matches our dashboard aesthetic)
-- Covers entire United States
-- Vector tiles (for performance and customization)
-- Self-hosted (for air-gapped environments and cost control)
-- Zoom level 12 sufficient for strategic/overview maps
-
-Current state: Mixed basemaps (some use Mapbox, some use OpenStreetMap, some have no basemap).
-
-## Decision
-
-We will generate and self-host **dark-themed OpenMapTiles** covering the entire United States at zoom level 12, and use this as the default basemap for all geospatial dashboards.
-
-## Options Considered
-
-### Option 1: Mapbox (SaaS)
-
-**Description**: Use Mapbox's hosted vector tiles with dark style.
-
-**Pros**:
-- Fully managed, no operational overhead
-- High-quality tiles and styles
-- Easy integration
-
-**Cons**:
-- Cost at scale (per-request pricing)
-- Not suitable for air-gapped environments
-- Vendor lock-in
-
-### Option 2: OpenStreetMap Raster Tiles
-
-**Description**: Use standard OSM raster tiles (e.g., from tile.openstreetmap.org).
-
-**Pros**:
-- Free, no operational overhead
-- Simple integration
-
-**Cons**:
-- Not dark-themed
-- Raster tiles (larger, less customizable)
-- Rate limiting on public servers
-- Not self-hosted
-
-### Option 3: OpenMapTiles Dark (Self-Hosted) (Chosen)
-
-**Description**: Generate dark-themed OpenMapTiles for US at z≤12, host via tileserver-gl.
-
-**Pros**:
-- Self-hosted (air-gapped compatible, no per-request costs)
-- Dark theme matches dashboard aesthetic
-- Vector tiles (smaller, more customizable)
-- Full control over style and data
-
-**Cons**:
-- Requires initial tile generation (time and storage)
-- Operational overhead (hosting tileserver)
-- Limited to z≤12 (may need higher zoom for specific regions)
-
-## Rationale
-
-We chose Option 3 because:
-- **Air-gapped environments** require self-hosted solutions
-- **Dark theme** matches our dashboard aesthetic and reduces eye strain
-- **Cost control**: No per-request fees, one-time generation cost
-- **Zoom level 12** is sufficient for strategic/overview maps (can generate higher zoom for specific regions if needed)
-
-The operational overhead is acceptable given cost and flexibility benefits.
-
-## Consequences
-
-### Positive
-
-- Consistent basemap across all dashboards
-- Dark theme reduces eye strain
-- Self-hosted enables air-gapped deployments
-- No per-request costs
-
-### Negative
-
-- Initial tile generation requires time and storage
-- Operational overhead for tileserver
-- Limited to z≤12 (may need region-specific higher zoom tiles)
-
-### Neutral / Trade-offs
-
-- Self-hosted vs SaaS: More control, more operational burden
-- z≤12 vs higher zoom: Balances coverage with storage/performance
-
-## Related Documents
-
-- **[Tutorial: Generating Dark OpenMapTiles for US at Z12](../../tutorials/database-data-engineering/openmaptiles-us-dark-z12.md)**
-- **[Best Practices: Geospatial Data Engineering](../database-data/geospatial-data-engineering.md)**
-
-## Implementation Notes
-
-- Generate US-wide tiles at z≤12 using OpenMapTiles toolchain
-- Deploy tileserver-gl for tile serving
-- Create dark style JSON
-- Integrate into Grafana and custom web apps
-- Document tile generation process for future updates
-```
-
-### Example 4: Require SBOM + Trivy Scan for All Docker Images Before Release
-
-```markdown
-# ADR-0003: SBOM and Trivy Policy for Docker Images
-
-**Status**: Accepted
-
-**Date**: 2024-01-12
-
-**Deciders**: Security Team, DevOps Team, Architecture Working Group
-
-**Tags**: security, containers, compliance
-
----
-
-## Context
-
-We need to ensure container security and compliance with supply-chain security requirements. Current state:
-- Docker images built and deployed without security scanning
-- No Software Bill of Materials (SBOM) for containers
-- No automated CVE detection or mitigation
-- Compliance requirements (SOC2, SLSA) demand supply-chain visibility
-
-Requirements:
-- Generate SBOMs for all Docker images
-- Scan images for CVEs before release
-- Automate CVE mitigation where possible
-- Fail builds on critical vulnerabilities
-
-## Decision
-
-We will require **SBOM generation** and **Trivy scanning** for all Docker images before release. Builds will fail on CRITICAL or HIGH severity CVEs unless explicitly allowed via ignore list.
-
-## Options Considered
-
-### Option 1: Manual Scanning (Current State)
-
-**Description**: Developers manually run security scans before releases.
-
-**Pros**:
-- No tooling changes required
-- Flexible process
-
-**Cons**:
-- Inconsistent application
-- Easy to forget or skip
-- No automation
-- No SBOM generation
-
-### Option 2: GitHub/GitLab Native Scanners
-
-**Description**: Use built-in security scanning (e.g., GitHub Dependabot, GitLab Container Scanning).
-
-**Pros**:
-- Integrated into existing workflows
-- No additional tooling
-
-**Cons**:
-- Limited to specific platforms
-- Less control over policies
-- May not generate SBOMs in required format
-- Less flexible for custom workflows
-
-### Option 3: Trivy + Syft (Chosen)
-
-**Description**: Use Trivy for vulnerability scanning and Syft for SBOM generation, integrated into CI/CD pipelines.
-
-**Pros**:
-- Comprehensive scanning (vulnerabilities, misconfigurations, secrets)
-- SBOM generation in multiple formats (CycloneDX, SPDX)
-- Flexible policy enforcement
-- Works with any CI/CD system
-- Can automate CVE mitigation
-
-**Cons**:
-- Requires CI/CD integration
-- Operational overhead for policy management
-- May require base image updates to fix CVEs
-
-## Rationale
-
-We chose Option 3 because:
-- **Comprehensive scanning**: Trivy covers vulnerabilities, misconfigurations, and secrets
-- **SBOM generation**: Syft produces SBOMs in standard formats (CycloneDX, SPDX) for compliance
-- **Flexible policies**: Can enforce different policies per environment (warn in dev, fail in prod)
-- **Automation potential**: Can integrate with dependency update tools (Renovate, Dependabot) for auto-fix
-- **Platform-agnostic**: Works with any CI/CD system, not tied to specific platforms
-
-The operational overhead is acceptable given security and compliance benefits.
-
-## Consequences
-
-### Positive
-
-- Comprehensive security scanning before release
-- SBOMs enable supply-chain visibility and compliance
-- Automated policy enforcement reduces human error
-- Can integrate with automated CVE mitigation
-
-### Negative
-
-- CI/CD pipeline changes required
-- May block releases until CVEs are fixed
-- Operational overhead for policy management
-- Requires base image and dependency updates
-
-### Neutral / Trade-offs
-
-- Security vs velocity: May slow releases, but improves security posture
-- Automation vs flexibility: More automated, but requires policy definition
-
-## Related Documents
-
-- **[Best Practices: SBOMs, Trivy Scans, and Automated CVE Mitigation](../docker-infrastructure/docker-sbom-trivy-cve-mitigation.md)**
-- **[Tutorial: Setting Up Trivy in CI/CD](../../tutorials/docker-infrastructure/trivy-ci-cd-setup.md)**
-
-## Implementation Notes
-
-- Integrate Trivy and Syft into CI/CD pipelines
-- Configure severity thresholds (fail on CRITICAL/HIGH)
-- Set up SBOM artifact storage
-- Create ignore list process for unavoidable CVEs
-- Document policy exceptions and approval process
-```
-
-## Tooling & Automation
-
-### Pre-commit / CI Checks
-
-**Validate ADR Structure**:
-
-Create a simple script to validate ADR files:
-
-```bash
-#!/bin/bash
-# scripts/validate-adr.sh
-
-adr_file="$1"
-
-# Check required fields
-required_fields=("Status" "Date" "Deciders")
-missing_fields=()
-
-for field in "${required_fields[@]}"; do
- if ! grep -q "^\*\*${field}\*\*:" "$adr_file"; then
- missing_fields+=("$field")
- fi
-done
-
-if [ ${#missing_fields[@]} -ne 0 ]; then
- echo "ERROR: Missing required fields in $adr_file: ${missing_fields[*]}"
- exit 1
-fi
-
-# Check status is valid
-status=$(grep "^\*\*Status\*\*:" "$adr_file" | sed 's/.*: //')
-valid_statuses=("Proposed" "Accepted" "Superseded" "Rejected" "Deprecated")
-if [[ ! " ${valid_statuses[@]} " =~ " ${status} " ]]; then
- echo "ERROR: Invalid status '$status' in $adr_file"
- exit 1
-fi
-
-# Check naming convention
-filename=$(basename "$adr_file")
-if ! [[ "$filename" =~ ^[0-9]{4}-.*\.md$ ]]; then
- echo "ERROR: ADR filename must match pattern NNNN-kebab-case-title.md"
- exit 1
-fi
-
-echo "✓ ADR $adr_file is valid"
-```
-
-**GitHub Actions Example**:
-
-```yaml
-# .github/workflows/validate-adrs.yml
-name: Validate ADRs
-
-on:
- pull_request:
- paths:
- - 'docs/adr/*.md'
-
-jobs:
- validate:
- runs-on: ubuntu-latest
- steps:
- - uses: actions/checkout@v3
- - name: Validate ADRs
- run: |
- for adr in docs/adr/*.md; do
- if [[ "$adr" != "docs/adr/index.md" ]] && [[ "$adr" != "docs/adr/_templates"* ]]; then
- ./scripts/validate-adr.sh "$adr"
- fi
- done
-```
-
-### ADR Index Generation
-
-**Auto-Generate Index**:
-
-```bash
-#!/bin/bash
-# scripts/generate-adr-index.sh
-
-index_file="docs/adr/index.md"
-adr_dir="docs/adr"
-
-cat > "$index_file" << 'EOF'
-# Architecture Decision Records
-
-This directory contains Architecture Decision Records (ADRs) for this project.
-
-See [ADR Best Practices](../../best-practices/architecture-design/adr-decision-governance.md) for guidelines on creating and maintaining ADRs.
-
-## ADR Index
-
-| Number | Title | Status | Date | Deciders |
-|--------|-------|--------|------|----------|
-EOF
-
-# Extract ADR metadata
-for adr in "$adr_dir"/*.md; do
- if [[ "$adr" == "$index_file" ]] || [[ "$adr" == *"/_templates/"* ]]; then
- continue
- fi
-
- number=$(basename "$adr" | cut -d'-' -f1)
- title=$(grep "^# ADR-" "$adr" | sed 's/^# ADR-[0-9]*: //')
- status=$(grep "^\*\*Status\*\*:" "$adr" | sed 's/.*: //')
- date=$(grep "^\*\*Date\*\*:" "$adr" | sed 's/.*: //')
- deciders=$(grep "^\*\*Deciders\*\*:" "$adr" | sed 's/.*: //' | sed 's/|/\\|/g')
-
- if [ -n "$title" ]; then
- echo "| $number | [$title]($(basename "$adr")) | $status | $date | $deciders |" >> "$index_file"
- fi
-done
-
-echo "✓ Generated ADR index at $index_file"
-```
-
-**Make Target**:
-
-```makefile
-# Makefile
-
-.PHONY: adr-index
-adr-index:
- @./scripts/generate-adr-index.sh
-
-.PHONY: validate-adrs
-validate-adrs:
- @for adr in docs/adr/*.md; do \
- if [[ "$$adr" != "docs/adr/index.md" ]] && [[ "$$adr" != *"/_templates/"* ]]; then \
- ./scripts/validate-adr.sh "$$adr"; \
- fi \
- done
-```
-
-### Tagging and Labeling
-
-**GitHub/GitLab Labels**:
-
-Create labels for ADR-related issues and PRs:
-- `adr` - Architecture Decision Record
-- `architecture` - Architectural change
-- `adr/proposed` - ADR under review
-- `adr/accepted` - ADR accepted
-
-**PR Template**:
-
-```markdown
-## Type of Change
-
-- [ ] Bug fix
-- [ ] New feature
-- [ ] ADR (Architecture Decision Record)
-- [ ] Documentation
-- [ ] Other
-
-## ADR Details (if applicable)
-
-- ADR Number: ADR-XXXX
-- Status: Proposed | Accepted
-- Related Issues: #123
-```
-
-### Dashboards and Search
-
-**ADR Metadata Extraction**:
-
-For advanced use cases, extract ADR metadata into a searchable format:
-
-```python
-# scripts/extract-adr-metadata.py
-import re
-import json
-from pathlib import Path
-
-adr_dir = Path("docs/adr")
-adrs = []
-
-for adr_file in adr_dir.glob("*.md"):
- if adr_file.name == "index.md":
- continue
-
- content = adr_file.read_text()
-
- adr = {
- "number": re.search(r"ADR-(\d+)", content).group(1) if re.search(r"ADR-(\d+)", content) else None,
- "title": re.search(r"# ADR-\d+: (.+)", content).group(1) if re.search(r"# ADR-\d+: (.+)", content) else None,
- "status": re.search(r"\*\*Status\*\*: (.+)", content).group(1).strip() if re.search(r"\*\*Status\*\*: (.+)", content) else None,
- "date": re.search(r"\*\*Date\*\*: (.+)", content).group(1).strip() if re.search(r"\*\*Date\*\*: (.+)", content) else None,
- "tags": re.findall(r"\*\*Tags\*\*: (.+)", content)[0].split(", ") if re.search(r"\*\*Tags\*\*: (.+)", content) else [],
- "file": str(adr_file.relative_to(Path(".")))
- }
-
- adrs.append(adr)
-
-# Output as JSON for search/indexing
-print(json.dumps(adrs, indent=2))
-```
-
-## Common Pitfalls & Anti-Patterns
-
-### Pitfall 1: Writing ADRs Long After Decision
-
-**Symptom**: ADR is written months or years after the decision was made.
-
-**Why It's a Problem**:
-- Context is lost or misremembered
-- Alternatives considered are forgotten
-- Rationale becomes post-hoc justification
-
-**How to Prevent**:
-- Write ADR **during or immediately after** decision-making
-- Make ADR part of the decision process, not an afterthought
-- If retroactively documenting, acknowledge uncertainty in context section
-
-### Pitfall 2: Giant ADRs Combining Many Decisions
-
-**Symptom**: ADR covers multiple unrelated systems or decisions (e.g., "Database, Message Queue, and API Gateway Strategy v3").
-
-**Why It's a Problem**:
-- Hard to find specific decisions
-- Difficult to supersede individual parts
-- Violates single-responsibility principle
-
-**How to Prevent**:
-- Split ADRs by decision, not by system
-- Use "Related ADRs" section to link related decisions
-- If decisions are tightly coupled, combine only if they form a single coherent pattern
-
-### Pitfall 3: ADRs with No Recorded Alternatives
-
-**Symptom**: ADR states the decision but doesn't document what alternatives were considered.
-
-**Why It's a Problem**:
-- No understanding of trade-offs
-- Future engineers can't evaluate if decision is still valid
-- Missed opportunity to learn from rejected options
-
-**How to Prevent**:
-- Always document at least 2-3 alternatives
-- Even if one option is "obvious," document why others were rejected
-- Use "Options Considered" section to force explicit comparison
-
-### Pitfall 4: Never Marking ADRs as Superseded or Deprecated
-
-**Symptom**: Old ADRs remain "Accepted" even when decisions have changed.
-
-**Why It's a Problem**:
-- Confusion about which decision is current
-- New engineers follow outdated guidance
-- No clear history of decision evolution
-
-**How to Prevent**:
-- Quarterly ADR review meetings
-- When creating new ADR, check if it supersedes an old one
-- Update old ADR status to "Superseded" with link to new ADR
-- Use "Related ADRs" section to maintain decision lineage
-
-### Pitfall 5: ADRs That Aren't Linked to Implementation
-
-**Symptom**: ADR exists but no code, config, or documentation references it.
-
-**Why It's a Problem**:
-- ADR becomes "zombie document" (exists but unused)
-- Decision may not actually be followed
-- No way to verify implementation
-
-**How to Prevent**:
-- Link ADRs from best-practices documents
-- Reference ADRs in design docs and RFCs
-- Include ADR numbers in PR descriptions when implementing decisions
-- Use ADR review to identify unlinked ADRs
-
-### Pitfall 6: Vague or Incomplete Context
-
-**Symptom**: ADR context section is too brief or doesn't explain the problem.
-
-**Why It's a Problem**:
-- Future readers can't understand why decision was made
-- Can't evaluate if decision is still relevant
-- Missing information leads to poor decisions
-
-**How to Prevent**:
-- Context section should be 2-4 paragraphs
-- Explain the problem, constraints, and business/technical drivers
-- Include relevant background (e.g., "We previously used X, but...")
-- Review context section in ADR review process
-
-## Implementation Roadmap
-
-### Phase 1: Seed (Weeks 1-2)
-
-**Goal**: Establish ADR practice with initial examples.
-
-**Tasks**:
-1. **Create ADR directory structure**:
- ```bash
- mkdir -p docs/adr/_templates
- ```
-
-2. **Add templates**:
- - Copy standard ADR template to `docs/adr/_templates/0001-template-standard.md`
- - Copy lightweight ADR template to `docs/adr/_templates/0002-template-lightweight.md`
-
-3. **Write 2-3 retrospective ADRs**:
- - Document already-decided major items:
- - Observability stack (Grafana + Prometheus + Loki)
- - Database auditing (PgAudit + PgCron)
- - SBOM/Trivy policy
- - Use "Accepted" status
- - Include as much context as you remember
-
-4. **Create ADR index**:
- - Manually create `docs/adr/index.md` with table of ADRs
- - Link to ADR best practices document
-
-**Deliverables**:
-- ADR directory with templates
-- 2-3 initial ADRs
-- ADR index page
-
-### Phase 2: Integrate (Weeks 3-4)
-
-**Goal**: Make ADRs part of standard workflow.
-
-**Tasks**:
-1. **Update contribution guidelines**:
- - Add section on "When to Write an ADR"
- - Include link to ADR best practices
- - Add ADR checklist to PR template
-
-2. **Require ADR for new major systems**:
- - When proposing new service/component, require ADR
- - Gate PRs on ADR review for architectural changes
-
-3. **Link ADRs from best-practices docs**:
- - Add ADR references to relevant best-practices documents
- - Update existing docs to reference ADRs where applicable
-
-4. **Train team**:
- - Share ADR best practices document
- - Walk through examples in team meeting
- - Answer questions and gather feedback
-
-**Deliverables**:
-- Updated contribution guidelines
-- ADR references in best-practices docs
-- Team trained on ADR process
-
-### Phase 3: Normalize (Months 2-3)
-
-**Goal**: Make ADRs a natural part of decision-making.
-
-**Tasks**:
-1. **Quarterly ADR reviews**:
- - Schedule first review meeting
- - Review all ADRs for stale status
- - Update superseded/deprecated ADRs
-
-2. **CI checks on ADR structure**:
- - Add pre-commit hook or CI check
- - Validate ADR naming and required fields
- - Fail PRs with invalid ADRs
-
-3. **Cross-links from best-practices**:
- - Ensure all major best-practices docs reference relevant ADRs
- - Add "Related ADRs" sections where missing
-
-4. **Refine process**:
- - Gather feedback from team
- - Update ADR templates if needed
- - Adjust governance model based on experience
-
-**Deliverables**:
-- Quarterly review process established
-- CI validation in place
-- Complete cross-linking
-
-### Phase 4: Mature (Months 4+)
-
-**Goal**: ADRs are fully integrated and self-sustaining.
-
-**Tasks**:
-1. **Auto-generate ADR index**:
- - Create script to generate index from ADR files
- - Add to CI/CD or pre-commit hook
- - Keep index up-to-date automatically
-
-2. **Searchable decision registry**:
- - Extract ADR metadata (optional)
- - Create searchable index or dashboard
- - Enable filtering by status, tags, date
-
-3. **Integration into onboarding**:
- - Add ADR section to onboarding docs
- - Point new engineers to ADR index
- - Explain how to find "why" decisions were made
-
-4. **Continuous improvement**:
- - Regular feedback collection
- - Refine templates and process
- - Share learnings with other teams
-
-**Deliverables**:
-- Auto-generated ADR index
-- Searchable decision registry (optional)
-- ADRs in onboarding process
-
-## Summary
-
-### Key Takeaways
-
-1. **ADRs capture "why"** alongside "how" (best-practices docs)
-2. **One decision per ADR** keeps them focused and maintainable
-3. **Git-based workflow** integrates ADRs into existing PR/review process
-4. **Quarterly reviews** prevent stale ADRs and maintain decision lineage
-5. **Cross-linking** connects ADRs to implementation guides and runbooks
-
-### Adoption Checklist
-
-- [ ] Create `/docs/adr/` directory with templates
-- [ ] Write 2-3 retrospective ADRs for major decisions
-- [ ] Create ADR index page
-- [ ] Update contribution guidelines
-- [ ] Link ADRs from relevant best-practices docs
-- [ ] Train team on ADR process
-- [ ] Schedule first quarterly review
-- [ ] Add CI validation for ADR structure
-- [ ] Integrate ADRs into onboarding
-
-### Next Steps
-
-1. **Start with Phase 1**: Create directory, add templates, write initial ADRs
-2. **Get feedback**: Share with team and iterate on process
-3. **Normalize**: Make ADRs part of standard workflow
-4. **Mature**: Automate and integrate into broader documentation system
-
-**Remember**: ADRs are a tool, not a goal. The goal is better decision-making and knowledge sharing. Start simple, iterate based on feedback, and focus on value over process.
-
-## See Also
-
-- **[Architecture & Design Best Practices](index.md)** - Overview of architecture patterns
-- **[Documentation Best Practices](documentation.md)** - General documentation guidelines
-- **[Git Workflows & Collaboration](../../best-practices/git/git-workflows-collaboration.md)** - Version control patterns
-
----
-
-*This guide provides a complete framework for Architecture Decision Records. Start with Phase 1, gather feedback, and iterate. The goal is better decisions, not perfect documentation.*
-
diff --git a/docs/best-practices/architecture-design/api-gateway-architecture.md b/docs/best-practices/architecture-design/api-gateway-architecture.md
index a3d5507..baa5533 100644
--- a/docs/best-practices/architecture-design/api-gateway-architecture.md
+++ b/docs/best-practices/architecture-design/api-gateway-architecture.md
@@ -1,6 +1,5 @@
# API Gateway Architecture: Best Practices
-**Objective**: Establish comprehensive API gateway architecture patterns for routing, security, rate limiting, and API management. When you need API gateway design, when you want unified API entry points, when you need API management—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/api-governance-interface-stability.md b/docs/best-practices/architecture-design/api-governance-interface-stability.md
index 58c2448..5bbb1be 100644
--- a/docs/best-practices/architecture-design/api-governance-interface-stability.md
+++ b/docs/best-practices/architecture-design/api-governance-interface-stability.md
@@ -1,6 +1,5 @@
# API Governance, Backward Compatibility Rules, and Cross-Language Interface Stability: Best Practices
-**Objective**: Establish comprehensive API governance that ensures backward compatibility, interface stability, and consistent patterns across Python, Go, Rust, and Postgres APIs. When you need API versioning, when you want interface stability, when you need cross-language coherence—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/architecture-fitness-functions-governance.md b/docs/best-practices/architecture-design/architecture-fitness-functions-governance.md
index 33d6d73..f3507d5 100644
--- a/docs/best-practices/architecture-design/architecture-fitness-functions-governance.md
+++ b/docs/best-practices/architecture-design/architecture-fitness-functions-governance.md
@@ -1,6 +1,5 @@
# Architectural Fitness Functions and Governance: Measuring and Evolving Architecture Quality
-**Objective**: Master production-grade architectural fitness functions and governance across distributed systems, databases, ML pipelines, and polyglot microservices. When you need to measure, evaluate, monitor, and evolve architecture quality—this guide provides complete patterns and implementations.
## Introduction
@@ -2036,7 +2035,6 @@ class LLMScorecardGenerator:
## See Also
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Decision recording
- **[Repository Standardization](repository-standardization-and-governance.md)** - Repository governance
- **[System Resilience](../operations-monitoring/system-resilience-and-concurrency.md)** - Resilience patterns
diff --git a/docs/best-practices/architecture-design/cache-topology-architecture.md b/docs/best-practices/architecture-design/cache-topology-architecture.md
index 97c9787..a57f540 100644
--- a/docs/best-practices/architecture-design/cache-topology-architecture.md
+++ b/docs/best-practices/architecture-design/cache-topology-architecture.md
@@ -1,6 +1,5 @@
# Cache-Topology Architecture: Best Practices
-**Objective**: Establish comprehensive multi-tier cache topology patterns that optimize performance, reduce latency, and manage cache hierarchies across edge, application, and data layers. When you need cache topology, when you want multi-tier caching, when you need cache strategy—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/caching-performance.md b/docs/best-practices/architecture-design/caching-performance.md
index 47b8c38..32d48aa 100644
--- a/docs/best-practices/architecture-design/caching-performance.md
+++ b/docs/best-practices/architecture-design/caching-performance.md
@@ -1,6 +1,5 @@
# Caching & Performance Best Practices
-**Objective**: Caches are accelerants. Done right, they make pipelines scream. Done wrong, they give you stale lies.
Caches are accelerants. Done right, they make pipelines scream. Done wrong, they give you stale lies.
@@ -902,4 +901,3 @@ Caching requires understanding performance characteristics, invalidation strateg
---
-*This guide provides the complete machinery for caching and performance. The patterns scale from simple database queries to complex geospatial computations, from basic API responses to advanced ML model inference.*
diff --git a/docs/best-practices/architecture-design/capacity-planning-and-workload-modeling.md b/docs/best-practices/architecture-design/capacity-planning-and-workload-modeling.md
index 550dab7..f00d8ad 100644
--- a/docs/best-practices/architecture-design/capacity-planning-and-workload-modeling.md
+++ b/docs/best-practices/architecture-design/capacity-planning-and-workload-modeling.md
@@ -1,6 +1,5 @@
# Holistic Capacity Planning, Scaling Economics, and Workload Modeling: Best Practices
-**Objective**: Establish comprehensive capacity planning frameworks that model workloads, predict resource needs, and optimize scaling economics across CPU/GPU clusters, data systems, and distributed services. When you need to plan capacity, when you want to model workloads, when you need scaling economics—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/ci-cd-pipelines.md b/docs/best-practices/architecture-design/ci-cd-pipelines.md
index 583b9cc..c665d0a 100644
--- a/docs/best-practices/architecture-design/ci-cd-pipelines.md
+++ b/docs/best-practices/architecture-design/ci-cd-pipelines.md
@@ -1,6 +1,5 @@
# Testing & CI/CD Pipelines Best Practices (2025 Edition)
-**Objective**: A broken pipeline is silent chaos. Here's how to build pipelines that are fast, reproducible, and hard to kill.
A broken pipeline is silent chaos. Here's how to build pipelines that are fast, reproducible, and hard to kill.
@@ -992,4 +991,3 @@ CI/CD pipelines require understanding testing strategies, deployment patterns, a
---
-*This guide provides the complete machinery for CI/CD pipelines. The patterns scale from simple unit tests to complex multi-environment deployments, from basic automation to advanced GitOps workflows.*
diff --git a/docs/best-practices/architecture-design/cognitive-load-developer-experience.md b/docs/best-practices/architecture-design/cognitive-load-developer-experience.md
index cda5b9c..8d8b9fc 100644
--- a/docs/best-practices/architecture-design/cognitive-load-developer-experience.md
+++ b/docs/best-practices/architecture-design/cognitive-load-developer-experience.md
@@ -1,6 +1,5 @@
# Cognitive Load Management and Developer Experience Architecture: Best Practices for Complex Technical Ecosystems
-**Objective**: Master production-grade cognitive load management and developer experience patterns across distributed systems, polyglot microservices, and complex technical ecosystems. When you need to minimize cognitive load, maximize developer experience, and enable effective reasoning about architecture—this guide provides complete patterns and implementations.
## Introduction
@@ -2217,8 +2216,6 @@ class CognitiveRiskReview:
- **[Architectural Fitness Functions](architecture-fitness-functions-governance.md)** - Architecture measurement
- **[Repository Standardization](repository-standardization-and-governance.md)** - Repository governance
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Decision recording
-
---
*This guide provides a complete framework for cognitive load management and developer experience. Start with measuring cognitive load, implement DX patterns, reduce complexity, and continuously improve. The goal is systems that remain mentally tractable and enable effective reasoning about architecture.*
diff --git a/docs/best-practices/architecture-design/cost-aware-architecture-and-efficiency-governance.md b/docs/best-practices/architecture-design/cost-aware-architecture-and-efficiency-governance.md
index a98a67c..2372da1 100644
--- a/docs/best-practices/architecture-design/cost-aware-architecture-and-efficiency-governance.md
+++ b/docs/best-practices/architecture-design/cost-aware-architecture-and-efficiency-governance.md
@@ -1,6 +1,5 @@
# Cost-Aware Architecture & Resource-Efficiency Governance: Best Practices
-**Objective**: Establish comprehensive cost governance frameworks that measure, optimize, and control resource costs across Kubernetes clusters, data systems, ML pipelines, and geospatial workloads. When you need to optimize costs, when you want to rightsize resources, when you need capacity planning—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/data-mesh-architecture.md b/docs/best-practices/architecture-design/data-mesh-architecture.md
index 8a068ec..9160fae 100644
--- a/docs/best-practices/architecture-design/data-mesh-architecture.md
+++ b/docs/best-practices/architecture-design/data-mesh-architecture.md
@@ -1,6 +1,5 @@
# Data Mesh Architecture: Best Practices
-**Objective**: Establish comprehensive data mesh architecture that enables domain-driven data ownership, decentralized data products, and federated governance. When you need domain ownership, when you want decentralized data, when you need federated governance—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/documentation.md b/docs/best-practices/architecture-design/documentation.md
index 2fd3acd..d936fdf 100644
--- a/docs/best-practices/architecture-design/documentation.md
+++ b/docs/best-practices/architecture-design/documentation.md
@@ -1,6 +1,5 @@
# Documentation Best Practices (2025 Edition)
-**Objective**: Docs aren't optional—they're survival kits. Without them, your brilliant pipelines collapse into archaeology.
Docs aren't optional—they're survival kits. Without them, your brilliant pipelines collapse into archaeology.
@@ -1183,4 +1182,3 @@ Documentation requires understanding information architecture, automation patter
---
-*This guide provides the complete machinery for documentation. The patterns scale from simple README files to complex multi-site documentation systems, from basic automation to advanced CI/CD integration.*
diff --git a/docs/best-practices/architecture-design/environment-config-governance.md b/docs/best-practices/architecture-design/environment-config-governance.md
index 714af90..56db1b8 100644
--- a/docs/best-practices/architecture-design/environment-config-governance.md
+++ b/docs/best-practices/architecture-design/environment-config-governance.md
@@ -1,6 +1,5 @@
# Cross-Environment Configuration Strategy, Drift Prevention, and Multi-Cluster State Management: Best Practices
-**Objective**: Master production-grade configuration governance across multiple environments (dev, test, staging, prod, air-gapped, HPC, RKE2, edge). When you need to prevent drift, ensure consistency, and manage configuration lifecycle—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/architecture-design/event-driven-architecture.md b/docs/best-practices/architecture-design/event-driven-architecture.md
index 0cc966e..fe173b1 100644
--- a/docs/best-practices/architecture-design/event-driven-architecture.md
+++ b/docs/best-practices/architecture-design/event-driven-architecture.md
@@ -1,6 +1,5 @@
# Event-Driven Architecture: Design, Deployment, and Hardening Best Practices
-**Objective**: Master production-grade event-driven architecture across Kubernetes, messaging systems, databases, and distributed services. When you need to build resilient, scalable event-driven systems with proper observability and security—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/architecture-design/index.md b/docs/best-practices/architecture-design/index.md
index a925e4b..a12c134 100644
--- a/docs/best-practices/architecture-design/index.md
+++ b/docs/best-practices/architecture-design/index.md
@@ -1,6 +1,5 @@
# Architecture & Design Best Practices
-**Objective**: Master senior-level architecture and design patterns for production systems. When you need to build robust, scalable architectures, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## System Architecture
@@ -42,8 +41,6 @@
## Documentation & Governance
- **[Documentation Best Practices](documentation.md)** - Writing, organizing, and sustaining useful docs for data + devops projects
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Architecture Decision Records (ADRs) and decision governance framework
-
## Data Serialization
- **[Protocol Buffers with Python](protobuf-python.md)** - Production-ready data serialization and microservices communication
@@ -55,4 +52,3 @@
---
-*These best practices provide the complete machinery for building production-ready architectures. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/architecture-design/multi-cloud-federation-portability.md b/docs/best-practices/architecture-design/multi-cloud-federation-portability.md
index d4414bb..3b52b5f 100644
--- a/docs/best-practices/architecture-design/multi-cloud-federation-portability.md
+++ b/docs/best-practices/architecture-design/multi-cloud-federation-portability.md
@@ -1,6 +1,5 @@
# Multi-Cloud Federation & Portability Architecture: Best Practices
-**Objective**: Establish comprehensive multi-cloud federation and portability patterns that enable workload portability, vendor independence, and unified operations across AWS, GCP, Azure, and on-premises infrastructure. When you need cloud portability, when you want vendor independence, when you need unified multi-cloud operations—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/multi-region-dr-strategy.md b/docs/best-practices/architecture-design/multi-region-dr-strategy.md
index d408b09..ac7a5a7 100644
--- a/docs/best-practices/architecture-design/multi-region-dr-strategy.md
+++ b/docs/best-practices/architecture-design/multi-region-dr-strategy.md
@@ -1,6 +1,5 @@
# Multi-Region, Multi-Cluster Disaster Recovery, Failover Topologies, and Data Sovereignty: Best Practices
-**Objective**: Establish comprehensive disaster recovery strategies across multiple regions, clusters, and cloud providers. When you need cross-region failover, when you want data sovereignty compliance, when you need air-gapped DR—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/polyglot-interoperability-design.md b/docs/best-practices/architecture-design/polyglot-interoperability-design.md
index 45a72af..3e8e0e8 100644
--- a/docs/best-practices/architecture-design/polyglot-interoperability-design.md
+++ b/docs/best-practices/architecture-design/polyglot-interoperability-design.md
@@ -1,6 +1,5 @@
# Polyglot Interoperability Design: Best Practices
-**Objective**: Establish comprehensive polyglot interoperability patterns that enable seamless integration between Python, Go, Rust, Postgres, and other systems. When you need cross-language integration, when you want polyglot systems, when you need interoperability—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/protobuf-python.md b/docs/best-practices/architecture-design/protobuf-python.md
index 5754a62..6744811 100644
--- a/docs/best-practices/architecture-design/protobuf-python.md
+++ b/docs/best-practices/architecture-design/protobuf-python.md
@@ -1,6 +1,5 @@
# Protocol Buffers with Python: Production-Ready Data Serialization
-**Objective**: Master Protocol Buffers (protobufs) for efficient, type-safe data serialization in Python applications. When you need high-performance data exchange, when you want type-safe APIs, when you're building microservices that communicate efficiently—protobufs become your weapon of choice.
Protocol Buffers provide the foundation for efficient data serialization and type-safe communication. Without proper understanding of schema design, code generation, and serialization patterns, you're building inefficient systems that miss the power of binary serialization and schema evolution. This guide shows you how to wield protobufs with the precision of a senior Python engineer.
@@ -1374,4 +1373,3 @@ Protocol Buffers provide the foundation for efficient, type-safe data serializat
---
-*This guide provides the complete machinery for mastering Protocol Buffers with Python. The patterns scale from simple message serialization to complex microservices communication, from basic schemas to advanced production deployment.*
diff --git a/docs/best-practices/architecture-design/rdf-owl-metadata-automation.md b/docs/best-practices/architecture-design/rdf-owl-metadata-automation.md
index abee9d8..3d0ccd0 100644
--- a/docs/best-practices/architecture-design/rdf-owl-metadata-automation.md
+++ b/docs/best-practices/architecture-design/rdf-owl-metadata-automation.md
@@ -1,6 +1,5 @@
# RDF/OWL Metadata Automation: Dynamic Knowledge Graphs with Friend-of-a-Friend Associations
-**Objective**: Master RDF/OWL for automated metadata handling and dynamic ontological associations. When you need to automatically discover relationships between entities, when you're building intelligent knowledge graphs, when you want to leverage semantic reasoning for metadata enrichment—RDF/OWL becomes your weapon of choice.
RDF/OWL metadata automation is the foundation of intelligent knowledge management. Without proper understanding of semantic web technologies, you're building static, disconnected metadata systems that miss the power of automated reasoning and relationship discovery. This guide shows you how to wield RDF/OWL with the precision of a semantic web engineer.
@@ -888,4 +887,3 @@ RDF/OWL metadata automation provides the foundation for intelligent knowledge ma
---
-*This guide provides the complete machinery for mastering RDF/OWL metadata automation. The patterns scale from simple ontologies to complex knowledge graphs, from basic reasoning to advanced inference.*
diff --git a/docs/best-practices/architecture-design/reference-architecture-diagrams.md b/docs/best-practices/architecture-design/reference-architecture-diagrams.md
index 56a32fd..46b3b16 100644
--- a/docs/best-practices/architecture-design/reference-architecture-diagrams.md
+++ b/docs/best-practices/architecture-design/reference-architecture-diagrams.md
@@ -1,6 +1,5 @@
# Reference Architecture Diagrams: Best Practices
-**Objective**: Establish comprehensive reference architecture diagrams that document system topologies, component relationships, and architectural patterns. When you need reference architectures, when you want system documentation, when you need architectural blueprints—this guide provides the complete framework.
## Introduction
@@ -23,7 +22,6 @@ Reference architecture diagrams are essential for understanding, communicating,
**Related Documents**:
This document integrates with:
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Decision documentation
- **[Documentation](documentation.md)** - Documentation patterns
- **[System-Wide Naming, Taxonomy, and Structural Vocabulary Governance](system-taxonomy-governance.md)** - Naming standards
@@ -98,7 +96,6 @@ diagram_standards:
## See Also
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Decisions
- **[Documentation](documentation.md)** - Documentation
- **[System-Wide Naming, Taxonomy, and Structural Vocabulary Governance](system-taxonomy-governance.md)** - Naming
diff --git a/docs/best-practices/architecture-design/repository-standardization-and-governance.md b/docs/best-practices/architecture-design/repository-standardization-and-governance.md
index e031e9d..efac884 100644
--- a/docs/best-practices/architecture-design/repository-standardization-and-governance.md
+++ b/docs/best-practices/architecture-design/repository-standardization-and-governance.md
@@ -1,6 +1,5 @@
# Repository Standardization, Templates, and Lifecycle Governance: Best Practices for Polyglot Ecosystems
-**Objective**: Master production-grade repository standardization across Python, Go, Rust, Docker, Kubernetes, and data pipelines. When you need to eliminate snowflake repos, ensure consistency, and enable automated governance—this guide provides complete patterns and implementations.
## Introduction
@@ -2171,7 +2170,6 @@ class LLMDocsIngester:
## See Also
-- **[ADR and Technical Decision Governance](adr-decision-governance.md)** - Decision recording
- **[Configuration Management](../operations-monitoring/configuration-management.md)** - Config governance
- **[Release Management](../operations-monitoring/release-management-and-progressive-delivery.md)** - Deployment practices
diff --git a/docs/best-practices/architecture-design/secrets-management.md b/docs/best-practices/architecture-design/secrets-management.md
index b7cff6d..108d299 100644
--- a/docs/best-practices/architecture-design/secrets-management.md
+++ b/docs/best-practices/architecture-design/secrets-management.md
@@ -1,6 +1,5 @@
# Secrets Management Best Practices (2025 Edition)
-**Objective**: Secrets are dangerous. They leak, they rot, they get copied into Slack and Git commits. Here's how to lock them down without breaking your developer flow.
Secrets are dangerous. They leak, they rot, they get copied into Slack and Git commits. Here's how to lock them down without breaking your developer flow.
@@ -799,4 +798,3 @@ Secrets management requires understanding security risks, access patterns, and l
---
-*This guide provides the complete machinery for secrets management. The patterns scale from simple environment variables to complex enterprise secret stores, from basic security to advanced threat protection.*
diff --git a/docs/best-practices/architecture-design/service-decomposition-strategy.md b/docs/best-practices/architecture-design/service-decomposition-strategy.md
index 70cea8a..81eb99f 100644
--- a/docs/best-practices/architecture-design/service-decomposition-strategy.md
+++ b/docs/best-practices/architecture-design/service-decomposition-strategy.md
@@ -1,6 +1,5 @@
# Service Decomposition Strategy: Best Practices
-**Objective**: Establish comprehensive service decomposition strategies that guide when and how to decompose monolithic systems into microservices, bounded contexts, and domain services. When you need decomposition guidance, when you want domain boundaries, when you need service boundaries—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/streaming-architecture-patterns.md b/docs/best-practices/architecture-design/streaming-architecture-patterns.md
index c591e33..b66e38f 100644
--- a/docs/best-practices/architecture-design/streaming-architecture-patterns.md
+++ b/docs/best-practices/architecture-design/streaming-architecture-patterns.md
@@ -1,6 +1,5 @@
# Streaming Architecture Patterns: SAGA, CQRS, and Outbox: Best Practices
-**Objective**: Establish comprehensive streaming architecture patterns including SAGA for distributed transactions, CQRS for read/write separation, and Outbox for reliable event publishing. When you need distributed transactions, when you want read/write separation, when you need reliable events—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/system-taxonomy-governance.md b/docs/best-practices/architecture-design/system-taxonomy-governance.md
index 015e99d..a5a60aa 100644
--- a/docs/best-practices/architecture-design/system-taxonomy-governance.md
+++ b/docs/best-practices/architecture-design/system-taxonomy-governance.md
@@ -1,6 +1,5 @@
# System-Wide Naming, Taxonomy, and Structural Vocabulary Governance: Best Practices
-**Objective**: Establish enterprise-wide naming conventions, domain taxonomy, and structural vocabulary that serve as the "lingua franca" across all systems, services, databases, and codebases. When you need consistent naming, when you want to reduce cognitive load, when you need cross-system clarity—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/architecture-design/temporal-governance-and-time-synchronization.md b/docs/best-practices/architecture-design/temporal-governance-and-time-synchronization.md
index ccca81e..4ca9e3c 100644
--- a/docs/best-practices/architecture-design/temporal-governance-and-time-synchronization.md
+++ b/docs/best-practices/architecture-design/temporal-governance-and-time-synchronization.md
@@ -1,6 +1,5 @@
# Temporal Governance and Time Synchronization: Best Practices for Distributed Systems
-**Objective**: Master production-grade time synchronization across Kubernetes, databases, ML pipelines, and distributed systems. When you need to ensure causality, prevent clock drift, and maintain temporal consistency—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/architecture/cost-aware-systems.md b/docs/best-practices/architecture/cost-aware-systems.md
index f64105b..c154827 100644
--- a/docs/best-practices/architecture/cost-aware-systems.md
+++ b/docs/best-practices/architecture/cost-aware-systems.md
@@ -1,6 +1,5 @@
# Cost-Aware System Architecture
-**Objective**: Design systems that account for operational cost as an architectural constraint, not an afterthought.
## Cost as an architectural constraint
diff --git a/docs/best-practices/creative-fun/celery-best-practices.md b/docs/best-practices/creative-fun/celery-best-practices.md
index ad27dff..ce6259b 100644
--- a/docs/best-practices/creative-fun/celery-best-practices.md
+++ b/docs/best-practices/creative-fun/celery-best-practices.md
@@ -1,6 +1,5 @@
# Celery in Production: Picking the Right Jobs, Writing Safe Tasks, Running It Like You Mean It
-**Objective**: Master Celery task queues for reliable distributed processing in Python applications. When you need to offload HTTP requests, when you want to process background jobs, when you're building scalable data pipelines—Celery becomes your weapon of choice.
Celery is a distributed task queue that provides reliable at-least-once execution with retries, scheduling, and monitoring. It's not a streaming processor or OLAP engine—it's designed for offloading work from web requests, processing scheduled jobs, and handling idempotent side effects at scale.
@@ -989,4 +988,3 @@ Celery requires understanding both distributed systems patterns and Python concu
---
-*This guide provides the complete machinery for Celery in production. The patterns scale from simple HTTP offloading to complex distributed workflows, from basic task execution to advanced orchestration.*
diff --git a/docs/best-practices/creative-fun/idempotency-and-dedup.md b/docs/best-practices/creative-fun/idempotency-and-dedup.md
index f3dae0e..0d8ba30 100644
--- a/docs/best-practices/creative-fun/idempotency-and-dedup.md
+++ b/docs/best-practices/creative-fun/idempotency-and-dedup.md
@@ -1,6 +1,5 @@
# Idempotency & De-dup: Designing Operations That Don't Double-Fire
-**Objective**: Master idempotency and de-duplication patterns to prevent duplicate operations in distributed systems. When you need safe retries, when you want to prevent double-billing, when you're building reliable APIs and data pipelines—idempotency becomes your weapon of choice.
Retries happen (clients, proxies, workers). Without idempotency, you print money twice—or delete the same row twice. Build request identity, dedupe state, and committed outcomes into every layer.
@@ -584,4 +583,3 @@ Idempotency requires understanding both system reliability and data consistency
---
-*This guide provides the complete machinery for idempotency and de-duplication. The patterns scale from simple HTTP APIs to complex distributed systems, from basic retry handling to advanced stream processing.*
diff --git a/docs/best-practices/creative-fun/index.md b/docs/best-practices/creative-fun/index.md
index 628a3f4..1098973 100644
--- a/docs/best-practices/creative-fun/index.md
+++ b/docs/best-practices/creative-fun/index.md
@@ -1,6 +1,5 @@
# Creative & Fun Best Practices
-**Objective**: Master creative and fun patterns for production systems. When you need to build engaging, creative solutions, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## Opinions
@@ -15,4 +14,3 @@
---
-*These best practices provide the complete machinery for building production-ready creative systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/creative-fun/latex.md b/docs/best-practices/creative-fun/latex.md
index a699122..f4db6b9 100644
--- a/docs/best-practices/creative-fun/latex.md
+++ b/docs/best-practices/creative-fun/latex.md
@@ -1,6 +1,5 @@
# LaTeX Workflows Best Practices
-**Objective**: Master production-grade LaTeX workflows for reproducible technical writing, academic papers, and engineering documentation. When you need to write professional technical documents, when you want to create beautiful diagrams and figures, when you're responsible for maintaining consistent formatting across teams—LaTeX workflows become your weapon of choice.
LaTeX isn't just typesetting—it's reproducible writing. Done poorly, it's chaos. Done well, it's a publishing pipeline as solid as any codebase.
@@ -631,4 +630,3 @@ LaTeX workflows require understanding document structure, build reproducibility,
---
-*This guide provides the complete machinery for LaTeX workflows. The patterns scale from simple documents to complex academic papers, from basic formatting to advanced document generation.*
diff --git a/docs/best-practices/creative-fun/time-hygiene.md b/docs/best-practices/creative-fun/time-hygiene.md
index b9455c6..9645d9d 100644
--- a/docs/best-practices/creative-fun/time-hygiene.md
+++ b/docs/best-practices/creative-fun/time-hygiene.md
@@ -1,6 +1,5 @@
# Time Hygiene: UTC Everywhere, Monotonic Durations, and DST-Proof Scheduling
-**Objective**: Master temporal hygiene for production systems that handle time correctly. When you need reliable scheduling, when you want to prevent DST disasters, when you're building data pipelines that span timezones—time hygiene becomes your weapon of choice.
Time is the most misunderstood aspect of system design. Proper time handling prevents DST disasters, enables reliable scheduling, and maintains data integrity across timezones. This guide shows you how to wield time with the precision of a paranoid SRE, covering everything from UTC storage to monotonic clocks and DST-proof scheduling.
@@ -806,4 +805,3 @@ Time handling requires understanding both temporal mechanics and system design p
---
-*This guide provides the complete machinery for time hygiene. The patterns scale from simple timestamp storage to complex distributed systems, from basic scheduling to advanced temporal data processing.*
diff --git a/docs/best-practices/creative-fun/yaml-recipe-format.md b/docs/best-practices/creative-fun/yaml-recipe-format.md
index aa02585..9d80414 100644
--- a/docs/best-practices/creative-fun/yaml-recipe-format.md
+++ b/docs/best-practices/creative-fun/yaml-recipe-format.md
@@ -1,6 +1,5 @@
# YAML Recipe Format: Why Your Kitchen Needs Indentation
-**Objective**: Master YAML as the superior format for culinary operations. When you need to version control your grandmother's secret sauce, when you want to lint your pancake recipes, when you're building a CI/CD pipeline for your chili—YAML recipe format becomes your weapon of choice.
Other people use index cards. We use declarative configs.
@@ -1014,4 +1013,3 @@ YAML recipe format provides the foundation for modern culinary operations. When
---
-*This guide provides the complete machinery for mastering YAML recipe format. The patterns scale from simple breakfast recipes to complex multi-course meals, from basic cooking to advanced culinary operations.*
diff --git a/docs/best-practices/data-governance/data-freshness-sla-governance.md b/docs/best-practices/data-governance/data-freshness-sla-governance.md
index ec459e0..1a90854 100644
--- a/docs/best-practices/data-governance/data-freshness-sla-governance.md
+++ b/docs/best-practices/data-governance/data-freshness-sla-governance.md
@@ -1,6 +1,5 @@
# Data Freshness, SLA/SLO Governance, and Pipeline Reliability Contracts: Best Practices
-**Objective**: Establish comprehensive data freshness governance with SLA/SLO frameworks for ETL pipelines, real-time streaming, geospatial processing, and data serving layers. When you need to ensure data freshness, when you want to define reliability contracts, when you need pipeline SLOs—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/data-governance/data-lineage-contracts.md b/docs/best-practices/data-governance/data-lineage-contracts.md
index 5a88c48..5ca715a 100644
--- a/docs/best-practices/data-governance/data-lineage-contracts.md
+++ b/docs/best-practices/data-governance/data-lineage-contracts.md
@@ -1,6 +1,5 @@
# Cross-System Data Lineage, Inter-Service Metadata Contracts & Provenance Enforcement
-**Objective**: Establish data lineage and contract enforcement across pipelines, services, and storage so that provenance is traceable, schema and semantics are agreed, and violations are detectable. When you need to trace data from source to consumption, enforce contracts between producers and consumers, or prove compliance—this guide provides the patterns and integration points.
## Introduction
diff --git a/docs/best-practices/data-governance/data-quality-sla-validation-observability.md b/docs/best-practices/data-governance/data-quality-sla-validation-observability.md
index dc30b72..bdf552f 100644
--- a/docs/best-practices/data-governance/data-quality-sla-validation-observability.md
+++ b/docs/best-practices/data-governance/data-quality-sla-validation-observability.md
@@ -1,6 +1,5 @@
# Data Quality SLAs, Validation Layers, and Observability for Tabular, Geospatial, and ML Data: Best Practices
-**Objective**: Establish comprehensive data quality governance with SLAs, multi-layer validation, and observability for tabular, geospatial, and ML data. When you need data quality assurance, when you want quality SLAs, when you need validation observability—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/data-governance/data-retention-archival-lifecycle-governance.md b/docs/best-practices/data-governance/data-retention-archival-lifecycle-governance.md
index 76039f7..0754ff2 100644
--- a/docs/best-practices/data-governance/data-retention-archival-lifecycle-governance.md
+++ b/docs/best-practices/data-governance/data-retention-archival-lifecycle-governance.md
@@ -1,6 +1,5 @@
# Data Retention, Archival Strategy, Lifecycle Governance & Cold Storage Patterns: Best Practices
-**Objective**: Establish comprehensive data retention and archival strategies that govern data lifecycle from hot to frozen storage, ensuring compliance, cost optimization, and operational efficiency. When you need retention policies, when you want archival strategies, when you need lifecycle governance—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/data-governance/data-validation-and-contract-governance.md b/docs/best-practices/data-governance/data-validation-and-contract-governance.md
index de30bee..ccebe35 100644
--- a/docs/best-practices/data-governance/data-validation-and-contract-governance.md
+++ b/docs/best-practices/data-governance/data-validation-and-contract-governance.md
@@ -1,6 +1,5 @@
# Data Validation and Contract Governance: Best Practices for Polyglot Distributed Systems
-**Objective**: Master production-grade data validation and contract governance across Postgres, DuckDB, MLflow, Parquet, ETL pipelines, and distributed systems. When you need to ensure data quality, prevent silent corruption, and maintain contract consistency—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/data-governance/index.md b/docs/best-practices/data-governance/index.md
index ae18139..0246cd5 100644
--- a/docs/best-practices/data-governance/index.md
+++ b/docs/best-practices/data-governance/index.md
@@ -1,6 +1,5 @@
# Data Governance Best Practices
-**Objective**: Master production-grade data governance for distributed analytics systems. When you need to ensure data quality, track lineage, enforce contracts, and maintain reproducibility across Postgres, Parquet, MLflow, and ETL pipelines—these best practices become your foundation.
This collection provides comprehensive guides for metadata management, schema governance, data provenance, and data contracts. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies.
diff --git a/docs/best-practices/data-governance/metadata-provenance-contracts.md b/docs/best-practices/data-governance/metadata-provenance-contracts.md
index 0037de2..d286ad3 100644
--- a/docs/best-practices/data-governance/metadata-provenance-contracts.md
+++ b/docs/best-practices/data-governance/metadata-provenance-contracts.md
@@ -1,6 +1,5 @@
# Metadata Standards, Schema Governance & Data Provenance Contracts: Best Practices for Distributed Analytics Systems
-**Objective**: Master production-grade metadata management, schema governance, and data provenance for distributed analytics ecosystems. When you need to ensure data quality, track lineage, enforce contracts, and maintain reproducibility across Postgres, Parquet, MLflow, and ETL pipelines—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/data-processing/spark/scaling-spark.md b/docs/best-practices/data-processing/spark/scaling-spark.md
index 857f481..cbeea49 100644
--- a/docs/best-practices/data-processing/spark/scaling-spark.md
+++ b/docs/best-practices/data-processing/spark/scaling-spark.md
@@ -1,6 +1,5 @@
# Scaling Spark Clusters Correctly
-**Objective**: Cluster sizing, executor strategy, and partitioning so Spark jobs run efficiently without waste or bottlenecks.
## Spark architecture overview
diff --git a/docs/best-practices/data-processing/spark/spark-modern-architecture.md b/docs/best-practices/data-processing/spark/spark-modern-architecture.md
index 889ef47..4e2a941 100644
--- a/docs/best-practices/data-processing/spark/spark-modern-architecture.md
+++ b/docs/best-practices/data-processing/spark/spark-modern-architecture.md
@@ -1,6 +1,5 @@
# Spark in Modern Data Architectures
-**Objective**: Place Spark in context among lakehouse, object storage, and modern analytics tools so you can choose and integrate it correctly.
## Historical role of Spark
diff --git a/docs/best-practices/data-processing/spark/spark-on-kubernetes.md b/docs/best-practices/data-processing/spark/spark-on-kubernetes.md
index a4c8179..03707a9 100644
--- a/docs/best-practices/data-processing/spark/spark-on-kubernetes.md
+++ b/docs/best-practices/data-processing/spark/spark-on-kubernetes.md
@@ -1,6 +1,5 @@
# Running Spark on Kubernetes
-**Objective**: Operational best practices for running Apache Spark on Kubernetes: architecture, deployment modes, resources, and storage.
## Why Kubernetes for Spark
diff --git a/docs/best-practices/data-processing/spark/spark-performance-tuning.md b/docs/best-practices/data-processing/spark/spark-performance-tuning.md
index c2b18cf..81f3dba 100644
--- a/docs/best-practices/data-processing/spark/spark-performance-tuning.md
+++ b/docs/best-practices/data-processing/spark/spark-performance-tuning.md
@@ -1,6 +1,5 @@
# Spark Performance Tuning
-**Objective**: Tune Spark jobs and data layout for throughput and stability: data layout, execution plans, memory, and join strategies.
## Data layout matters
diff --git a/docs/best-practices/data-processing/spark/when-to-use-spark.md b/docs/best-practices/data-processing/spark/when-to-use-spark.md
index cb4b042..d476f1c 100644
--- a/docs/best-practices/data-processing/spark/when-to-use-spark.md
+++ b/docs/best-practices/data-processing/spark/when-to-use-spark.md
@@ -1,6 +1,5 @@
# When to Use Apache Spark (and When Not To)
-**Objective**: Help engineers decide when Spark is the right tool and when alternatives are a better fit.
## What Spark is designed for
diff --git a/docs/best-practices/data/metadata-control-plane.md b/docs/best-practices/data/metadata-control-plane.md
index 6dd4d0e..e944324 100644
--- a/docs/best-practices/data/metadata-control-plane.md
+++ b/docs/best-practices/data/metadata-control-plane.md
@@ -1,6 +1,5 @@
# Metadata as a Control Plane
-**Objective**: Explain how metadata governs data systems and how to treat it as infrastructure rather than documentation.
## Metadata vs data
diff --git a/docs/best-practices/data/reproducible-data-pipelines.md b/docs/best-practices/data/reproducible-data-pipelines.md
index c008efc..0d8f0f7 100644
--- a/docs/best-practices/data/reproducible-data-pipelines.md
+++ b/docs/best-practices/data/reproducible-data-pipelines.md
@@ -1,6 +1,5 @@
# Reproducible Data Pipelines
-**Objective**: Define practices for building deterministic, auditable data pipelines that produce the same outputs from the same inputs and support governance and debugging.
## Why reproducibility matters
@@ -119,6 +118,6 @@ Avoid pipelines that depend on undocumented inputs, unversioned code, or overwri
- [ETL Pipeline Design](../database-data/etl-pipeline-design.md) — production ETL patterns and orchestration
- [Parquet](../database-data/parquet.md) and [GeoParquet](../database-data/geoparquet.md) — format and layout best practices
- [Data Lineage, Contracts & Provenance](../database-data/data-lineage-contracts.md) — lineage and provenance enforcement
-- [Data Pipeline Quality, Governance & Observability](../../tutorials/best-practices-integration/data-pipeline-quality-governance-observability.md) — quality and observability in pipelines
+- [Metadata as control plane](metadata-control-plane.md) — lineage and operational metadata
- [Why Most Data Pipelines Fail](../../deep-dives/why-most-data-pipelines-fail.md) — failure modes and structural fixes
- [Lakehouse vs Warehouse vs Database](../../deep-dives/lakehouse-vs-warehouse-vs-database.md) — where pipelines land data
diff --git a/docs/best-practices/database-data/ai-ml-geospatial-knowledge-graph.md b/docs/best-practices/database-data/ai-ml-geospatial-knowledge-graph.md
index 3259eb6..d7ed10b 100644
--- a/docs/best-practices/database-data/ai-ml-geospatial-knowledge-graph.md
+++ b/docs/best-practices/database-data/ai-ml-geospatial-knowledge-graph.md
@@ -1,6 +1,5 @@
# Best Practices for Designing an AI-Ready, ML-Enabled Geospatial Knowledge Graph
-**Objective**: Establish comprehensive best practices for designing, building, operating, and evolving an AI-ready, ML-enabled Geospatial Knowledge Graph (GeoKG) suitable for production environments, geospatial analytics, and advanced ML/RAG workflows. When you need to combine geospatial data with knowledge graphs, semantic layers, and ML systems—this guide provides the complete framework.
## Abstract
diff --git a/docs/best-practices/database-data/aws-serverless-geospatial.md b/docs/best-practices/database-data/aws-serverless-geospatial.md
index fcddf0e..633f50c 100644
--- a/docs/best-practices/database-data/aws-serverless-geospatial.md
+++ b/docs/best-practices/database-data/aws-serverless-geospatial.md
@@ -1,6 +1,5 @@
# AWS Serverless Geospatial Processing
-**Objective**: Build scalable, cost-effective geospatial processing pipelines using AWS serverless services.
Serverless architecture eliminates infrastructure management while providing automatic scaling and pay-per-use pricing. This guide covers building production-ready geospatial processing systems on AWS.
diff --git a/docs/best-practices/database-data/data-lake-governance.md b/docs/best-practices/database-data/data-lake-governance.md
index cdb95ad..082bf43 100644
--- a/docs/best-practices/database-data/data-lake-governance.md
+++ b/docs/best-practices/database-data/data-lake-governance.md
@@ -1,6 +1,5 @@
# Best Practices for Data Lake Governance
-**Objective**: Master data lake governance for trustworthy, auditable, compliant data at scale. When you need to enforce data quality, when you're building compliance workflows, when you need to scale to billions of objects—data lake governance becomes your weapon of choice.
Data lake governance is the foundation of trustworthy data infrastructure. Without proper governance, data lakes become data swamps with shadow data, schema drift, and compliance violations. This guide shows you how to design and enforce governance with the precision of a data architect who has been burned before.
@@ -1003,4 +1002,3 @@ Data lake governance provides the foundation for trustworthy, auditable, complia
---
-*This guide provides the complete machinery for mastering data lake governance. The patterns scale from development to production, from simple schemas to enterprise-grade compliance frameworks.*
diff --git a/docs/best-practices/database-data/database-migrations.md b/docs/best-practices/database-data/database-migrations.md
index c46fb4a..28f0933 100644
--- a/docs/best-practices/database-data/database-migrations.md
+++ b/docs/best-practices/database-data/database-migrations.md
@@ -1,6 +1,5 @@
# Database Migrations & Schema Evolution
-**Objective**: Databases are living systems. Schema changes are inevitable. Handle them without breaking prod, corrupting data, or waking ops in the night.
Databases are living systems. Schema changes are inevitable. Handle them without breaking prod, corrupting data, or waking ops in the night.
@@ -1002,4 +1001,3 @@ Database migrations require understanding schema evolution, data safety, and pro
---
-*This guide provides the complete machinery for database migrations. The patterns scale from simple table changes to complex geospatial schema evolution, from basic Alembic usage to advanced zero-downtime strategies.*
diff --git a/docs/best-practices/database-data/etl-pipeline-design.md b/docs/best-practices/database-data/etl-pipeline-design.md
index 37361da..ae21c62 100644
--- a/docs/best-practices/database-data/etl-pipeline-design.md
+++ b/docs/best-practices/database-data/etl-pipeline-design.md
@@ -1,6 +1,5 @@
# ETL Pipeline Design
-**Objective**: Build robust, scalable ETL pipelines for geospatial data processing using modern orchestration tools.
ETL pipelines are the backbone of data engineering. This guide covers designing production-ready ETL systems that handle geospatial data at scale with reliability and performance.
diff --git a/docs/best-practices/database-data/geoparquet-data-warehouses.md b/docs/best-practices/database-data/geoparquet-data-warehouses.md
index 8777199..fc0b935 100644
--- a/docs/best-practices/database-data/geoparquet-data-warehouses.md
+++ b/docs/best-practices/database-data/geoparquet-data-warehouses.md
@@ -1,6 +1,5 @@
# GeoParquet Data Warehouses
-**Objective**: Build modern geospatial data warehouses using GeoParquet format for optimal performance and interoperability.
GeoParquet represents the future of geospatial data storage—columnar, compressed, and cross-platform compatible. This guide covers building production-ready geospatial data warehouses that scale.
diff --git a/docs/best-practices/database-data/geospatial-benchmarking.md b/docs/best-practices/database-data/geospatial-benchmarking.md
index 5fb8713..0b9b1df 100644
--- a/docs/best-practices/database-data/geospatial-benchmarking.md
+++ b/docs/best-practices/database-data/geospatial-benchmarking.md
@@ -1,6 +1,5 @@
# Best Practices for Geospatial Benchmarking under CPU/GPU Stress
-**Objective**: Master geospatial benchmarking methodologies for high-performance computing environments. When you need to stress-test spatial algorithms, when you're optimizing for GPU acceleration, when you need to validate performance under load—geospatial benchmarking becomes your weapon of choice.
Geospatial benchmarking is the foundation of performance validation for spatial computing. Without proper benchmarking, you're flying blind into production with algorithms that may fail under load, scale poorly, or consume excessive resources. This guide shows you how to design and execute comprehensive geospatial benchmarks with the precision of an HPC engineer.
@@ -1052,4 +1051,3 @@ Geospatial benchmarking under CPU/GPU stress provides the foundation for perform
---
-*This guide provides the complete machinery for mastering geospatial benchmarking under CPU/GPU stress. The patterns scale from development to production, from simple operations to enterprise-grade spatial computing systems.*
diff --git a/docs/best-practices/database-data/index.md b/docs/best-practices/database-data/index.md
index 8174dab..f986171 100644
--- a/docs/best-practices/database-data/index.md
+++ b/docs/best-practices/database-data/index.md
@@ -1,6 +1,5 @@
# Database & Data Management Best Practices
-**Objective**: Master senior-level database and data management patterns for production systems. When you need to build robust, scalable data systems, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## PostgreSQL & High Availability
@@ -36,4 +35,3 @@
---
-*These best practices provide the complete machinery for building production-ready data systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/database-data/lake-vs-lakehouse-vs-warehouse.md b/docs/best-practices/database-data/lake-vs-lakehouse-vs-warehouse.md
index 32ed32d..a70a2ec 100644
--- a/docs/best-practices/database-data/lake-vs-lakehouse-vs-warehouse.md
+++ b/docs/best-practices/database-data/lake-vs-lakehouse-vs-warehouse.md
@@ -1,6 +1,5 @@
# Choosing Data Lakes, Lakehouses, and Warehouses: Trade-offs, Patterns, and Free Stacks
-**Objective**: Master the architectural trade-offs between data lakes, lakehouses, and warehouses for modern data infrastructure. When you need to choose the right data architecture, when you're building scalable analytics platforms, when you need to balance cost, performance, and governance—data architecture becomes your weapon of choice.
Data architecture decisions determine the fate of your analytics platform. Without proper understanding of the trade-offs, you'll build expensive, slow, or ungovernable systems. This guide shows you how to choose and implement the right architecture with the precision of a principal data architect.
@@ -951,4 +950,3 @@ Data architecture decisions determine the fate of your analytics platform. When
---
-*This guide provides the complete machinery for choosing and implementing data architectures. The patterns scale from development to production, from simple analytics to enterprise-grade data platforms.*
diff --git a/docs/best-practices/database-data/patroni-postgres-ha.md b/docs/best-practices/database-data/patroni-postgres-ha.md
index 3c1e0a9..4c79322 100644
--- a/docs/best-practices/database-data/patroni-postgres-ha.md
+++ b/docs/best-practices/database-data/patroni-postgres-ha.md
@@ -1,6 +1,5 @@
# Patroni PostgreSQL HA: The Art of High Availability
-**Objective**: Master Patroni to build bulletproof PostgreSQL clusters that survive hardware failures, network partitions, and catastrophic disasters. When your database becomes the single point of failure, when downtime costs thousands per minute, when data loss is not an option—Patroni becomes your weapon of choice.
PostgreSQL high availability is the bridge between single-node databases and enterprise-grade resilience. Without proper HA setup, you're flying blind into production with databases that could fail in ways that destroy your business. This guide shows you how to wield Patroni with the precision of a seasoned database engineer.
@@ -910,4 +909,3 @@ Patroni PostgreSQL HA is the foundation of reliable database deployments. When c
---
-*This tutorial provides the complete machinery for mastering Patroni PostgreSQL HA. The patterns scale from development to production, from simple clusters to enterprise-grade deployments.*
diff --git a/docs/best-practices/database-data/semantic-layer-engineering.md b/docs/best-practices/database-data/semantic-layer-engineering.md
index 4e5cb31..5b3481a 100644
--- a/docs/best-practices/database-data/semantic-layer-engineering.md
+++ b/docs/best-practices/database-data/semantic-layer-engineering.md
@@ -1,6 +1,5 @@
# Semantic Layer Engineering, Domain Models, and Knowledge Graph Alignment: Best Practices
-**Objective**: Establish enterprise semantic layers that bridge business concepts to physical storage across geospatial, infrastructure, ML/AI, and data domains. When you need domain models, when you want knowledge graph alignment, when you need semantic versioning—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/diagrams/svg-workflow-generation.md b/docs/best-practices/diagrams/svg-workflow-generation.md
index 09ea658..504e518 100644
--- a/docs/best-practices/diagrams/svg-workflow-generation.md
+++ b/docs/best-practices/diagrams/svg-workflow-generation.md
@@ -231,7 +231,7 @@ A future GitHub Actions workflow might look like:
## Example: Data Platform Workflow
-The following Mermaid source represents the full data platform workflow for this site — sensors through serving, with governance and observability cross-cuts.
+The following Mermaid source is a full data platform workflow — sensors through serving, with governance and observability cross-cuts.
**Mermaid source** (`docs/assets/diagrams/workflows/data-platform-workflow.mmd`):
@@ -316,6 +316,6 @@ flowchart LR
---
!!! tip "See also"
- - [Mermaid → SVG Workflow Pipeline (Tutorial)](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md) — step-by-step guide to setting up and running the render pipeline
- - [Diagram Style Guide](../../diagrams/style-guide.md) — Mermaid snippet library and formatting conventions
- - [ADR 0015: Standardize Diagrams on Mermaid](../../adr/0015-diagrams-mermaid.md) — the architectural decision record that established this standard
+ - [Rendering Mermaid diagrams to SVG](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md) — step-by-step setup and render pipeline
+ - [Systems Diagramming Best Practices](systems-diagramming-best-practices.md) — layering, boundaries, and entropy control
+ - [Layered Systems Diagrams tutorial](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md) — end-to-end IoT → lakehouse example
diff --git a/docs/best-practices/diagrams/systems-diagramming-best-practices.md b/docs/best-practices/diagrams/systems-diagramming-best-practices.md
index cb06e1e..61c70b5 100644
--- a/docs/best-practices/diagrams/systems-diagramming-best-practices.md
+++ b/docs/best-practices/diagrams/systems-diagramming-best-practices.md
@@ -9,7 +9,6 @@ tags:
# Systems Diagramming Best Practices (Mermaid, SVG, and Diagram Discipline)
-**Objective**: Define a clear doctrine for systems diagramming: conceptual vs physical diagrams, layering, system boundary clarity, and diagram entropy control so that diagrams remain accurate, readable, and maintainable.
---
@@ -110,13 +109,13 @@ The diagram makes the trust boundary explicit: Internet → DMZ → Internal. Us
| Fast iteration and co-editing in docs | Reuse in other tools; consistent layout |
| Diagram lives in the repo with the doc | You need a single export for external use |
-### Repo convention
+### Mermaid source and SVG artifacts
-- **Source**: `.mmd` files under `docs/assets/diagrams/` (e.g. `examples/iot-mqtt-lakehouse-context.mmd`).
-- **Artifact**: `.svg` committed **beside** the `.mmd` (same directory, same basename).
-- **Render**: Run the site’s render script from `tools/diagrams/` (e.g. `npm run render:all`) so SVG is regenerated from source. Never edit the SVG by hand; it will be overwritten.
+- **Source**: version-controlled `.mmd` files (one diagram per file).
+- **Artifact**: `.svg` committed beside the `.mmd` (same directory, same basename) when static docs must render without a Mermaid runtime.
+- **Render**: regenerate SVG from source with `@mermaid-js/mermaid-cli` or an equivalent script. Do not hand-edit generated SVG; it will be overwritten on the next render.
-See [Generating Complex Workflow Diagrams as SVG](svg-workflow-generation.md) and the [Layered Systems Diagrams tutorial](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md) for the full pipeline.
+See [Generating Complex Workflow Diagrams as SVG](svg-workflow-generation.md) and the [Layered Systems Diagrams tutorial](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md) for a worked example.
---
@@ -153,7 +152,7 @@ See [Generating Complex Workflow Diagrams as SVG](svg-workflow-generation.md) an
- **Proprietary formats vs committed SVG**: Prefer committing SVG (and, when possible, Mermaid source) so the doc build and future readers don’t depend on a SaaS tool.
- **Collaboration vs reproducibility**: For one-off workshops, Lucid or draw.io may be faster. For long-lived docs, Mermaid + rendered SVG gives reproducibility and version control.
-This site standardizes on **Mermaid as source** and **SVG as artifact** via `tools/diagrams/`. See the [Mermaid → SVG Workflow Pipeline](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md) and the [Diagram Style Guide](../../diagrams/style-guide.md).
+For documentation that ships as static HTML, **Mermaid as source** and **committed SVG as artifact** is a practical default. See the [Mermaid → SVG workflow tutorial](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md) for setup and rendering steps.
---
@@ -284,7 +283,7 @@ flowchart LR
F -.-> H[Observability]
```
-The full example—source `.mmd` files and rendered SVG artifacts—lives under `docs/assets/diagrams/examples/` and is used in the [Layered Systems Diagrams: Mermaid Source → SVG Artifact](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md) tutorial. Follow that tutorial to generate the SVGs and embed them in MkDocs.
+A full IoT → MQTT → lakehouse example—with `.mmd` sources and rendered SVG—is in the [Layered Systems Diagrams: Mermaid → SVG](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md) tutorial.
---
@@ -293,7 +292,6 @@ The full example—source `.mmd` files and rendered SVG artifacts—lives under
!!! tip "See also"
- **[Layered Systems Diagrams: Mermaid Source → SVG Artifact](../../tutorials/diagrams/layered-systems-diagrams-mermaid-to-svg.md)** — End-to-end tutorial with the IoT → MQTT → Lakehouse example and SVG embedding.
- - **[Generating Complex Workflow Diagrams as SVG](svg-workflow-generation.md)** — Mermaid-first, artifact-driven workflow and repository conventions.
- - **[Mermaid → SVG Workflow Pipeline](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md)** — Existing pipeline for rendering `.mmd` to `.svg` with `tools/diagrams/`.
- - **[Diagram Style Guide](../../diagrams/style-guide.md)** — When to use Mermaid vs SVG, orientation, subgraphs, and accessibility.
+ - **[Generating Complex Workflow Diagrams as SVG](svg-workflow-generation.md)** — Mermaid-first, artifact-driven workflow and layout conventions.
+ - **[Rendering Mermaid diagrams to SVG](../../tutorials/diagrams/mermaid-to-svg-workflow-pipeline.md)** — Set up `mmdc` and render `.mmd` to `.svg`.
- **Deep dives**: [Observability vs Monitoring](../../deep-dives/observability-vs-monitoring.md), [Why Most Data Pipelines Fail](../../deep-dives/why-most-data-pipelines-fail.md) — Systems context that diagrams should reflect.
diff --git a/docs/best-practices/docker-infrastructure/ansible-inventory-management.md b/docs/best-practices/docker-infrastructure/ansible-inventory-management.md
index 71a6950..3c12c85 100644
--- a/docs/best-practices/docker-infrastructure/ansible-inventory-management.md
+++ b/docs/best-practices/docker-infrastructure/ansible-inventory-management.md
@@ -1,6 +1,5 @@
# Ansible Inventory Management Best Practices
-**Objective**: Master Ansible inventory design for scalable, maintainable automation. When you need to manage complex infrastructure, when you want consistent configuration across environments, when you're building enterprise automation—inventory management becomes your weapon of choice.
Ansible inventory is the foundation of automation. Proper inventory design enables scalable deployments, environment separation, and maintainable configuration. This guide shows you how to wield inventory management with the precision of a DevOps engineer.
@@ -639,4 +638,3 @@ Ansible inventory management requires understanding both infrastructure patterns
---
-*This guide provides the complete machinery for Ansible inventory management. The patterns scale from simple single-environment setups to complex multi-environment deployments, from basic host management to advanced dynamic inventory and secrets management.*
diff --git a/docs/best-practices/docker-infrastructure/ansible-performance-optimization.md b/docs/best-practices/docker-infrastructure/ansible-performance-optimization.md
index 874b9e8..82184a8 100644
--- a/docs/best-practices/docker-infrastructure/ansible-performance-optimization.md
+++ b/docs/best-practices/docker-infrastructure/ansible-performance-optimization.md
@@ -1,6 +1,5 @@
# Ansible Performance Optimization Best Practices
-**Objective**: Master Ansible performance patterns for enterprise-scale automation. When you need to optimize automation workflows, when you want to reduce execution time and resource usage, when you're building high-performance infrastructure—performance optimization becomes your weapon of choice.
Ansible performance is critical for enterprise automation. Proper optimization enables faster deployments, reduced resource usage, and scalable automation. This guide shows you how to wield performance optimization with the precision of a systems engineer.
@@ -569,4 +568,3 @@ Ansible performance optimization requires understanding both automation patterns
---
-*This guide provides the complete machinery for Ansible performance optimization. The patterns scale from basic task optimization to advanced enterprise scaling, from simple caching to complex resource management.*
diff --git a/docs/best-practices/docker-infrastructure/ansible-playbook-design.md b/docs/best-practices/docker-infrastructure/ansible-playbook-design.md
index b9d6358..1012453 100644
--- a/docs/best-practices/docker-infrastructure/ansible-playbook-design.md
+++ b/docs/best-practices/docker-infrastructure/ansible-playbook-design.md
@@ -1,6 +1,5 @@
# Ansible Playbook Design Best Practices
-**Objective**: Master Ansible playbook architecture for maintainable, scalable automation. When you need to build complex automation workflows, when you want reusable and testable code, when you're creating enterprise-grade automation—playbook design becomes your weapon of choice.
Ansible playbooks are the heart of automation. Proper playbook design enables maintainable code, reusable components, and scalable automation. This guide shows you how to wield playbook design with the precision of a DevOps engineer.
@@ -685,4 +684,3 @@ Ansible playbook design requires understanding both automation patterns and soft
---
-*This guide provides the complete machinery for Ansible playbook design. The patterns scale from simple single-task playbooks to complex multi-role deployments, from basic automation to advanced enterprise patterns.*
diff --git a/docs/best-practices/docker-infrastructure/ansible-security-hardening.md b/docs/best-practices/docker-infrastructure/ansible-security-hardening.md
index 05d8cd4..3303682 100644
--- a/docs/best-practices/docker-infrastructure/ansible-security-hardening.md
+++ b/docs/best-practices/docker-infrastructure/ansible-security-hardening.md
@@ -1,6 +1,5 @@
# Ansible Security Hardening Best Practices
-**Objective**: Master Ansible security patterns for enterprise-grade automation. When you need to secure automation workflows, when you want to protect sensitive data and credentials, when you're building compliant infrastructure—security hardening becomes your weapon of choice.
Ansible security is critical for enterprise automation. Proper security practices prevent credential exposure, ensure compliance, and protect sensitive infrastructure. This guide shows you how to wield security hardening with the precision of a security engineer.
@@ -609,4 +608,3 @@ Ansible security hardening requires understanding both automation patterns and s
---
-*This guide provides the complete machinery for Ansible security hardening. The patterns scale from basic credential protection to advanced enterprise security, from simple access controls to complex compliance frameworks.*
diff --git a/docs/best-practices/docker-infrastructure/conda-to-docker-migration.md b/docs/best-practices/docker-infrastructure/conda-to-docker-migration.md
index 24bf4a5..00afbc0 100644
--- a/docs/best-practices/docker-infrastructure/conda-to-docker-migration.md
+++ b/docs/best-practices/docker-infrastructure/conda-to-docker-migration.md
@@ -1,6 +1,5 @@
# Migrating Conda to Slim Docker Images
-**Objective**: Transform bloated conda environments into production-ready, slim Docker images. Eliminate conda bloat while maintaining package compatibility and build reproducibility.
When your conda environment is dragging gigabytes of unnecessary packages, you're paying for storage and transfer costs with your soul. This guide shows how to migrate from conda to Docker while keeping images lean and production-ready.
@@ -693,4 +692,3 @@ Migrating from conda to slim Docker images requires understanding your dependenc
---
-*This tutorial provides the complete machinery for migrating conda environments to slim Docker images. The patterns scale from development to production, from megabytes to terabytes.*
diff --git a/docs/best-practices/docker-infrastructure/docker-sbom-trivy-cve-mitigation.md b/docs/best-practices/docker-infrastructure/docker-sbom-trivy-cve-mitigation.md
index aca7c82..c16c284 100644
--- a/docs/best-practices/docker-infrastructure/docker-sbom-trivy-cve-mitigation.md
+++ b/docs/best-practices/docker-infrastructure/docker-sbom-trivy-cve-mitigation.md
@@ -1,6 +1,5 @@
# Best Practices for SBOMs, Trivy Scans, and Automated CVE Mitigation in Docker Images
-**Objective**: Master production-grade container security with SBOM generation, vulnerability scanning, and automated CVE mitigation. When you need to secure containerized workloads, maintain compliance, and reduce manual security toil—these best practices become your foundation.
## Introduction
@@ -1580,5 +1579,4 @@ This approach transforms container security from a manual, error-prone process i
---
-*This guide provides the complete machinery for production-grade container security. The patterns scale from single images to enterprise registries, from manual processes to fully automated remediation workflows.*
diff --git a/docs/best-practices/docker-infrastructure/index.md b/docs/best-practices/docker-infrastructure/index.md
index 9200d01..6b35796 100644
--- a/docs/best-practices/docker-infrastructure/index.md
+++ b/docs/best-practices/docker-infrastructure/index.md
@@ -1,6 +1,5 @@
# Docker & Infrastructure Best Practices
-**Objective**: Master senior-level infrastructure and containerization patterns for production systems. When you need to build robust, scalable infrastructure, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## Containerization & Orchestration
@@ -25,4 +24,3 @@
---
-*These best practices provide the complete machinery for building production-ready infrastructure systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/docker-infrastructure/jinja-best-practices.md b/docs/best-practices/docker-infrastructure/jinja-best-practices.md
index bbc0d8b..a92b244 100644
--- a/docs/best-practices/docker-infrastructure/jinja-best-practices.md
+++ b/docs/best-practices/docker-infrastructure/jinja-best-practices.md
@@ -1,6 +1,5 @@
# Jinja Best Practices: Architecture, Power Tricks, and Safety
-**Objective**: Master Jinja templating for production-grade text generation. When you need to generate configuration files, documentation, or any structured text from data, when you want maintainable and secure templating, when you're building code generation or documentation systems—Jinja becomes your weapon of choice.
Jinja is a compiler for text. Treat it like code: testable, linted, secure, and fast. This guide shows you how to wield Jinja with the precision of a senior template sorcerer, covering everything from basic patterns to advanced security and performance optimization.
@@ -830,4 +829,3 @@ Jinja templating requires understanding both template patterns and security impl
---
-*This guide provides the complete machinery for Jinja templating. The patterns scale from simple text substitution to complex document generation, from basic security to advanced performance optimization.*
diff --git a/docs/best-practices/docker-infrastructure/nginx-best-practices.md b/docs/best-practices/docker-infrastructure/nginx-best-practices.md
index 46b127f..1135a8a 100644
--- a/docs/best-practices/docker-infrastructure/nginx-best-practices.md
+++ b/docs/best-practices/docker-infrastructure/nginx-best-practices.md
@@ -1,6 +1,5 @@
# NGINX Best Practices: Patterns, Hardening, and Multi-API Replication
-**Objective**: Master production-grade NGINX configuration for reverse proxying, load balancing, SSL termination, and multi-API gateways. When you need to front multiple services, replicate traffic for testing, and harden your edge layer—these best practices become your foundation.
## Introduction
@@ -1763,5 +1762,4 @@ This maturity path takes you from a basic reverse proxy to a production-grade AP
---
-*This guide provides the complete machinery for production-grade NGINX configuration. The patterns scale from simple reverse proxies to complex multi-API gateways with traffic mirroring, from single instances to load-balanced clusters.*
diff --git a/docs/best-practices/docker-infrastructure/nginx-production.md b/docs/best-practices/docker-infrastructure/nginx-production.md
index f597138..8488a0a 100644
--- a/docs/best-practices/docker-infrastructure/nginx-production.md
+++ b/docs/best-practices/docker-infrastructure/nginx-production.md
@@ -1,6 +1,5 @@
# Nginx Production Best Practices
-**Objective**: Build bulletproof web infrastructure with Nginx. Handle traffic spikes, secure your applications, and optimize for performance while maintaining operational sanity.
When your web application needs to handle millions of requests, serve static content at lightning speed, and protect against every attack vector imaginable, Nginx becomes your first line of defense. This guide shows you how to configure Nginx for production workloads that demand reliability, security, and performance.
@@ -1183,4 +1182,3 @@ Nginx is the foundation of modern web infrastructure. When configured properly,
---
-*This tutorial provides the complete machinery for building production-ready Nginx infrastructure. The patterns scale from development to production, from single machines to enterprise deployments.*
diff --git a/docs/best-practices/docker-infrastructure/tmux-advanced.md b/docs/best-practices/docker-infrastructure/tmux-advanced.md
index 37d034a..f8de532 100644
--- a/docs/best-practices/docker-infrastructure/tmux-advanced.md
+++ b/docs/best-practices/docker-infrastructure/tmux-advanced.md
@@ -1,6 +1,5 @@
# Advanced Tmux Workflows (2025 Edition)
-**Objective**: Master production-grade tmux workflows for distributed geospatial and data engineering environments. When you need to orchestrate complex multi-service workflows, when you want to maintain persistent development environments, when you're debugging across multiple systems—advanced tmux becomes your weapon of choice.
tmux is more than pane-splitting—it's your personal mission control.
@@ -702,4 +701,3 @@ Advanced tmux requires understanding distributed systems, workflow orchestration
---
-*This guide provides the complete machinery for advanced tmux workflows. The patterns scale from simple pane management to complex orchestration, from local development to distributed team collaboration.*
diff --git a/docs/best-practices/esp32/index.md b/docs/best-practices/esp32/index.md
index d3eeef7..fd52268 100644
--- a/docs/best-practices/esp32/index.md
+++ b/docs/best-practices/esp32/index.md
@@ -6,7 +6,6 @@ tags:
# 🔌 Embedded Systems & ESP32
-**Objective**: Build firmware that is safe, power-efficient, and maintainable. These guides cover the real decisions you make when designing embedded systems around the ESP32 — not just "how to blink an LED" but how to architect, power, secure, and ship firmware responsibly.
## Guides in this section
diff --git a/docs/best-practices/geospatial/geospatial-system-design.md b/docs/best-practices/geospatial/geospatial-system-design.md
index 74ccfe7..5bdd4ab 100644
--- a/docs/best-practices/geospatial/geospatial-system-design.md
+++ b/docs/best-practices/geospatial/geospatial-system-design.md
@@ -1,6 +1,5 @@
# Geospatial System Architecture
-**Objective**: Establish foundational patterns for designing geospatial data systems: primitives, indexing, pipelines, storage, and performance.
## Core geospatial primitives
@@ -84,4 +83,4 @@ For benchmarking and comparison see [Geospatial Benchmarking](../database-data/g
- [Reproducible Data Pipelines](../data/reproducible-data-pipelines.md) — determinism in spatial pipelines
- [Geospatial File Format Choices](../../deep-dives/geospatial-file-format-choices.md) — format tradeoffs
- [The Operational Geometry of Spatial Systems](../../deep-dives/the-operational-geometry-of-spatial-systems.md) — indexing and geometry
-- [Geospatial Data Mesh, Cost & Capacity](../../tutorials/best-practices-integration/geospatial-data-mesh-cost-capacity.md) — data mesh and cost in geospatial
+- [Cost-aware system architecture](../architecture/cost-aware-systems.md) — cost trade-offs in system design
diff --git a/docs/best-practices/git/git-production.md b/docs/best-practices/git/git-production.md
index 4cfdbcf..69158f0 100644
--- a/docs/best-practices/git/git-production.md
+++ b/docs/best-practices/git/git-production.md
@@ -1,6 +1,5 @@
# Git Production Best Practices
-**Objective**: Master Git for enterprise-scale development. Handle monorepos, submodules, and complex workflows while maintaining code quality and team collaboration.
When your codebase spans multiple applications, shared libraries, and deployment environments, Git becomes more than version control—it becomes the foundation of your development workflow. This guide shows you how to structure repositories, manage dependencies, and maintain code quality at scale.
@@ -920,4 +919,3 @@ Git is the foundation of modern development workflows. When configured properly,
---
-*This tutorial provides the complete machinery for building production-ready Git workflows. The patterns scale from development to production, from single developers to enterprise teams.*
diff --git a/docs/best-practices/git/git-workflows-collaboration.md b/docs/best-practices/git/git-workflows-collaboration.md
index b6cc1c0..edc3e09 100644
--- a/docs/best-practices/git/git-workflows-collaboration.md
+++ b/docs/best-practices/git/git-workflows-collaboration.md
@@ -1,6 +1,5 @@
# Git Workflows and Collaboration Patterns
-**Objective**: Master Git workflows for enterprise-grade collaboration and code quality. When you need to coordinate multiple developers, when you want to maintain code quality and project stability, when you're building scalable development processes—Git workflows become your weapon of choice.
Git workflows are the foundation of modern software development. Proper workflow design enables seamless collaboration, maintains code quality, and prevents integration nightmares. This guide shows you how to wield Git workflows with the precision of a senior software engineer, covering everything from basic branching to advanced collaboration patterns.
@@ -654,4 +653,3 @@ Git workflows require understanding both version control mechanics and team coll
---
-*This guide provides the complete machinery for Git workflows and collaboration. The patterns scale from simple feature development to complex enterprise processes, from basic branching to advanced automation and security.*
diff --git a/docs/best-practices/git/gitflow-best-practices.md b/docs/best-practices/git/gitflow-best-practices.md
index 2ac498e..9e0b8c3 100644
--- a/docs/best-practices/git/gitflow-best-practices.md
+++ b/docs/best-practices/git/gitflow-best-practices.md
@@ -1,6 +1,5 @@
# Modern GitFlow Best Practices: Branching, Releases, Hotfixes, and CI/CD Governance
-**Objective**: Master production-grade GitFlow workflows for enterprise software development. When you need structured branching, controlled releases, hotfix processes, and CI/CD governance—these best practices become your foundation.
## Introduction
@@ -1349,5 +1348,4 @@ main (trunk)
---
-*This guide provides the complete machinery for production-grade GitFlow workflows. The patterns scale from small teams to enterprise organizations, from simple releases to complex multi-version maintenance.*
diff --git a/docs/best-practices/git/index.md b/docs/best-practices/git/index.md
index 1c527df..0289ebd 100644
--- a/docs/best-practices/git/index.md
+++ b/docs/best-practices/git/index.md
@@ -1,6 +1,5 @@
# Git Best Practices
-**Objective**: Master Git for enterprise-scale development, collaboration, and code quality. When you need to structure repositories, manage complex workflows, coordinate team collaboration, and maintain code quality at scale—these Git best practices become your foundation.
This collection provides comprehensive guides for Git repository management, workflow design, and team collaboration patterns. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies.
diff --git a/docs/best-practices/go/go-api-design.md b/docs/best-practices/go/go-api-design.md
index d321cd5..65d419f 100644
--- a/docs/best-practices/go/go-api-design.md
+++ b/docs/best-practices/go/go-api-design.md
@@ -1,6 +1,5 @@
# Go API Design Best Practices
-**Objective**: Master senior-level Go API design patterns for production systems. When you need to build robust, scalable RESTful APIs, when you want to follow proven methodologies, when you need enterprise-grade API design patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1103,4 +1102,3 @@ func (h *Handler) ListUsers(w http.ResponseWriter, r *http.Request) {
---
-*This guide provides the complete machinery for building production-ready RESTful APIs in Go applications. Each pattern includes implementation examples, validation strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-cicd-pipelines.md b/docs/best-practices/go/go-cicd-pipelines.md
index f215f9a..5f94615 100644
--- a/docs/best-practices/go/go-cicd-pipelines.md
+++ b/docs/best-practices/go/go-cicd-pipelines.md
@@ -1,6 +1,5 @@
# Go CI/CD Pipelines Best Practices
-**Objective**: Master senior-level Go CI/CD pipeline patterns for production systems. When you need to build robust, automated deployment pipelines, when you want to ensure code quality and security, when you need enterprise-grade CI/CD patterns—these best practices become your weapon of choice.
## Core Principles
@@ -883,4 +882,3 @@ pipeline {
---
-*This guide provides the complete machinery for building robust CI/CD pipelines for Go applications. Each pattern includes implementation examples, security considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-concurrency-optimization.md b/docs/best-practices/go/go-concurrency-optimization.md
index 83727d4..6ce43e3 100644
--- a/docs/best-practices/go/go-concurrency-optimization.md
+++ b/docs/best-practices/go/go-concurrency-optimization.md
@@ -1,6 +1,5 @@
# Go Concurrency Optimization Best Practices
-**Objective**: Master senior-level Go concurrency optimization patterns for production systems. When you need to build high-performance concurrent applications, when you want to optimize existing concurrent code, when you need enterprise-grade concurrency optimization patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1034,4 +1033,3 @@ stats := metrics.GetStats()
---
-*This guide provides the complete machinery for optimizing concurrency in Go applications. Each pattern includes implementation examples, performance strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-concurrency-patterns.md b/docs/best-practices/go/go-concurrency-patterns.md
index 2f31ec3..5c24674 100644
--- a/docs/best-practices/go/go-concurrency-patterns.md
+++ b/docs/best-practices/go/go-concurrency-patterns.md
@@ -1,6 +1,5 @@
# Go Concurrency Patterns Best Practices
-**Objective**: Master senior-level Go concurrency patterns for production systems. When you need to build robust, scalable concurrent applications, when you want to follow proven methodologies, when you need enterprise-grade concurrency patterns—these best practices become your weapon of choice.
## Core Principles
@@ -832,4 +831,3 @@ if !limiter.Allow() {
---
-*This guide provides the complete machinery for building production-ready concurrent Go applications. Each pattern includes implementation examples, testing strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-containerization.md b/docs/best-practices/go/go-containerization.md
index 5527429..35c7a3c 100644
--- a/docs/best-practices/go/go-containerization.md
+++ b/docs/best-practices/go/go-containerization.md
@@ -1,6 +1,5 @@
# Go Containerization Best Practices
-**Objective**: Master senior-level Go containerization patterns for production systems. When you need to build efficient, secure containers, when you want to optimize container performance, when you need enterprise-grade containerization patterns—these best practices become your weapon of choice.
## Core Principles
@@ -963,4 +962,3 @@ services:
---
-*This guide provides the complete machinery for containerizing Go applications efficiently and securely. Each pattern includes implementation examples, security considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-data-processing.md b/docs/best-practices/go/go-data-processing.md
index 79b89f1..06c403a 100644
--- a/docs/best-practices/go/go-data-processing.md
+++ b/docs/best-practices/go/go-data-processing.md
@@ -1,6 +1,5 @@
# Go Data Processing Best Practices
-**Objective**: Master senior-level Go data processing patterns for production systems. When you need to build efficient data pipelines, when you want to process large datasets, when you need enterprise-grade data processing patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1212,4 +1211,3 @@ eventStream.Start()
---
-*This guide provides the complete machinery for building efficient data processing pipelines in Go applications. Each pattern includes implementation examples, performance considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-database-patterns.md b/docs/best-practices/go/go-database-patterns.md
index 4ecc015..b72a9e5 100644
--- a/docs/best-practices/go/go-database-patterns.md
+++ b/docs/best-practices/go/go-database-patterns.md
@@ -1,6 +1,5 @@
# Go Database Patterns Best Practices
-**Objective**: Master senior-level Go database patterns for production systems. When you need to build robust, scalable database applications, when you want to optimize database performance, when you need enterprise-grade database patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1246,4 +1245,3 @@ err := orm.Create(ctx, user)
---
-*This guide provides the complete machinery for building robust database applications in Go. Each pattern includes implementation examples, performance considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-dev-environment.md b/docs/best-practices/go/go-dev-environment.md
index 7c07ea5..bd1e66f 100644
--- a/docs/best-practices/go/go-dev-environment.md
+++ b/docs/best-practices/go/go-dev-environment.md
@@ -1,6 +1,5 @@
# Go Development Environment Best Practices
-**Objective**: Master senior-level Go development environment setup and operation across macOS, Linux, and Windows (WSL and native). Copy-paste runnable, auditable, and production-ready.
## Core Principles
@@ -720,4 +719,3 @@ make benchmark # Run benchmarks
---
-*This guide provides the complete machinery for setting up a production-ready Go development environment. Each pattern includes configuration examples, tooling setup, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/go/go-error-handling.md b/docs/best-practices/go/go-error-handling.md
index 35e1131..9ac5181 100644
--- a/docs/best-practices/go/go-error-handling.md
+++ b/docs/best-practices/go/go-error-handling.md
@@ -1,6 +1,5 @@
# Go Error Handling Best Practices
-**Objective**: Master senior-level Go error handling patterns for production systems. When you need to build robust, maintainable applications, when you want to follow proven methodologies, when you need enterprise-grade error handling patterns—these best practices become your weapon of choice.
## Core Principles
@@ -723,4 +722,3 @@ logger.LogError(err, map[string]interface{}{
---
-*This guide provides the complete machinery for building production-ready error handling in Go applications. Each pattern includes implementation examples, testing strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-geospatial-development.md b/docs/best-practices/go/go-geospatial-development.md
index 66c280f..c09e51a 100644
--- a/docs/best-practices/go/go-geospatial-development.md
+++ b/docs/best-practices/go/go-geospatial-development.md
@@ -1,6 +1,5 @@
# Go Geospatial Development Best Practices
-**Objective**: Master senior-level Go geospatial development patterns for production systems. When you need to build location-aware applications, when you want to process spatial data efficiently, when you need enterprise-grade geospatial patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1134,4 +1133,3 @@ profile, err := profiler.ProfileQuery(ctx, "SELECT * FROM locations WHERE ST_DWi
---
-*This guide provides the complete machinery for building geospatial applications in Go. Each pattern includes implementation examples, performance considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-memory-management.md b/docs/best-practices/go/go-memory-management.md
index d122fa2..dd80848 100644
--- a/docs/best-practices/go/go-memory-management.md
+++ b/docs/best-practices/go/go-memory-management.md
@@ -1,6 +1,5 @@
# Go Memory Management Best Practices
-**Objective**: Master senior-level Go memory management patterns for production systems. When you need to optimize memory usage, when you want to understand Go's garbage collector, when you need enterprise-grade memory management patterns—these best practices become your weapon of choice.
## Core Principles
@@ -844,4 +843,3 @@ debug.SetMemoryLimit(512 * 1024 * 1024) // 512MB limit
---
-*This guide provides the complete machinery for optimizing memory usage in Go applications. Each pattern includes implementation examples, monitoring strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-microservices.md b/docs/best-practices/go/go-microservices.md
index ee0cb5c..181f795 100644
--- a/docs/best-practices/go/go-microservices.md
+++ b/docs/best-practices/go/go-microservices.md
@@ -1,6 +1,5 @@
# Go Microservices Best Practices
-**Objective**: Master senior-level Go microservices patterns for production systems. When you need to build robust, scalable distributed systems, when you want to follow proven methodologies, when you need enterprise-grade microservices patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1041,4 +1040,3 @@ defer span.Finish()
---
-*This guide provides the complete machinery for building production-ready microservices in Go applications. Each pattern includes implementation examples, testing strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-monitoring-observability.md b/docs/best-practices/go/go-monitoring-observability.md
index 6f2052b..f304567 100644
--- a/docs/best-practices/go/go-monitoring-observability.md
+++ b/docs/best-practices/go/go-monitoring-observability.md
@@ -1,6 +1,5 @@
# Go Monitoring & Observability Best Practices
-**Objective**: Master senior-level Go monitoring and observability patterns for production systems. When you need to build comprehensive monitoring solutions, when you want to implement distributed tracing, when you need enterprise-grade observability patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1255,4 +1254,3 @@ alertManager.AddRule(AlertRule{
---
-*This guide provides the complete machinery for implementing comprehensive monitoring and observability in Go applications. Each pattern includes implementation examples, integration strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-performance-tuning.md b/docs/best-practices/go/go-performance-tuning.md
index ee743a6..754559f 100644
--- a/docs/best-practices/go/go-performance-tuning.md
+++ b/docs/best-practices/go/go-performance-tuning.md
@@ -1,6 +1,5 @@
# Go Performance Tuning Best Practices
-**Objective**: Master senior-level Go performance tuning patterns for production systems. When you need to build high-performance applications, when you want to optimize existing code, when you need enterprise-grade performance patterns—these best practices become your weapon of choice.
## Core Principles
@@ -1072,4 +1071,3 @@ metrics.RecordRequest(duration, false)
---
-*This guide provides the complete machinery for optimizing Go applications for maximum performance. Each pattern includes implementation examples, benchmarking strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-testing-best-practices.md b/docs/best-practices/go/go-testing-best-practices.md
index d744466..dba533e 100644
--- a/docs/best-practices/go/go-testing-best-practices.md
+++ b/docs/best-practices/go/go-testing-best-practices.md
@@ -1,6 +1,5 @@
# Go Testing Best Practices
-**Objective**: Master senior-level Go testing patterns for production systems. When you need to build robust, maintainable test suites, when you want to follow proven methodologies, when you need enterprise-grade testing patterns—these best practices become your weapon of choice.
## Core Principles
@@ -800,4 +799,3 @@ func BenchmarkPackageName_FunctionName(b *testing.B) {
---
-*This guide provides the complete machinery for building production-ready test suites in Go applications. Each pattern includes implementation examples, testing strategies, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/go-web-services.md b/docs/best-practices/go/go-web-services.md
index 44f0102..ee5ce7c 100644
--- a/docs/best-practices/go/go-web-services.md
+++ b/docs/best-practices/go/go-web-services.md
@@ -1,6 +1,5 @@
# Go Web Services Best Practices
-**Objective**: Master senior-level Go web service patterns for production systems. When you need to build robust, scalable HTTP services, when you want to follow proven methodologies, when you need enterprise-grade web service patterns—these best practices become your weapon of choice.
## Core Principles
@@ -940,4 +939,3 @@ server.RunWithGracefulShutdown()
---
-*This guide provides the complete machinery for building production-ready web services in Go applications. Each pattern includes implementation examples, security considerations, and real-world usage patterns for enterprise deployment.*
diff --git a/docs/best-practices/go/index.md b/docs/best-practices/go/index.md
index de1b9ca..2f84762 100644
--- a/docs/best-practices/go/index.md
+++ b/docs/best-practices/go/index.md
@@ -1,6 +1,5 @@
# Go Development Best Practices
-**Objective**: Master senior-level Go development patterns for production systems. When you need to build robust, scalable Go applications, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## Core Go Development
@@ -36,4 +35,3 @@
---
-*These best practices provide the complete machinery for building production-ready Go systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/index.md b/docs/best-practices/index.md
index efe9c37..1439d7b 100644
--- a/docs/best-practices/index.md
+++ b/docs/best-practices/index.md
@@ -1,170 +1,54 @@
# Best Practices
-Notes on how systems behave in the wild: trade-offs, failure modes, governance, and the parts of operations that survive contact with production.
+Notes on how systems behave in the wild: trade-offs, failure modes, and the parts of operations that survive contact with production.
These pages answer *why* and *when* more often than *follow these steps in order*. Code shows up when it clarifies a pattern — the point is judgment, not a runbook.
-## 🐍 Python Development
+## Start here
-Comprehensive guides for Python development, covering core Python, web development, and production patterns.
+- [Geospatial system architecture](geospatial/geospatial-system-design.md)
+- [Failure-oriented system design](operations/failure-oriented-design.md)
+- [System resilience and concurrency](operations-monitoring/system-resilience-and-concurrency.md)
+- [PostGIS patterns](postgres/postgis-best-practices.md)
+- [Parquet and GeoParquet](database-data/parquet.md) · [GeoParquet](database-data/geoparquet.md)
-- **[Python Development Overview](python/index.md)** - Core Python, web development, and production patterns
+## Languages and runtimes
-## 🦀 Rust Development
+- [Python](python/index.md)
+- [Rust](rust/index.md)
+- [Go](go/index.md)
+- [R](r/index.md)
-Systems programming and high-performance Rust development patterns.
+## Data and platforms
-- **[Rust Development Overview](rust/index.md)** - Systems programming and high-performance patterns
+- [Database and data management](database-data/index.md)
+- [PostgreSQL](postgres/index.md)
+- [Machine learning and AI](ml-ai/index.md)
+- [Docker and infrastructure](docker-infrastructure/index.md)
+- [Git and version control](git/index.md)
-## 🐹 Go Development
+## Architecture and operations
-High-performance systems programming and web services with Go.
+- [Architecture and design](architecture-design/index.md)
+- [Operations and monitoring](operations-monitoring/index.md)
+- [Data governance](data-governance/index.md)
+- [Security](security/index.md)
-- **[Go Development Overview](go/index.md)** - Systems programming, web services, and performance optimization
+## Embedded and home lab
-## 📊 R Development
+- [ESP32 and embedded](esp32/index.md)
+- [Power electronics](embedded/index.md)
+- [Home automation](home-automation/index.md)
-Statistical computing, data analysis, and reproducible research with R.
+## Diagrams
-- **[R Development Overview](r/index.md)** - Data science, statistical modeling, and reproducible research
+- [Systems diagramming](diagrams/systems-diagramming-best-practices.md)
+- [Mermaid → SVG workflow](diagrams/svg-workflow-generation.md)
-## 🐳 Docker & Infrastructure
+## Creative reference
-Containerization, orchestration, and infrastructure automation patterns.
-
-- **[Docker & Infrastructure Overview](docker-infrastructure/index.md)** - Containerization, orchestration, and automation
-
-## 📝 Git & Version Control
-
-Version control, repository management, and collaboration workflows.
-
-- **[Git & Version Control Overview](git/index.md)** - Repository structure, workflows, and collaboration patterns
-
-## 🗄️ Database & Data Management
-
-Database optimization, data architecture, and governance patterns.
-
-- **[Database & Data Management Overview](database-data/index.md)** - PostgreSQL, data architecture, and governance
-- **GeoParquet** — [Spatial Parquet done right](database-data/geoparquet.md): WKB geometry, GeoParquet metadata & CRS rules, object-storage serving (Garage/MinIO/AWS), and `parquet_s3_fdw` wiring with row-group filtering
-- **[PostgreSQL Development Overview](postgres/index.md)** - PostgreSQL optimization, performance, and enterprise patterns
-
-## 🤖 Machine Learning & AI
-
-ML operations, AI integration, and data science patterns.
-
-- **[Machine Learning & AI Overview](ml-ai/index.md)** - ML operations, AI integration, and data science
-
-## 🏗️ Architecture & Design
-
-System architecture, knowledge management, and data serialization patterns.
-
-- **[Architecture & Design Overview](architecture-design/index.md)** - System architecture, knowledge management, and serialization
-
-### 🎯 Integrated Best Practices Suite
-
-A cohesive suite of five deeply interrelated best-practices documents that form an integrated "constellation" of practices:
-
-1. **[System-Wide Naming, Taxonomy, and Structural Vocabulary Governance](architecture-design/system-taxonomy-governance.md)** - The foundational "lingua franca" that enables all other practices
-2. **[Cross-System Data Lineage, Inter-Service Metadata Contracts & Provenance Enforcement](database-data/data-lineage-contracts.md)** - Comprehensive data lineage and provenance tracking
-3. **[Secure-by-Design Lifecycle Architecture Across Polyglot Systems](security/secure-by-design-polyglot.md)** - Security as a lifecycle concern across all languages
-4. **[Observability as Architecture: Unified Telemetry Models](operations-monitoring/unified-observability-architecture.md)** - Unified observability across all systems
-5. **[Developer Experience (DX) as Infrastructure: Golden Paths, Tooling Ecosystems & Workflow Automation](python/dx-architecture-and-golden-paths.md)** - DX patterns that integrate all practices
-
-These documents are designed to work together, with each referencing the others and showing how they combine to reduce cognitive load, operational entropy, and system complexity.
-
-### 🎯 Architectural Resilience & Data Governance Suite
-
-A complementary suite of four deeply interrelated best-practices documents covering critical architectural gaps:
-
-1. **[Chaos Engineering, Fault Injection, and Reliability Validation](operations-monitoring/chaos-engineering-governance.md)** - Safe fault injection and systematic reliability validation across all system layers
-2. **[Multi-Region, Multi-Cluster Disaster Recovery, Failover Topologies, and Data Sovereignty](architecture-design/multi-region-dr-strategy.md)** - Comprehensive DR strategies with RTO/RPO frameworks and data sovereignty
-3. **[Semantic Layer Engineering, Domain Models, and Knowledge Graph Alignment](database-data/semantic-layer-engineering.md)** - Enterprise semantic layers with RDF/OWL integration and entity resolution
-4. **[ML Systems Architecture: Feature Stores, Model Serving, Experiment Governance, and Cross-System Reproducibility](ml-ai/ml-systems-architecture-governance.md)** - Complete ML lifecycle architecture with reproducibility and governance
-
-These documents form a cohesive framework for resilience, data governance, and ML systems, with each document cross-referencing the others and integrating with the foundational taxonomy and observability practices.
-
-### 🎯 System Efficiency & Data Reliability Suite
-
-A focused suite of three best-practices documents covering critical operational and governance gaps:
-
-1. **[Cost-Aware Architecture & Resource-Efficiency Governance](architecture-design/cost-aware-architecture-and-efficiency-governance.md)** - Comprehensive cost governance frameworks for measuring, optimizing, and controlling resource costs across all system layers
-2. **[Data Freshness, SLA/SLO Governance, and Pipeline Reliability Contracts](data-governance/data-freshness-sla-governance.md)** - Data freshness governance with SLA/SLO frameworks for ETL pipelines, real-time streaming, and geospatial processing
-3. **[Secure Computes, Sandboxing, and Multi-Tenant Isolation for Polyglot Systems](security/secure-sandboxing-and-multi-tenant-isolation.md)** - Comprehensive sandboxing and multi-tenant isolation patterns across all system components
-
-These documents address critical operational concerns: cost efficiency, data reliability, and secure isolation, with each integrating into the broader best-practices ecosystem.
-
-### 🎯 Foundational Architecture & Operations Suite
-
-A critical suite of three best-practices documents covering fundamental architectural and operational gaps:
-
-1. **[Holistic Capacity Planning, Scaling Economics, and Workload Modeling](architecture-design/capacity-planning-and-workload-modeling.md)** - Systematic capacity planning frameworks for modeling workloads, predicting resource needs, and optimizing scaling economics
-2. **[Data Retention, Archival Strategy, Lifecycle Governance & Cold Storage Patterns](data-governance/data-retention-archival-lifecycle-governance.md)** - Comprehensive data retention and archival strategies governing data lifecycle from hot to frozen storage
-3. **[Operational Risk Modeling, Blast Radius Reduction & Failure Domain Architecture](operations-monitoring/blast-radius-risk-modeling.md)** - Risk modeling frameworks for identifying failure domains, modeling blast radius, and designing containment strategies
-
-These documents address foundational concerns: capacity planning, data lifecycle, and operational risk—essential for any mature, coherent technical architecture.
-
-### 🎯 Enterprise Architecture & Governance Suite
-
-A comprehensive suite of six best-practices documents covering critical enterprise-grade architectural and operational gaps:
-
-1. **[Cross-Domain Identity Federation, AuthZ/AuthN Architecture & Identity Propagation Models](security/identity-federation-authz-authn-architecture.md)** - Unified identity federation architecture across all system layers and environments
-2. **[Cross-Environment Configuration Drift Prevention, Promotion Workflows & Release Channels](operations-monitoring/environment-promotion-drift-governance.md)** - Controlled environment promotion governance with drift prevention and release channel management
-3. **[Data Quality SLAs, Validation Layers, and Observability for Tabular, Geospatial, and ML Data](data-governance/data-quality-sla-validation-observability.md)** - Comprehensive data quality governance with SLAs and multi-layer validation
-4. **[API Governance, Backward Compatibility Rules, and Cross-Language Interface Stability](architecture-design/api-governance-interface-stability.md)** - API governance ensuring backward compatibility and interface stability across languages
-5. **[Secret Supply Chains, Encryption Lifecycle Management & Cryptographic Rotation Strategy](security/encryption-lifecycle-and-crypto-rotation.md)** - Encryption lifecycle governance with key rotation and cryptographic supply chain management
-6. **[Observability-Driven Development (ODD), Telemetry-First Coding Practices, and Preemptive Debugging Architecture](operations-monitoring/observability-driven-development.md)** - Embedding observability as a first-class design input with preemptive debugging patterns
-
-These documents address enterprise concerns: identity federation, environment promotion, data quality, API governance, encryption lifecycle, and observability-driven development—essential for mature, coherent distributed systems.
-
-## 🔧 Operations & Monitoring
-
-Performance monitoring, logging, testing, and security patterns.
-
-- **[Operations & Monitoring Overview](operations-monitoring/index.md)** - Performance, logging, testing, and security
-
-## 🔌 Embedded Systems & ESP32
-
-Safe, power-efficient, and maintainable firmware patterns for ESP32-based projects.
-
-- **[Embedded Systems Overview](esp32/index.md)** - Programming architecture, power management, safety, security, sensors, and e-ink displays
-- **[Programming Architecture](esp32/esp32-programming-architecture.md)** - Event loops, FreeRTOS tasks, state machines, ISR safety, memory discipline
-- **[Power Management & Deep Sleep](esp32/power-management-and-deep-sleep.md)** - Sleep modes, RTC memory, battery safety, sub-100 µA design
-- **[Hardware & Electrical Safety](esp32/esp32-hardware-and-electrical-safety.md)** - 3.3 V logic, GPIO limits, level shifting, LiPo safety
-- **[Embedded Security & OTA](esp32/embedded-security-and-ota.md)** - NVS secrets, secure boot, OTA updates, WiFi hygiene
-- **[Sensor Integration](esp32/sensor-integration-best-practices.md)** - I2C/SPI wiring, ADC caveats, filtering, calibration
-- **[E-Ink Display Integration](esp32/e-ink-display-best-practices.md)** - Partial vs full refresh, ghosting, hibernate, frame buffers
-- **[MQTT Security](esp32/mqtt-security-best-practices.md)** - Broker hardening, ACLs, TLS, topic design, credential storage for IoT
-- **[LoRa Best Practices (SX127x)](esp32/lora-best-practices-sx127x.md)** - Spreading factor, duty cycle compliance, antenna layout, power strategy, AES encryption for raw LoRa
-- **[Safety Checklist (Printable)](esp32/esp32-safety-checklist-printable.md)** - Pre-build checklist covering electrical, power, firmware, RF, and physical safety — print to PDF
-- **[ESP32-S3 and C3 Notes](esp32/esp32-s3-and-c3-architecture-notes.md)** - USB native, RISC-V, BLE 5.0, variant selection guide, DS/HMAC security peripherals
-
-## 🔋 Power Electronics & Embedded Hardware
-
-Cross-platform hardware engineering patterns — applies to ESP32, RP2040, and any microcontroller project.
-
-- **[Power Electronics Overview](embedded/index.md)** - Section overview
-- **[Power Electronics for ESP32](embedded/power-electronics-for-esp32.md)** - Logic-level MOSFETs, high/low-side switching, LiPo charging circuits, solar input, buck vs LDO, brownout protection
-
-## 🏠 Home Automation & MQTT
-
-Secure home automation architecture patterns.
-
-- **[Home Automation Overview](home-automation/index.md)** - Security-focused guides for home automation systems
-- **[Home Assistant Security Best Practices](home-automation/home-assistant-security-best-practices.md)** - Access control, TLS, reverse proxy, IoT VLAN, zero-trust mindset
-
-## 🖼️ Diagrams
-
-Systems diagramming doctrine, layering, boundaries, and Mermaid → SVG workflow.
-
-- **[Systems Diagramming Best Practices](diagrams/systems-diagramming-best-practices.md)** — Conceptual vs logical vs physical, layering (L0–L3), boundary clarity, entropy control, pattern library
-- **[Generating Complex Workflow Diagrams as SVG](diagrams/svg-workflow-generation.md)** — Mermaid-first, artifact-driven workflow and repository conventions
-
-## 🎨 Creative & Fun
-
-Creative solutions, opinions, and fun content patterns.
-
-- **[Creative & Fun Overview](creative-fun/index.md)** - Creative solutions, opinions, and fun content
+- [Creative and fun patterns](creative-fun/index.md) — idempotency, Celery notes, and similar (not the same as [Just for Fun tutorials](../tutorials/just-for-fun/index.md))
---
-*Pick a section that matches the problem; cross-links inside each guide point to related tutorials when you need a hands-on walkthrough.*
\ No newline at end of file
+*Pick a section that matches the problem; cross-links inside each guide point to [tutorials](../tutorials/index.md) when you need a hands-on walkthrough.*
diff --git a/docs/best-practices/ml-ai/embeddings-and-vector-databases.md b/docs/best-practices/ml-ai/embeddings-and-vector-databases.md
index 72294f3..7589789 100644
--- a/docs/best-practices/ml-ai/embeddings-and-vector-databases.md
+++ b/docs/best-practices/ml-ai/embeddings-and-vector-databases.md
@@ -1,6 +1,5 @@
# Embeddings, Vector Databases, and Embedding LLMs — A Survival Guide
-**Objective**: Master production-grade embeddings, vector databases, and embedding LLMs for search, RAG, and semantic applications. When you need to build scalable semantic search, when you want to implement production RAG systems, when you're deploying embedding models at scale—vector databases become your weapon of choice.
Vector databases provide the foundation for semantic search and retrieval-augmented generation. Without proper understanding of embeddings, vector operations, and production deployment, you're building fragile systems that miss the power of semantic understanding. This guide shows you how to wield vector databases with the precision of a senior search engineer.
@@ -1228,4 +1227,3 @@ Vector databases provide the foundation for semantic search and RAG systems. Whe
---
-*This guide provides the complete machinery for mastering embeddings, vector databases, and embedding LLMs. The patterns scale from simple similarity search to complex RAG systems, from basic retrieval to advanced semantic understanding.*
diff --git a/docs/best-practices/ml-ai/index.md b/docs/best-practices/ml-ai/index.md
index 83773ba..3490d87 100644
--- a/docs/best-practices/ml-ai/index.md
+++ b/docs/best-practices/ml-ai/index.md
@@ -1,6 +1,5 @@
# Machine Learning & AI Best Practices
-**Objective**: Master senior-level machine learning and AI patterns for production systems. When you need to build robust, scalable ML systems, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## ML Systems Architecture
@@ -20,4 +19,3 @@
---
-*These best practices provide the complete machinery for building production-ready ML and AI systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/ml-ai/mcp-fastapi-stack.md b/docs/best-practices/ml-ai/mcp-fastapi-stack.md
index 121296b..34b5814 100644
--- a/docs/best-practices/ml-ai/mcp-fastapi-stack.md
+++ b/docs/best-practices/ml-ai/mcp-fastapi-stack.md
@@ -1,6 +1,5 @@
# Python MCP Done Right: Tools, Router, and LLM Integration with FastAPI
-**Objective**: Master Model Context Protocol (MCP) for building secure, scalable tool servers with intelligent routing and LLM integration. When you need to expose tools to LLMs safely, when you want to build policy-driven tool orchestration, when you're creating production-ready AI toolchains—MCP becomes your weapon of choice.
MCP provides the foundation for secure, scalable tool integration with LLMs. Without proper understanding of protocol design, security boundaries, and orchestration patterns, you're building vulnerable systems that miss the power of controlled tool execution. This guide shows you how to wield MCP with the precision of a senior systems engineer.
@@ -965,4 +964,3 @@ MCP provides the foundation for secure, scalable AI tool integration. When used
---
-*This guide provides the complete machinery for mastering MCP with Python and FastAPI. The patterns scale from simple tool servers to complex AI orchestration systems, from basic security to advanced production deployment.*
diff --git a/docs/best-practices/ml-ai/ml-systems-architecture-governance.md b/docs/best-practices/ml-ai/ml-systems-architecture-governance.md
index 88c4b23..706c0cf 100644
--- a/docs/best-practices/ml-ai/ml-systems-architecture-governance.md
+++ b/docs/best-practices/ml-ai/ml-systems-architecture-governance.md
@@ -1,6 +1,5 @@
# ML Systems Architecture: Feature Stores, Model Serving, Experiment Governance, and Cross-System Reproducibility
-**Objective**: Establish comprehensive ML systems architecture covering the full lifecycle from data ingestion to model deployment and monitoring. When you need feature stores, when you want experiment governance, when you need reproducibility—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/ml-ai/onnx-model-optimization.md b/docs/best-practices/ml-ai/onnx-model-optimization.md
index c4469ab..4dd818e 100644
--- a/docs/best-practices/ml-ai/onnx-model-optimization.md
+++ b/docs/best-practices/ml-ai/onnx-model-optimization.md
@@ -1,6 +1,5 @@
# ONNX Model Optimization: Production-Ready Machine Learning Deployment
-**Objective**: Master ONNX model creation, optimization, and deployment for production machine learning systems. When you need cross-platform model deployment, when you're optimizing for inference performance, when you need to standardize model formats across frameworks—ONNX becomes your weapon of choice.
ONNX model optimization is the foundation of production machine learning deployment. Without proper ONNX understanding, you're building inefficient models, struggling with framework compatibility, and missing the power of cross-platform deployment. This guide shows you how to wield ONNX with the precision of a machine learning engineer.
@@ -1142,4 +1141,3 @@ ONNX model optimization provides the foundation for production machine learning
---
-*This guide provides the complete machinery for mastering ONNX model optimization. The patterns scale from development to production, from simple conversions to complex deployment architectures.*
diff --git a/docs/best-practices/ml-ai/prompting-llms.md b/docs/best-practices/ml-ai/prompting-llms.md
index 39a7e7d..98ce2f1 100644
--- a/docs/best-practices/ml-ai/prompting-llms.md
+++ b/docs/best-practices/ml-ai/prompting-llms.md
@@ -1,6 +1,5 @@
# Best Practices: Prompting LLMs (with Agent Instructions)
-**Objective**: Prompting isn't magic—it's controlled constraint. Your job: pin context, force structure, and require self-checks. The sections below include a Mermaid framing map, a table of prompt transformations, and a strict instruction block an agentic LLM can execute.
## Prompt Framing Map
@@ -243,4 +242,3 @@ Miss any one and the model improvises; give all six and it builds.
---
-*This guide provides the complete machinery for effective LLM prompting. Each technique is production-ready, copy-paste runnable, and designed for reliable, structured outputs with proper safety and verification.*
diff --git a/docs/best-practices/ml-ai/r-data-exploration.md b/docs/best-practices/ml-ai/r-data-exploration.md
index 1279462..61c1c64 100644
--- a/docs/best-practices/ml-ai/r-data-exploration.md
+++ b/docs/best-practices/ml-ai/r-data-exploration.md
@@ -1,6 +1,5 @@
# R Data Exploration: Tidyverse vs data.table (and Why You Should Prefer data.table)
-**Objective**: Master R data exploration with a focus on production performance, scalability, and memory efficiency. When you need to handle large datasets, when you want maximum performance, when you're building production data pipelines—data.table becomes your weapon of choice.
Tidyverse feels nice, but data.table keeps the lights on when the data gets ugly. While tidyverse excels at readability and teaching, data.table dominates when scale, speed, or reproducibility matter. This guide shows you how to wield data.table with the precision of a senior R engineer.
@@ -871,4 +870,3 @@ R data exploration requires understanding both tidyverse and data.table ecosyste
---
-*This guide provides the complete machinery for mastering R data exploration. The patterns scale from simple data manipulation to complex production pipelines, from basic analysis to advanced performance optimization.*
diff --git a/docs/best-practices/ml-ai/vibe-to-agentic.md b/docs/best-practices/ml-ai/vibe-to-agentic.md
index 845e424..6febcc1 100644
--- a/docs/best-practices/ml-ai/vibe-to-agentic.md
+++ b/docs/best-practices/ml-ai/vibe-to-agentic.md
@@ -1,6 +1,5 @@
# Vibe Coding → Heavy Prompts → Agentic Execution (2025 Edition)
-**Objective**: Master the art of transforming creative "vibe coding" into structured, executable agentic workflows. Start with mood and intent; end with automated delivery. You are distilling chaos into contracts, and contracts into machines.
## The Pipeline (Mental Model)
@@ -286,4 +285,3 @@ Return original structure with redactions. If none found, return unchanged.
---
-*This guide provides the complete machinery for transforming creative vibes into structured, executable agentic workflows. Each pattern includes concrete prompts, schemas, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/operations-monitoring/blast-radius-risk-modeling.md b/docs/best-practices/operations-monitoring/blast-radius-risk-modeling.md
index d5762d0..4387160 100644
--- a/docs/best-practices/operations-monitoring/blast-radius-risk-modeling.md
+++ b/docs/best-practices/operations-monitoring/blast-radius-risk-modeling.md
@@ -1,6 +1,5 @@
# Operational Risk Modeling, Blast Radius Reduction & Failure Domain Architecture: Best Practices
-**Objective**: Establish comprehensive risk modeling frameworks that identify failure domains, model blast radius, and design containment strategies across clusters, databases, data pipelines, and ML systems. When you need to reduce risk, when you want to contain failures, when you need failure domain architecture—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/chaos-engineering-governance.md b/docs/best-practices/operations-monitoring/chaos-engineering-governance.md
index afbd59d..c1918a3 100644
--- a/docs/best-practices/operations-monitoring/chaos-engineering-governance.md
+++ b/docs/best-practices/operations-monitoring/chaos-engineering-governance.md
@@ -1,6 +1,5 @@
# Chaos Engineering, Fault Injection, and Reliability Validation Across Distributed Systems: Best Practices
-**Objective**: Establish safe, systematic chaos engineering practices for validating system resilience across distributed systems, databases, microservices, and data pipelines. When you need to test failure scenarios, when you want to validate SLOs under stress, when you need reproducible fault injection—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/configuration-drift-detection-prevention.md b/docs/best-practices/operations-monitoring/configuration-drift-detection-prevention.md
index b7c9b29..7a3f446 100644
--- a/docs/best-practices/operations-monitoring/configuration-drift-detection-prevention.md
+++ b/docs/best-practices/operations-monitoring/configuration-drift-detection-prevention.md
@@ -1,6 +1,5 @@
# Cross-Environment Configuration Drift Detection & Drift-Proof Deployment Best Practices
-**Objective**: Master production-grade configuration drift detection, prevention, and remediation across multi-environment distributed systems. When you need to ensure consistency across dev/stage/prod, prevent silent failures, and maintain reproducibility—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/configuration-management.md b/docs/best-practices/operations-monitoring/configuration-management.md
index 16bcc9a..bea5cbf 100644
--- a/docs/best-practices/operations-monitoring/configuration-management.md
+++ b/docs/best-practices/operations-monitoring/configuration-management.md
@@ -1,6 +1,5 @@
# Configuration Management, Secrets Lifecycle, and Multi-Environment Drift Control
-**Objective**: Master production-grade configuration management across multi-environment distributed systems. When you need to prevent config drift, manage secrets securely, validate configurations, and maintain consistency across dev/stage/prod—this guide provides complete patterns and implementations.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/environment-promotion-drift-governance.md b/docs/best-practices/operations-monitoring/environment-promotion-drift-governance.md
index 1f958ae..003d09d 100644
--- a/docs/best-practices/operations-monitoring/environment-promotion-drift-governance.md
+++ b/docs/best-practices/operations-monitoring/environment-promotion-drift-governance.md
@@ -1,6 +1,5 @@
# Cross-Environment Configuration Drift Prevention, Promotion Workflows & Release Channels: Best Practices
-**Objective**: Establish comprehensive environment promotion governance that controls configuration flow from dev → staging → prod → isolated environments, prevents drift, and ensures immutable, versioned deployments. When you need controlled promotion, when you want drift prevention, when you need release channels—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/grafana-prometheus-loki-observability.md b/docs/best-practices/operations-monitoring/grafana-prometheus-loki-observability.md
index 80a5f74..8d0b8d4 100644
--- a/docs/best-practices/operations-monitoring/grafana-prometheus-loki-observability.md
+++ b/docs/best-practices/operations-monitoring/grafana-prometheus-loki-observability.md
@@ -1,6 +1,5 @@
# Best Practices for Monitoring & Observability with Grafana, Prometheus, Loki, Node Exporter, and Structured Logging
-**Objective**: Master production-grade observability with Grafana, Prometheus, Loki, and Node Exporter. When you need comprehensive monitoring, structured logging, alerting, and dashboards at scale—these best practices become your foundation.
## Introduction
@@ -2066,5 +2065,4 @@ This observability stack provides the foundation for understanding your systems
---
-*This guide provides the complete machinery for production-grade observability. The patterns scale from single instances to enterprise deployments, from basic metrics to full distributed tracing.*
diff --git a/docs/best-practices/operations-monitoring/grafana.md b/docs/best-practices/operations-monitoring/grafana.md
index 703a76b..4b70011 100644
--- a/docs/best-practices/operations-monitoring/grafana.md
+++ b/docs/best-practices/operations-monitoring/grafana.md
@@ -1,6 +1,5 @@
# Grafana Best Practices
-**Objective**: Master production-grade Grafana deployment, configuration, and operations for enterprise observability. When you need to build reliable monitoring dashboards, when you want to scale observability across teams, when you're responsible for keeping the lights on at 3AM—Grafana best practices become your weapon of choice.
Grafana is not just a pretty dashboard tool. Treat it as production-critical software: upgrade it, provision it, secure it, observe it.
@@ -1115,4 +1114,3 @@ Grafana requires understanding distributed systems, observability patterns, and
---
-*This guide provides the complete machinery for production Grafana operations. The patterns scale from simple dashboards to complex enterprise monitoring, from basic alerting to advanced observability platforms.*
diff --git a/docs/best-practices/operations-monitoring/index.md b/docs/best-practices/operations-monitoring/index.md
index 9376b9e..5b0f606 100644
--- a/docs/best-practices/operations-monitoring/index.md
+++ b/docs/best-practices/operations-monitoring/index.md
@@ -1,6 +1,5 @@
# Operations & Monitoring Best Practices
-**Objective**: Master senior-level operations and monitoring patterns for production systems. When you need to build robust, scalable operational systems, when you want to follow proven methodologies, when you need enterprise-grade patterns—these best practices become your weapon of choice.
## Performance & Reliability
@@ -32,4 +31,3 @@
---
-*These best practices provide the complete machinery for building production-ready operational systems. Each guide includes architectural patterns, configuration examples, and real-world implementation strategies for enterprise deployment.*
diff --git a/docs/best-practices/operations-monitoring/logging-observability.md b/docs/best-practices/operations-monitoring/logging-observability.md
index bb76e17..5dda308 100644
--- a/docs/best-practices/operations-monitoring/logging-observability.md
+++ b/docs/best-practices/operations-monitoring/logging-observability.md
@@ -1,6 +1,5 @@
# Structured Logging & Observability Best Practices (2025 Edition)
-**Objective**: Master production-grade logging and observability for modern data and DevOps environments. When you need to debug complex distributed systems, when you want to trace data flows across services, when you're responsible for keeping systems alive at 3AM—structured logging becomes your weapon of choice.
Logs aren't for printing—they are lifelines. Make them structured, contextual, and queryable, or drown in noise.
@@ -819,4 +818,3 @@ Structured logging requires understanding distributed systems, observability pat
---
-*This guide provides the complete machinery for structured logging and observability. The patterns scale from simple application logging to complex distributed system monitoring, from basic log aggregation to advanced observability platforms.*
diff --git a/docs/best-practices/operations-monitoring/observability-driven-development.md b/docs/best-practices/operations-monitoring/observability-driven-development.md
index ceac368..a826b98 100644
--- a/docs/best-practices/operations-monitoring/observability-driven-development.md
+++ b/docs/best-practices/operations-monitoring/observability-driven-development.md
@@ -1,6 +1,5 @@
# Observability-Driven Development (ODD), Telemetry-First Coding Practices, and Preemptive Debugging Architecture: Best Practices
-**Objective**: Establish comprehensive observability-driven development practices that embed telemetry, logging, tracing, and metrics as first-class design inputs from day one. When you need observability-first design, when you want preemptive debugging, when you need telemetry standards—this guide provides the complete framework.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/operational-resilience-and-incident-response.md b/docs/best-practices/operations-monitoring/operational-resilience-and-incident-response.md
index d5af3d8..0aec479 100644
--- a/docs/best-practices/operations-monitoring/operational-resilience-and-incident-response.md
+++ b/docs/best-practices/operations-monitoring/operational-resilience-and-incident-response.md
@@ -1,6 +1,5 @@
# Operational Resilience and Incident Response: Best Practices for Distributed Systems
-**Objective**: Master production-grade operational resilience across RKE2, Postgres, Prefect, Redis, ML pipelines, and air-gapped clusters. When you need to survive incidents, maintain uptime, and recover gracefully—this guide provides complete operational playbooks and incident response frameworks.
## Introduction
diff --git a/docs/best-practices/operations-monitoring/release-management-and-progressive-delivery.md b/docs/best-practices/operations-monitoring/release-management-and-progressive-delivery.md
index be92ff4..1370eed 100644
--- a/docs/best-practices/operations-monitoring/release-management-and-progressive-delivery.md
+++ b/docs/best-practices/operations-monitoring/release-management-and-progressive-delivery.md
@@ -7,7 +7,6 @@ tags:
# Release Management, Change Governance, and Progressive Delivery: Best Practices for Distributed Systems
-**Objective**: Master production-grade release management and progressive delivery for distributed systems. When you need to safely deploy changes across applications, databases, data pipelines, and ML systems—this guide provides complete patterns and implementations.
## Introduction
@@ -101,7 +100,7 @@ graph TB
- Ticket creation with labels
**2. Design & ADR**:
-- Architecture Decision Record (see [ADR Guide](../architecture-design/adr-decision-governance.md))
+- Architecture Decision Record (see [Documentation](../architecture-design/documentation.md))
- Design review
- Risk classification
@@ -1875,7 +1874,6 @@ kubectl apply -f manifests/k8s/ -n prod
## See Also
-- **[ADR and Technical Decision Governance](../architecture-design/adr-decision-governance.md)** - Decision recording
- **[Configuration Management](configuration-management.md)** - Config governance
- **[System Resilience](../operations-monitoring/system-resilience-and-concurrency.md)** - Resilience patterns
- **[Testing Best Practices](../operations-monitoring/testing-best-practices.md)** - Testing strategies
diff --git a/docs/best-practices/operations-monitoring/secrets-config.md b/docs/best-practices/operations-monitoring/secrets-config.md
index 467f232..6a93d5b 100644
--- a/docs/best-practices/operations-monitoring/secrets-config.md
+++ b/docs/best-practices/operations-monitoring/secrets-config.md
@@ -1,6 +1,5 @@
# Secrets & Configuration Management Best Practices (2025 Edition)
-**Objective**: Master production-grade secrets and configuration management for modern data engineering and DevOps environments. When you need to secure sensitive data across environments, when you want to prevent credential exposure, when you're responsible for keeping secrets safe—proper secrets management becomes your weapon of choice.
Secrets are the sharpest knives in your system. If you don't sheath them properly, you'll cut yourself—or worse, everyone else.
@@ -810,4 +809,3 @@ Secret management requires understanding security patterns, access control, and
---
-*This guide provides the complete machinery for secret and configuration management. The patterns scale from simple environment variables to complex enterprise secret stores, from basic access control to advanced security operations.*
diff --git a/docs/best-practices/operations-monitoring/system-resilience-and-concurrency.md b/docs/best-practices/operations-monitoring/system-resilience-and-concurrency.md
index fb6d1be..c44fe1f 100644
--- a/docs/best-practices/operations-monitoring/system-resilience-and-concurrency.md
+++ b/docs/best-practices/operations-monitoring/system-resilience-and-concurrency.md
@@ -1,2390 +1,110 @@
-# System Resilience, Rate Limiting, Concurrency Control & Backpressure: Best Practices for Distributed Systems
+# System resilience, concurrency, and backpressure
-**Objective**: Master production-grade resilience patterns for distributed systems. When you need to prevent cascading failures, control concurrency, implement rate limiting, handle backpressure, and maintain SLOs under load—this guide provides complete patterns and implementations.
+Distributed systems fail when work enters faster than you can finish it. Connection pools exhaust, queues grow without bound, and one slow dependency stalls everything upstream. This page is about **bounding concurrency** and **signaling overload** before you run out of memory or threads.
-## Introduction
+It is not a catalog of every resilience pattern. It focuses on what I reach for first: limits, backpressure, and honest degradation.
-Distributed systems face constant pressure: traffic spikes, resource exhaustion, network failures, and cascading errors. Without proper resilience patterns, systems fail catastrophically. This guide provides a complete framework for building resilient, high-performance distributed systems.
+## When to care
-**What This Guide Covers**:
-- Concurrency control (threads, async, worker pools)
-- Rate limiting (token bucket, leaky bucket, sliding windows)
-- Backpressure mechanisms
-- Load shedding and graceful degradation
-- Circuit breakers and bulkheads
-- SLO-driven design and error budgets
-- Autoscaling strategies
-- GIS and ML-specific resilience patterns
+- **Postgres-heavy services**: pool size is a hard ceiling. Slow queries hold connections; retries multiply load.
+- **Async Python / Go / Rust services**: goroutines or tasks are cheap until they are not — memory and file descriptors still cap you.
+- **GIS and batch jobs**: large geometries and cold caches create latency spikes that look like “random” timeouts unless you limit fan-out.
+- **ML inference**: GPU memory and batch queues need explicit bounds; “just add workers” often makes p95 worse.
-**Prerequisites**:
-- Understanding of distributed systems, microservices, and concurrency
-- Familiarity with Python async, Go concurrency, Rust async
-- Experience with Kubernetes, databases, and message queues
+If the workload is a once-a-day cron, you probably do not need a circuit breaker. You need a timeout and an alert when the job overruns.
-## Purpose & Importance
+## Concurrency: pick a boundary and enforce it
-### Why Resilience Patterns Are Essential
+**Threads / processes (sync code)**
+Use a fixed worker pool size derived from CPU and I/O, not `number_of_cpus * 100`. For DB-bound work, size against **pool connections**, not cores.
-**Postgres-Heavy Systems**:
-- Connection pool exhaustion causes cascading failures
-- Slow queries block all requests
-- Replication lag causes read inconsistencies
-- Without rate limiting, databases are overwhelmed
+**Async (Python asyncio, etc.)**
+Semaphores limit in-flight requests. Without them, every accepted socket spawns unbounded task creation under load.
-**GIS Computation Pipelines**:
-- Spatial queries are CPU-intensive and slow
-- Tile generation can saturate compute resources
-- Isochrone calculations are expensive
-- Without backpressure, GIS services become unresponsive
+**Go**
+Worker pools with buffered channels; respect `context` cancellation on shutdown.
-**ML/ONNX Inference Flows**:
-- Model loading is expensive (cold starts)
-- GPU saturation causes queue buildup
-- Feature extraction can be slow
-- Without load shedding, inference latency degrades
+**Rust**
+`tokio::sync::Semaphore` or dedicated worker tasks; avoid unbounded `spawn` in handlers.
-**Real-Time Dashboards**:
-- Grafana queries can overwhelm databases
-- NiceGUI WebSocket connections consume resources
-- Real-time updates create thundering herd
-- Without rate limiting, dashboards become unusable
-
-**ETL & Streaming Workloads**:
-- Large datasets overwhelm memory
-- Streaming backpressure causes data loss
-- Pipeline failures cascade downstream
-- Without circuit breakers, entire pipelines fail
-
-**Microservice Architectures**:
-- Service failures cascade across dependencies
-- Network partitions cause timeouts
-- Retry storms overwhelm services
-- Without isolation, one service failure takes down all
-
-**Air-Gapped Clusters**:
-- Limited compute resources
-- No external fallbacks
-- Resource contention is severe
-- Without load shedding, systems become unusable
-
-### Failure Cascade Without Resilience
-
-```mermaid
-graph TD
- Start["Traffic Spike"] --> API["API Service
(No Rate Limit)"]
- API -->|"Unlimited Requests"| PG["Postgres
(Connection Exhaustion)"]
- API -->|"Unlimited Requests"| Redis["Redis
(Memory Exhaustion)"]
- API -->|"Unlimited Requests"| ML["ML Service
(GPU Saturation)"]
-
- PG -->|"Slow Queries"| API2["API Service
(Request Timeout)"]
- Redis -->|"Evictions"| API3["API Service
(Cache Misses)"]
- ML -->|"Queue Buildup"| API4["API Service
(Latency Spike)"]
-
- API2 -->|"Cascading Failures"| Down["System Down"]
- API3 -->|"Cascading Failures"| Down
- API4 -->|"Cascading Failures"| Down
-
- style Start fill:#ffcccc
- style Down fill:#ff0000
-```
-
-### Failure Cascade With Resilience
-
-```mermaid
-graph TD
- Start["Traffic Spike"] --> RL["Rate Limiter
(Token Bucket)"]
- RL -->|"Throttled Requests"| API["API Service
(Controlled Load)"]
-
- API -->|"Pooled Connections"| PG["Postgres
(Connection Pool)"]
- API -->|"Cached Responses"| Redis["Redis
(TTL + Eviction)"]
- API -->|"Queued Requests"| ML["ML Service
(Worker Pool)"]
-
- PG -->|"Slow Query"| CB["Circuit Breaker
(Opens)"]
- CB -->|"Fallback"| Cache["Cache Response"]
-
- ML -->|"Queue Full"| BP["Backpressure
(Reject New)"]
- BP -->|"Load Shed"| LS["Load Shedder
(Drop Low Priority)"]
-
- Cache -->|"Graceful Degradation"| Stable["System Stable"]
- LS -->|"Graceful Degradation"| Stable
-
- style Start fill:#ffcccc
- style Stable fill:#ccffcc
-```
-
-## Core Concepts
-
-### Concurrency Control
-
-#### Thread-Level Concurrency
-
-**Python Threading**:
-
-```python
-# concurrency/thread_pool.py
-from concurrent.futures import ThreadPoolExecutor
-import threading
-
-class ThreadPoolController:
- def __init__(self, max_workers: int = 10):
- self.max_workers = max_workers
- self.executor = ThreadPoolExecutor(max_workers=max_workers)
- self.semaphore = threading.Semaphore(max_workers)
-
- def submit(self, func, *args, **kwargs):
- """Submit task with concurrency control"""
- with self.semaphore:
- return self.executor.submit(func, *args, **kwargs)
-
- def shutdown(self, wait=True):
- """Shutdown thread pool"""
- self.executor.shutdown(wait=wait)
-
-# Usage
-pool = ThreadPoolController(max_workers=10)
-future = pool.submit(expensive_operation, arg1, arg2)
-result = future.result()
-```
-
-**Go Goroutine Pool**:
-
-```go
-// concurrency/worker_pool.go
-package concurrency
-
-import (
- "context"
- "sync"
-)
-
-type WorkerPool struct {
- workers int
- jobQueue chan func()
- wg sync.WaitGroup
-}
-
-func NewWorkerPool(workers int, queueSize int) *WorkerPool {
- return &WorkerPool{
- workers: workers,
- jobQueue: make(chan func(), queueSize),
- }
-}
-
-func (wp *WorkerPool) Start(ctx context.Context) {
- for i := 0; i < wp.workers; i++ {
- wp.wg.Add(1)
- go func() {
- defer wp.wg.Done()
- for {
- select {
- case job := <-wp.jobQueue:
- job()
- case <-ctx.Done():
- return
- }
- }
- }()
- }
-}
-
-func (wp *WorkerPool) Submit(job func()) error {
- select {
- case wp.jobQueue <- job:
- return nil
- default:
- return ErrQueueFull
- }
-}
-
-func (wp *WorkerPool) Stop() {
- close(wp.jobQueue)
- wp.wg.Wait()
-}
-```
-
-**Rust Tokio Concurrency**:
-
-```rust
-// concurrency/tokio_pool.rs
-use tokio::sync::Semaphore;
-use std::sync::Arc;
-
-pub struct ConcurrencyLimiter {
- semaphore: Arc,
-}
-
-impl ConcurrencyLimiter {
- pub fn new(max_concurrent: usize) -> Self {
- Self {
- semaphore: Arc::new(Semaphore::new(max_concurrent)),
- }
- }
-
- pub async fn execute(&self, f: F) -> T
- where
- F: Future