Imported from infrabytes/homelab-kubernetes (
AGENTS.md). Install upstream withnpx skills add infrabytes/homelab-kubernetes. Copyright stays with the author.
AGENTS.md
Guidance for AI coding agents working in this repository.
Project overview
GitOps-driven homelab Kubernetes cluster. A Talos Linux cluster runs on Proxmox VE; it is provisioned with Terragrunt + OpenTofu (infra/) and applications are delivered by ArgoCD from this same repo (platform/ + apps/).
Stack: Talos Linux - Kubernetes - Cilium (kube-proxy-free, L2 LB, Gateway API, WireGuard encryption, Hubble) - ArgoCD - cert-manager (Let's Encrypt DNS-01) - external-dns - VictoriaMetrics + Loki + Grafana (self-hosted observability) - SOPS/age - Renovate.
Repository layout
infra/ Terragrunt/OpenTofu units: cluster -> viewer-kubeconfig, addons -> argocd-config
(grafana-cloud-config is order-independent, runs in parallel)
env.hcl ALL unit inputs centralized (versions, nodes, secrets)
root.hcl shared remote_state (S3 backend on SeaweedFS, pbkdf2-encrypted) +
artifact-sync hooks (kubeconfig/talosconfig <-> bucket `artifacts/` prefix)
secrets.sops.yaml single SOPS-encrypted secrets file (never plaintext)
scripts/ sync-artifacts.sh: bucket <-> /var/tmp/homelab-artifacts (symlinked into cluster/artifacts/)
cluster/ Talos cluster + Cilium; talos_machine/talos_cluster drive
in-place Talos + Kubernetes upgrades; writes
artifacts/kubeconfig + talosconfig
viewer-kubeconfig/ Mints the view-only client cert + kubeconfig (CSR API, no CA key extraction)
addons/ Installs ArgoCD, cert-manager, external-dns, OpenBao namespace + seal Secret, ARC namespaces
grafana-cloud-config/
Grafana Cloud watchdog (grafana/grafana provider): Talos
folder, PVE-maintenance dashboard + rules, cluster-heartbeat
rule; no cluster dependency
argocd-config/ ArgoCD bootstrap ApplicationSet (app-of-appsets)
argocd/appsets/ committed ApplicationSets (platform, apps, pdeu, tenants), applied via the
Terraform bootstrap ApplicationSet
argocd/tenants/ tenant list (source of truth for the tenants ApplicationSet + the
openbao postStart tenant config)
charts/ in-repo Helm charts rendered by ApplicationSets (tenant-access)
platform/ ArgoCD-managed cluster-level resources (network, issuer,
observability, networkpolicies, metrics-server,
kubelet-serving-cert-approver, homelab-runner +
cluster-viewer RBAC)
helm-charts/ one parent ArgoCD app (app-of-apps) for the Helm chart
Applications (cert-manager, csi-driver-nfs, external-dns,
Longhorn, ARC, vcluster, grafana-cloud, victoria-metrics,
loki, agent-sandbox, ...)
grafana-dashboards/ the GrafanaDashboard CRs grafana-operator syncs into
the local Grafana instance (chart-shipped plus vendored
JSON in in-repo ConfigMaps)
observability/ local Grafana stack: namespace, Grafana CR + datasource CRs,
LAN HTTPRoute + tailnet Ingress, GrafanaAlertRuleGroup CRs,
contact point/notification policy, heartbeat CronJob
apps/ ArgoCD-managed applications (one subdir per app)
ansible/ rolling PVE host maintenance playbook (proxmox01-04, windrunner timer)
.github/ CI workflows + scripts (pre-commit, PR preview diff)
.pre-commit-config.yaml the single lint/format gate
renovate.json dependency automation
Commands
pre-commit run --all-files # run every lint/format gate locally
cd infra && terragrunt validate --all
cd infra && terragrunt plan --all
cd infra && terragrunt apply --all # CAUTION: mutates the live cluster
cd infra && terragrunt destroy --all # CAUTION: destroys everything
# Infra changes are normally applied by Atlantis from PRs (automerge on):
# atlantis.yaml at the repo root defines the projects/workflow; the server
# runs on the windrunner VM (atlantis.icaninto.space, deploy files in
# ~/appdata/atlantis on that host, outside this repo).
# Apply consumes the plan written by plan (-out $PLANFILE), so a failed or
# stale plan blocks apply: comment `atlantis plan` again before `atlantis apply`.
# Atlantis plans and applies the *branch*, so a branch that predates another
# merge to the same unit re-applies the old values and silently reverts it —
# hence `execution_order_group`, and hence: rebase/re-plan a PR after anything
# else in its unit merged (a stale apply of #228 reverted #227's argo-cd sizing).
# Plans and applies are ordered by each project's `execution_order_group`
# (mirroring the Terragrunt DAG: cluster -> viewer-kubeconfig/addons ->
# argocd-config), with abort_on_execution_order_fail; a new unit needs a group.
# Never add `depends_on`: with groups it blocks the first global apply
# (runatlantis/atlantis#5791).
# Every project pins `terraform_distribution: opentofu` + `terraform_version`
# (matching the image's tofu): without a distribution Atlantis looks for
# Terraform, re-resolves the newest release its `required_version` allows on
# every plan, and fails when its downloader rejects HashiCorp's content type.
# Bump that pin and the image's tofu together.
# The image also carries kubectl (the cluster unit's converge.tf local-exec
# needs it) and is tagged atlantis-homelab:vX.Y.Z there.
sops infra/secrets.sops.yaml # edit secrets (re-encrypts on save)
# PVE host maintenance (dry run; canary = add --limit proxmox02, needs explicit go-ahead):
sops exec-env infra/secrets.sops.yaml \
'ansible-playbook -i ansible/inventory.yml ansible/proxmox-node-updates.yml --check --diff'
.github/scripts/test-renovate.py # local Renovate dry-run; verify dep extraction/updates (no branches/PRs)
Agent environment: OMP sessions run directly on the host (no sandbox). Tools and the SOPS age key are the host's; if a tool is missing from PATH, ask the user.
Validation & pre-commit
Every change must pass .pre-commit-config.yaml; CI runs it on push/PR on the self-hosted homelab-runner (.github/workflows/pre-commit.yaml, pinned tool versions, tools served from the pod-local cache hydrated from the shared seed volume; see README.md). PR runs are diff-scoped (--from-ref/--to-ref); pushes to main and PRs touching hook config run the full --all-files sweep. Jobs run fully concurrent (per-pod emptyDir caches; only the warm workflow writes the seed, atomically — it wipes the pod-local trees before rebuilding, so the seed never accumulates old versions). Notable hooks:
terragrunt_fmt+terraform_tflint(config:infra/cluster/.tflint.hcl) forinfra/- local
terragrunt-validatehook (.github/scripts/terragrunt-validate.sh):terragrunt validate --alloninfra/; skips without the SOPS age key (e.g. CI), enforcing locally where secrets decrypt yamllint(.yamllint.yaml; ignoressecrets.sops.yaml, 160-char lines)kubeconformonplatform/,apps/andargocd/YAMLansible-lintonansible/(config:ansible/.ansible-lint)- local
ansible-inventory-checkhook (.github/scripts/check-ansible-inventory.py): thepvegroup ofansible/inventory.ymland thenodes = { ... }block ofinfra/env.hclmust agree on inventory host ↔pve_node↔vm_id↔ Talos hostname (the env.hcl node IPs are the Talos VM addresses, not the PVE host IPs, so the check does not compare IPs) - local
argocd-apps-checkhook (.github/scripts/check-argocd-apps.py): for everyApplicationmanifest underplatform//apps/, pulls its helm/OCI chart attargetRevision— or clones a git-sourced chart and renders itspath(agent-sandbox) — and renders it withhelm template --include-crds(ArgoCD's default; opt-out viahelm.skipCrds) with its release name, namespace, values (skips if helm/PyYAML/git are missing). It also rejects duplicate mapping keys insidehelm.values(YAML last-key-wins silently drops the earlier block — that is how the alloy-logs resources vanished) and requires every rendered container and init container to declare a CPU request, a memory request and a memory limit; chart-internal containers with no values key are exempted, with a reason, in.github/scripts/resource-coverage-allowlist.yaml - local
tenant-config-checkhook (.github/scripts/check-tenants.py): the tenant list (argocd/tenants/tenants.json, members with an explicit tailnetidentityand optional GitHubuser) and the openbao postStartTENANTSmarker (<tenant>:<namespace>[:<login>[|<login>]]) must agree on tenant/namespace/logins; members are validated (identity charset, duplicate identities,@githubidentities matching their login); alsosh -ns the postStart script and renderscharts/tenant-accessper tenant, asserting the RoleBinding/ClusterRoleBinding/ServiceAccount/SecretStore and that both bindings' subjects equal the member identities (skips helm/PyYAML when missing) detect-secrets(baseline:.secrets.baseline); never add plaintext secrets. The baseline carries no result entries — known false positives are filtered instead:infra/secrets.sops.yamland.sops.yamlare excluded by file (encryption is enforced by thesops-encryptedhook), and lines containingpasswordKey:/secretKeyRef/argocdServerAdminPassword(chart key names, not credentials) are excluded by line. Keep it that way: drift-prone result entries get auto-rewritten by the hook on every partial commit. New false positive → extend the--exclude-lines/--exclude-filesregexes in the baseline'sfilters_used(regenerate viadetect-secrets scan --exclude-files '<regex>' --exclude-lines '<regex>').renovate-config-validatorforrenovate.jsonruff-check(astral-sh/ruff-pre-commit) for Python files- local
sops-encryptedhook;*.sops.yamlfiles must be encrypted
Renovate
renovate.json is the single source of truth for dependency scanning. Coverage today:
infra/env.hclversion pins → regex custom managers (talos_version,kubernetes_version,cilium_chart_version,gateway_api_crds_version). Every version field inenv.hclMUST have a matchingcustomManagersentry.- Terraform
helm_release(ArgoCD ininfra/addons/main.tf) →terraformmanager (helm datasource). Thegrafana/grafanaprovider pin ininfra/grafana-cloud-config/versions.tfis also auto-discovered by theterraformmanager — nocustomManagersentry needed for provider pins in.tffiles. - ArgoCD
Application/ApplicationSetmanifests underplatform/,apps/andargocd/(cert-manager, external-dns, spegel, the ARC charts, the committedplatform/apps/tenantsApplicationSets) →argocdmanager (helm datasource; OCI charts like cert-manager and spegel resolve via thedockerdatasource on quay.io/ghcr.io; a git-sourced chart like agent-sandbox resolves itstargetRevisionvia thegit-tagsdatasource). ThetenantsApplicationSet references an in-repo chart bypath, and in-repo charts (charts/*) carry no upstream version — neither needs a manager. - Agent Sandbox's controller image tag → regex custom manager scoped to
platform/helm-charts/agent-sandbox/application.yaml(dockerdatasource onregistry.k8s.io). The image tag must track the chart tag, so both pins are grouped into oneplatformbranch by thematchFileNames: platform/**package rule; never bump just one. - The hubble-observer chart is pinned twice: the chart Application's
targetRevision(OCI chart,dockerdatasource) and theoci.referenceof its dashboard CR inplatform/grafana-dashboards/flows-dashboard.yaml(regex custom manager,dockerdatasource on ghcr.io) — the dashboard JSON is pulled from the chart artifact, so both pins are the same version; theplatform/**rule groups them into one branch. Never bump just one. - The Cilium chart is pinned twice:
cilium_chart_versionininfra/env.hcl(cluster release) and thehubble-uiapp'stargetRevision. The standalone Hubble UI must match the deployed chart, so thematchDatasources: helm+matchPackageNames: ciliumrule groups both into onecilium chartbranch; never bump just one. - Raw manifests (
metrics-server,kubelet-serving-cert-approver) →kubernetesmanager (image + API versions). ansible/requirements.ymlcollection pins → the built-inansible-galaxymanager (galaxy-collectiondatasource on galaxy.ansible.com); no custom manager needed — thegalaxy(roles) datasource cannot resolve collections and fails lookup..github/workflows/pre-commit.yaml+.github/workflows/warm-tool-cache.yamlCLI pins (terragrunttg_version, tofutofu_version, tflinttflint_version; kubeconform + pyyaml stay pre-commit-only) → regex custom managers whosemanagerFilePatternscover both files;.pre-commit-config.yaml→pre-commitmanager. Theactions/setup-pythonpython-versionpin needs NO custom manager: the built-ingithub-actionsmanager already tracks it in every workflow file — don't add one.atlantis.yaml: the per-projectterraform_version(OpenTofu, matching the atlantis-homelab image) is tracked by the same regex custom manager as the workflows'tofu_version; never bump it alone..github/workflows/pr-preview.yaml+.github/workflows/warm-tool-cache.yaml: kubectl pin (Azure/setup-kubectl, must matchkubernetes_versioninenv.hcl), the helm CLI pin (Azure/setup-helm), and the in-vCluster Argo CD chart version (helm datasource,argoproj.github.io/argo-helm) → regex custom managers. Thehelm/helmCLI pin custom manager covers all three workflow files.- Runner CLI pins (the CLIs without setup actions that the
pr-previewworkflow runs on the self-hosted runner: vcluster, argocd-diff-preview, gh, kind) → regex custom managers matching the job-env*_VERSIONvalues in.github/workflows/pr-preview.yaml(github-releases datasource). Installed per-run by the workflow from those pins, so a bump is one env value. kubectl and helm are NOT pinned there; they come from the setup actions, which install into the pod-localRUNNER_TOOL_CACHE(/opt/tool-cache/local/runner, hydrated from the shared seed volume; seeREADME.md).
PR preview (.github/workflows/pr-preview.yaml): on PRs touching platform//apps/, a self-hosted ARC runner (homelab-runner) renders the base vs. target diff through a dedicated Argo CD instance that runs inside the vCluster (namespace argocd-preview, installed idempotently via helm upgrade --install) and posts one comment with output/diff.md via argocd-diff-preview. The diff is the entire PR gate — no preview deployment, no health checks (the former vCluster deploy/health stage with its per-app allowlist, chart mirroring, CRD supply, secret copy and cleanup job was retired; runtime issues surface post-merge via the ArgoCD sync + Grafana, rollback is a git revert). The runner SA arc-runner reaches the vCluster via the host Roles in platform/homelab-runner/rbac.yaml (argocd-diff-preview namespace access retired with the host diff instance; vcluster-connect reads the vcluster namespace incl. the vc-vcluster kubeconfig Secret). The runner PAT lives in the arc-runner-auth Secret (namespace arc-runners), synced by ESO from OpenBao (secret/arc-runner-auth); the addons unit only creates the namespace. argocd-diff-preview connects to the in-vCluster Argo CD through the vCluster kubeconfig (vcluster connect --server https://vcluster.vcluster.svc:443 --insecure), port-forwards to the server, and authenticates with the initial admin password from the cluster.
Runtime validation — none in the PR gate: chart apps and raw manifests are covered by the diff render plus pre-commit (check-argocd-apps.py renders every Application with its chart and values, kubeconform validates raw repo manifests). Deploy-time behavior (crashloops, bad images, webhook rejections) is observed after merge: ArgoCD syncs main automatically, Grafana monitors the result, rollback is a git revert. The vcluster chart (integrations.metricsServer.enabled: true) still serves the metrics API by proxying the host metrics-server.
When adding or removing a component:
- New version pin in
env.hcl(or any new*.hcl/workflow file) → add the corresponding custom manager; on removal, delete the entry. - New manifests under
platform/,apps/orargocd/→ auto-discovered by theargocd/kubernetesmanagers, no config change needed (just ship the manifest). - Removing an Application/manifest → no config change needed; remove the manifest only.
Verify renovate changes with .github/scripts/test-renovate.py before pushing (pinned renovate version; needs gh auth or GITHUB_TOKEN).
- Custom-manager file matching is
managerFilePatternsin Renovate 44.x (wasfileMatchbefore v44). Always validate with the pinned version (pre-commit run renovate-config-validator --all-files); a stalenpxcache runs an old renovate and reports false errors. - Local dry-runs need Node matching renovate's engine (44.x:
^24.11.0; use the pre-commitnode_env-ltsnode if the system node is too new) and a token viaRENOVATE_HOST_RULES(platform=localdoes not auto-injectRENOVATE_TOKEN).
Conventions
- Terragrunt: put unit inputs (and derived values) in
infra/env.hcl, onelocals.cluster/locals.viewer_kubeconfig/locals.addons/locals.grafana_cloud/locals.argocd_configmap. A unit'sterragrunt.hclonly wiresinputs = local.env.locals.<unit>and carries no values.apply --allrunsclusterfirst, thenviewer-kubeconfigandaddons(independent siblings), thenargocd-config;destroy --allreverses it. Providers resolve the cluster connection fromcluster/artifacts/kubeconfigviaenv.hcl, so noKUBECONFIGexport is needed. The credential files (kubeconfig,talosconfig,viewer-kubeconfig) live in the machine-invariantcredentials_dir(/var/tmp/homelab-artifacts— state stores each resource's absolute filename, so a checkout path would plan creates in every other workspace) with symlinks incluster/artifacts/; they sync with the state bucket'sartifacts/prefix viabefore_hook/after_hookinroot.hclrunningscripts/sync-artifacts.sh(download before plan/apply, upload after apply — a missing bucket object is a cache miss: plan shows create, apply regenerates and re-uploads; hook output is silent, contents never appear in Atlantis output). The cilium/gateway-api inline manifests stay local-only incluster/artifacts/(accepted plan noise when absent). Thegrafana-cloud-configunit is the exception: Grafana Cloud API via the grafana/grafana provider (stack service-account token from SOPS), nodependencies {}block, runs in parallel. - Secrets: one SOPS-encrypted file,
infra/secrets.sops.yaml(recipients in.sops.yaml). Edit only viasops. The age key is NOT in the repo. - ArgoCD:
platform/= cluster-scoped/admin resources;apps/= regular applications. Theplatform/appsApplicationSets are committed underargocd/appsets/and applied by the Terraform-managed bootstrap ApplicationSet (infra/argocd-config/): one intermediate Application per appset dir, nameappset-{{path.basename}}. They generate frommainwithautomatedsync (prune + selfHeal), so pushing tomaindeploys. Generated Applications apply server-side (syncOptions: ServerSideApply=truein the appset templates), which the 468 KB vendored Node Exporter Full ConfigMap needs — client-side apply stores that payload in the 256KB-cappedlast-applied-configurationannotation and fails; nested chart Applications keep their ownsyncPolicy. Addingapps/<name>/(or a newplatform/<name>/) is picked up automatically. Adding a new ApplicationSet = add a dir underargocd/appsets/(no Terraform change). Helm chartApplications go underplatform/helm-charts/<chart>/so theplatformApplicationSet generates a single parent app that applies them (avoids one outer app per chart). Tenant access is data-driven:argocd/tenants/tenants.jsonfeeds both the committedtenantsApplicationSet (onecharts/tenant-accessrender per tenant: RBAC, ESO ServiceAccount, namespaced OpenBao store) and the openbao postStartTENANTSmarker (mount, policies, k8s/OIDC roles); thetenant-config-checkhook fails when the two drift. The local Grafana stack is GitOps:platform/observability/carries theGrafana/GrafanaDatasource/GrafanaAlertRuleGroup/GrafanaContactPoint/GrafanaNotificationPolicy/GrafanaFolderCRs plus the CloudNativePGClusterthat backs Grafana's database, andplatform/grafana-dashboards/theGrafanaDashboardCRs — chart-shipped ones (imported from chart ConfigMaps or the OCI artifact) plus the in-repo ConfigMaps (Network Policies, the vendored grafana.com 1860 "Node Exporter Full"). No Terragrunt unit talks to the local Grafana;grafana-cloud-configkeeps only the Cloud watchdog (Talos folder, maintenance dashboard, maintenance + heartbeat rules). - Storage homes: NAS (
192.168.0.22:/nas) via thenfs-nasStorageClass from thecsi-driver-nfschart backs the observability TSDBs (VictoriaMetrics, Loki); Longhorn keeps cluster state (Grafana's CloudNativePG cluster, ARC tool cache) and itsstorageReservedis never lowered to fit a TSDB.nfs-nasusesreclaimPolicy: Retain(a PVC prune must not delete 90 days of data) andmountPermissions: "0777"(the driver creates the volume subdir as a squashed user; Loki writes as uid 10001). - Observability split: cluster metrics/logs go only to the in-cluster VictoriaMetrics/Loki (namespace
observability, both onnfs-nas); Alloy writes to those two endpoints; its onlymetricsTuningtrims are the kube-state-metrics NetworkPolicy include and the host-metrics allowlist union behind the Talos node dashboard. Grafana Cloud is the out-of-cluster watchdog: it holds the PVE-maintenance rules, thecluster-heartbeatdead-man rule (fed byplatform/observability/heartbeat-cronjob.yamlpushing into Cloud Loki) and Discord as the notification transport (hand-configuredGrafanaBotcontact point + its policy); it gets no cluster metrics/logs. Adaptive Metrics was retired with the free tier, so nografana-adaptive-metricsprovider is in play. - Versions: dependency pins live in
infra/env.hcl(talos_version,kubernetes_version,cilium_chart_version,gateway_api_crds_version),infra/*/versions.tf(provider pins),.github/workflows/pre-commit.yaml+.github/workflows/warm-tool-cache.yaml(CLI tools),.pre-commit-config.yaml,ansible/requirements.yml(collection pins), ArgoCDApplicationcharttargetRevisions, and the PR preview job env in.github/workflows/pr-preview.yaml(vcluster, argocd-diff-preview, gh, kind). Renovate drives bumps, so don't bump versions manually without a reason. When adding/removing a pinned dependency, updaterenovate.jsonper the Renovate section and verify with.github/scripts/test-renovate.py. - Resource sizing: derive requests/limits from the longest retained Grafana window — memory request = p95 rounded up to 64Mi (32Mi floor), memory limit = max(1.5x request, 1.25x observed peak), CPU request = p95 rounded up to 10m, and no CPU limits. State the measurement window in the values comment, and re-derive from the alerting history rather than hand-tuning. Exceptions are one-shot escalations with a recorded re-measure date (currently the ArgoCD application-controller 1Gi/2Gi and applicationset-controller 128Mi/256Mi, 2026-09-24), and the four new observability workloads (VictoriaMetrics, Loki, Grafana, csi-driver-nfs) whose initial estimates are dated 2026-09-16 and re-measured two weeks after cutover.
- Style: Conventional Commits —
type(scope): subject, imperative, lowercase (conventional-commitskill); Clean Code principles (clean-codeskill) — comments terse, why-not-what. - Documentation: keep the root
README.md(stack overview, bootstrap/day-2 workflows) current — reflect added/removed components, bootstrap-order and exposure changes there. Detailed ops docs stay ininfra/*/README.md.
Rules & guardrails
- Never push to
main, force-push, or rewrite history. All changes go through a branch + PR: create a dedicated git worktree under.worktrees/(or use thegithubtool'spr_checkout), commit there, push the branch, and open a PR tomain— no direct commits tomain, no asking first. Auth: gh CLI credential helper (gh auth git-credential, HTTPS, no SSH). - Never run
terragrunt apply/destroy/importagainst the live cluster unless the user explicitly asks. These are destructive, real-world operations. Note: opening a PR that touchesinfra/(orenv.hcl) lets Atlantis auto-apply on the PR and auto-merge — treat pushing a branch + PR as an apply trigger for infra units. - Never commit unencrypted secrets, private keys, or Terraform state.
.terraform/,.terragrunt-cache/,*.tfstate*, andartifacts/are gitignored, so don't force-add them. - Never edit
infra/secrets.sops.yamlas plaintext or decrypt it into a committed file. Re-encrypt withsops -e -i(CI rejects unencrypted*.sops.yaml). - Never delete or regenerate machine secrets (
talos_machine_secrets); the local state andartifacts/talosconfigcarry cluster identity. Losing them means the cluster can't be re-adopted. - Don't touch
artifacts/outputs (kubeconfig/talosconfig/viewer-kubeconfig); they are generated by thecluster/viewer-kubeconfigunits..terragrunt-cache/is transient, so ignore it. - Don't modify local tooling (host-side helpers) unless asked.
Further reading
README.md(repo root): stack overview, getting started, day-2 workflowsinfra/README.md: Terragrunt workflow, SOPS/age, unit orderinginfra/cluster/README.md: Talos provisioning, Cilium inline manifest, upgrades, pitfallsinfra/addons/README.md: ArgoCD bootstrap and adding appsinfra/argocd-config/README.md: ApplicationSets- Session continuity: work may span worktrees under
.worktrees/(gitignored) and multiple sessions — checkgit status/git branchand ask where a task left off before continuing.
