Imported from willpxxr/willpxxr-live (
AGENTS.md). Install upstream withnpx skills add willpxxr/willpxxr-live. Copyright stays with the author.
AGENTS.md
CLAUDE.mdis a symlink to this file — one source of truth; any tool that auto-loads either name gets the same content.
Infrastructure-as-code for willpxxr.com: cloud infrastructure via Terraform (remote
state in Terraform Cloud, org willpxxr, workspace willpxxr-live), Kubernetes
cluster configuration via ArgoCD (GitOps).
Agent rules (binding)
What the rest of this file describes is how things work; these are what you must do. They apply to any agent (human or LLM) working in this repo.
- main is live. Every push to
maintriggers a real Terraform apply (TFC) and a real cluster reconcile. There is no PR/plan gate. Run theverify-infra-changeskill before every push — before, not after. - Git ships, hands don't. No manual
kubectl apply,helm upgrade, orterraform applyagainst the cluster/TFC. Changes land via commits tomainonly; the documented one-time bootstraps (ArgoCD install, Terraform bootstrap objects) are the only exceptions. - Scope: work in
de/hetznerunless explicitly told otherwise.uk/prodis legacy being wound down — treat as read-only. - Secrets: never hard-code, log, or echo them. Terraform inputs come from
TFC workspace variables; cluster secrets only via the 1Password
ExternalSecretpattern (see "Sensitive information" and the app conventions below). - Decisions get recorded. A change needing a migration plan or spanning
multiple applies is an RFC-subtype WEP; a single non-obvious decision is an
ADR-subtype WEP — write it in
docs/wep/alongside the change (seedocs/wep/README.mdfor the split). - Follow the conventions below (network-policy posture, ExternalSecret
key layout,
movedblocks,terraform fmt, Auth0 scope naming). New apps get scaffolded via thenew-gitops-appskill, not improvised. - Verify against reality. CRD field names, defaults, and chart versions are confirmed against the live cluster's schemas or the pinned upstream source — never assumed from memory or from docs for a different version.
- Commit hygiene: commit only what you were asked to ship (no unrelated
working-tree changes), concise imperative message matching
git logstyle, no amending or force-pushing published history. - Keep this file current — stale docs are bugs (see the section below).
Clusters
de/hetzner(gitops/clusters/de/hetzner/cluster/) — the active cluster; essentially all current work happens here. Talos Linux on Hetzner Cloud (nbg1), provisioned via thehcloud-talosTerraform module (terraform/hetzner.tf). Cilium CNI (kube-proxy replacement, native routing, Hubble; ArgoCD-managed viaapps/cilium/since WEP-0011 — do not re-enable the module'sdeploy_cilium) — withsocketLB.hostNamespaceOnly: trueinapps/cilium/values.yaml, which is load-bearing: Cilium's socket-LB force-enables with kube-proxy replacement (cilium/cilium#47417), and with it active in pod namespaces, pod→ClusterIP traffic bypasses the netfilter DNAT hooks the Tailscale operator's proxies depend on — silently blackholing tailnet traffic to LoadBalancer Services. Envoy Gateway + Envoy AI Gateway for ingress/routing (all host routing + TLS termination; wildcard*.internal.willpxxr.comvia cert-manager DNS-01), cert-manager, external-dns (apps/external-dns/, syncs*.internal.willpxxr.comrecords to Cloudflare from Gateway HTTPRoutes + Services with theexternal-dns.alpha.kubernetes.io/hostnameannotation — theservicesource was added for the Valheim server's NodePort, WEP-0015), external-secrets (1Password backend), the Tailscale operator. Tailnet exposure is L3-only: oneloadBalancerClass: tailscaleLoadBalancer Service (the Envoy data plane); the per-hostname Tailscale L7 Ingresses were removed (WEP-0003) — their hostname machinery is what caused the 2026-07/08svc:gatewayoutage, so prefer not to bring them back. The Valheim dedicated server (apps/valheim/, WEP-0015) is the cluster's only public non-HTTP workload: atype: NodePortService (ports 32456-32458 UDP,externalTrafficPolicy: Local) with the hcloud firewall opened via the hcloud-talos module'sextra_firewall_rules. DNS (valheim.willpxxr.com) is ExternalDNS-managed via theservicesource + hostname annotation —externalTrafficPolicy: Localensures only the node running the pod is published. The MCP gateway fronts third-party MCP servers; its credential vault (apps/mcp-token-vault, WEP-0006) stores per-provider OAuth tokens (envelope-encrypted in Supabase Postgres) and injects them upstream — OAuth-needing MCP backends ride the vault's single path-dispatched proxy listener (/<provider>path on 8081; direct Envoy-side injection awaits upstreamcredentialOverridesupport for MCP backends, see WEP-0007) — credentials are connected out of band via its browser UI attokens.internal.willpxxr.com(Auth0token_vault:use), since MCP in-band elicitations are swallowed by the AI Gateway proxy.
Tech stack
- Terraform
>= 1.11(seeterraform/providers.tf'srequired_version) — remote backend (terraform/backend.tf), no localterraform applyexpected. - Providers:
cloudflare(v5 — seedocs/wep/0004-adr-cloudflare-provider-v5.md; token minting for downscoped tokens uses the account-token surface, which needs the automation token to hold Account/API-Tokens Read+Write),hcloud/talos(the active cluster),tailscale,onepassword,auth0,logtail(Better Stack),supabase(mcp-token-vault project, WEP-0006 -- static Management-API PAT; the API has no OIDC surface for the TFC workload-identity path Tailscale uses). Pluskubernetes/helm/kubectl/tlsfor the small set of bootstrap-only k8s objects Terraform manages directly (see below). - ArgoCD (self-managed app-of-apps, see
gitops/clusters/de/hetzner/cluster/argocd/) reconciles everything undergitops/. - CI (
.github/workflows/): OSSF Scorecard, dependency-review, Checkov (IaC security scanning), and a Packer build for the Talos node image. - pre-commit (
.pre-commit-config.yaml): gitleaks, end-of-file-fixer, trailing-whitespace.
Repository structure
.
├── terraform/ # ALL Terraform (TFC VCS-driven: workspace
│ │ # working dir = terraform/, trigger patterns
│ │ # terraform/** -- pushes without .tf changes
│ │ # don't trigger runs)
│ ├── backend.tf, providers.tf, variables.tf, data.tf, locals.tf, main.tf
│ ├── hetzner.tf # de/hetzner Talos cluster (hcloud-talos module)
│ ├── tailscale.tf # Tailscale ACL/OAuth client + bootstrap k8s namespaces/Secret
│ ├── auth0.tf # Auth0 clients/scopes for every Envoy Gateway SecurityPolicy
│ ├── argocd.tf # ArgoCD redis auth: random_password -> 1Password item (ESO-synced)
│ ├── betterstack*.tf # Better Stack API creds/source + dashboards/SLOs
│ ├── synthetic.tf, supabase.tf, internaldns.tf
│ └── oci.tf, moves.tf # Legacy OCI cluster + decommission `moved` blocks
├── packer/talos/ # Talos node snapshot image build
├── scripts/ # Helper scripts (gateway login, model sync, etc.)
├── services/ # In-cluster services with source in this repo (mcp-token-vault, WEP-0006; CI builds/pushes the image)
├── .opencode/ # opencode global config + ai-gateway-auth plugin (see "opencode config" below)
├── docs/wep/ # Willpxxr enhancement proposals (RFC + ADR subtypes) -- see docs/wep/README.md
├── docs/runbooks/ # Operational runbooks (e.g. docs/runbooks/k8s-upgrade.md for Kubernetes version bumps)
└── gitops/clusters/de/hetzner/cluster/
├── argocd/ # root-app (app-of-apps) + the single apps ApplicationSet
├── policy/ # cluster-wide Cilium policies (its own app)
└── apps/<name>/ # one directory per deployed component: app.yaml (+ values.yaml, manifests)
Development workflow
- Commit straight to
main— this is a single-developer homelab repo; PRs are pure overhead here and are not used. (Earlier history has some PR merges from before this was settled — that's not a convention to continue.)mainisn't branch-protected at the GitHub level either, which is consistent with that. - Committing to
mainwhile a feature-branch checkout is active (e.g. WEP-0005 staging on a branch): usegit worktree add /tmp/<name> mainand commit there. Do NOTgit checkout mainin the branch checkout — files tracked on main but absent from the branch are deleted from the working tree on switch (bit us twice withservices/), and stash-dances around modified files are fragile. - Terraform: Terraform Cloud applies on every push to
main(VCS-driven). No localterraform applyexpected. - GitOps: ArgoCD auto-syncs (selfHeal + prune) the Applications its
ApplicationSet renders from
gitops/clusters/de/hetzner/cluster/apps/*/app.yamlon push tomain. No manualkubectl apply. - A handful of k8s objects are created directly by Terraform rather than GitOps —
only for genuine bootstrap ordering (things ArgoCD/external-secrets themselves
depend on), e.g. the
external-secrets/tailscale/argocdnamespaces and the 1Password ESO service-account token Secret interraform/tailscale.tf. The one-time ArgoCD install itself was also bootstrap (helm, releaseargocd, which theargocdapp then adopts). Everything else lives ingitops/. - Because pushing to
maintriggers a real Terraform apply and a real ArgoCD sync with no PR/plan-only step in between to catch mistakes first, run theverify-infra-changeskill (see below) before pushing, not after something breaks.
opencode config
.opencode/ holds the canonical opencode global config: opencode.jsonc
repoints the built-in synthetic provider at the AI gateway
(https://ai.tailb40090.ts.net/v1), and plugin/ai-gateway-auth.ts registers
the Auth0 PKCE OAuth flow on it (opencode auth login synthetic, silent
refresh afterwards). ~/.config/opencode/ symlinks to these files, so edit
them here, not there. The Auth0 client_id in the plugin is a public PKCE
native client (no secret exists), so committing it is fine.
GitOps app conventions (de/hetzner cluster)
ArgoCD renders every app from a per-app app.yaml via the single
argocd/apps-applicationset.yaml (git files generator over
apps/*/app.yaml + policy/app.yaml). Each apps/<name>/ directory has:
app.yaml— the app's definition:name,namespace(omit for cluster-scoped apps),dir(cluster-relative path), and any combination of:helm:entries (repoURL/chart/version/releaseName+ optional per-entryvalues:file; for OCI, repoURL is the full artifact path and the registry is registered hostname-only inapps/argocd/values.yaml),kustomize:entries (repoURL/path/revision), andmanifests:(path globs, synced as plain manifests — nothing is synced implicitly). See WEP-0001 for the full schema.- Disabling an app — rename
app.yaml→app.yaml.disabled(and back to re-enable). The generator glob (apps/*/app.yaml) matches by filename, so a disabled app stops being generated and the ApplicationSet controller (defaultsyncpolicy) deletes the Application; theresources-finalizercascades the delete to the helm release and manifests. There is no in-fileenabled:flag — the generator matches filenames, not content, andtemplatePatchcannot prevent Application creation. values.yaml— helm values (never synced as a manifest; consumed via the$appvalueFiles reference).namespace.yaml+network-policy.yaml—CiliumNetworkPolicy, default-deny posture. Every namespace gets explicitallow-same-namespace+allow-dns-egress(+allow-kube-apiserver-egressif the workload talks to the API server — list every controller that watches resources, not just the obvious ones). Egress to a specific third-party SaaS host whose IPs aren't enumerable usestoEntities: [world]restricted to port 443, one rule per external dependency, each with adescriptionexplaining why.kube-systemis the one deliberate exception (allow-all— OVH-managed components whose requirements aren't documented). UI/exposure apps additionally need the Envoy data plane to reach their backend: add an egress rule toapps/envoy-gateway/network-policy.yaml'sallow-backend-egress(one per backend).externalsecret.yamlif it needs a secret:ExternalSecret(secretStoreRef: ClusterSecretStore/onepassword,refreshInterval: 6h— required and admission-enforced; 1Password service-account rate limits are tight, on-demand refresh via theforce-syncannotation) pulling from thekubernetes1Password vault, key convention<item-title>/credentials/<field>. When the secret's origin is another Terraform-managed provider resource rather than something typed in by hand, a matchingonepassword_itemresource writes it into that vault from the relevant.tffile (seeterraform/tailscale.tf,terraform/betterstack.tf,terraform/argocd.tf) — prefer this over asking a human to paste a secret into 1Password whenever the upstream service has a usable Terraform provider. When it doesn't (e.g. Synthetic's LLM-gateway API key,terraform/synthetic.tf), Terraform creates the item with a placeholder value pluslifecycle { ignore_changes = [section_map] }(so later applies don't revert the hand-pasted key) and a human pastes the real value into the item afterwards; the key layout still applies.- Ordering: none is encoded cross-app — ArgoCD apps whose CRDs aren't in place yet fail sync and converge via the retry policy. Within an app, Secrets/CRDs apply before CRs.
- Controller-mutated CRs (Envoy Gateway / AI Gateway kinds): the
controllers write defaults into specs at reconcile time; per-kind
ignoreDifferenceslive inapps/argocd/values.yaml(resource.customizations.ignoreDifferences.*— note the camelCasejqPathExpressions/jsonPointerskeys, and null-safe jq[]?iterations: iterating null silently disables the whole normalizer). Extend them when a new controller-defaulted kind appears. - ArgoCD's own chart lives in
apps/argocd/(releaseargocd, adopted from the one-time helm bootstrap) together with its exposure (Tailscale Ingress inenvoy-gateway-system+ HTTPRoute + Auth0 SecurityPolicy) — the app manages ArgoCD itself;dexand the redis init Job are disabled (the init Job deadlocks GitOps syncs; the redis auth is ESO-managed, seeterraform/argocd.tf).
Observability (otel-collector) — DISABLED
The observability stack is disabled as of 2026-09-03 (app dirs kept,
definitions renamed to app.yaml.disabled; see the disabling convention in the
GitOps app conventions above): the otel-collector-agent DaemonSet (host
metrics, kubelet stats, container log tailing), the otel-collector-gateway
Deployment (cluster metrics, k8s events, Prometheus scraping, OTLP receiver —
they exported to Better Stack), and Beyla (eBPF auto-instrumentation,
Grafana's OBI). The Better Stack source itself is still Terraform-managed
(terraform/betterstack.tf, logtail provider) — it simply receives nothing.
Re-enable by renaming the app.yaml.disabled files back to app.yaml.
Caveats while disabled:
apps/otel-collector/'s namespace/network-policies/externalsecret are Flux-era orphans (no ArgoCD Application ever synced them — noapp.yamlwas carried over in WEP-0001) — they needed a one-time manualkubectl delete ns otel-collectoronce the workloads dropped.- Envoy data-plane tracing is commented out in
apps/envoy-gateway-config/envoyproxy.yaml(thetelemetry.tracingblock targets the dropped collector; itsReferenceGrantwent with it) — kept inline to uncomment together with the collector's re-enable. - If Beyla comes back, remember its built-in
DefaultExcludeInstrumenthard-excludeskube-system(and other platform namespaces) regardless of thediscovery.instrumentglob (OBIpkg/obi/config.go) — components living there only ever get client-side spans.
Terraform conventions
- DNS records live in
terraform/locals.tfunderlocal.records; redirects underlocal.redirects; WAF rules underlocal.waf(VPN allow-list:local.lists.vpn). - Use
movedblocks interraform/moves.tfwhen renaming/moving resources, to avoid destructive replacement. - The Kubernetes version for the active cluster is
terraform/hetzner.tf'smodule.talos.kubernetes_version/talos_version(keep in sync withpacker/talos/talos.pkr.hcl's default). - Auth0 scope naming (
terraform/auth0.tf) is<resource>:<tier>, tier one ofget/admin/use— see the comment at the top ofterraform/auth0.tffor the full rationale before adding a new scope. Auth0 resource-server identifiers and all SecurityPolicy redirect URLs / audiences usehttps://<name>.internal.willpxxr.com(MCP carries a/mcppath suffix) — keepterraform/auth0.tfidentifiers and the matching SecurityPolicy fields in lockstep. - Run
terraform fmtbefore committing.
Sensitive information
All values below are Terraform Cloud workspace variables (var.*), sensitive,
never hard-coded:
var.cloudflare_api_tokenvar.hetzner_tokenvar.tailscale_bootstrap_oauth_client_idvar.onepassword_terraform_service_account_tokenvar.auth0_domain,var.auth0_mgmt_client_id,var.auth0_mgmt_client_secretvar.betterstack_api_tokenvar.supabase_token
Keeping this file current
This file is a live map of the repo, not a snapshot. When you learn something
material during a session — a new convention, a section here that's gone stale or
wrong, a new provider/app, a non-obvious decision and the reasoning behind it —
update the relevant section as part of that work, not as an afterthought. Edit in
place rather than appending a changelog: this file should describe what is true
now, not a history of edits (git log is the history). Architecture/workflow/
convention-level facts belong here; "why this exact line of code" belongs in a code
comment next to that code.
Project skills
.claude/skills/ holds skills scoped to this repo — they encode the repeatable
parts of working here so they don't have to be re-derived each session:
verify-infra-change— run before pushing tomain:terraform fmt, akubectl --dry-run=clientpass against the live cluster's CRDs for changed manifests, and ahelm templaterender for any changed appvalues.yaml.new-gitops-app— scaffolds a newapps/<name>/directory (app.yaml + values.yaml + manifests) following the conventions above; the ApplicationSet picks the directory up automatically.
If you find yourself doing the same multi-step verification or scaffolding twice, that's a sign it should become a skill (or an update to an existing one) rather than tribal knowledge re-derived every time. Skills should evolve the same way this file does — if you hit a case a skill doesn't handle, extend the skill, don't just route around it once and move on.
