Imported from zncdatadev/operator-go (
AGENTS.md). Install upstream withnpx skills add zncdatadev/operator-go. Copyright stays with the author.
AGENTS.md
Project Overview
operator-go is a Golang SDK/framework for building Kubernetes operators. It provides a reusable reconciliation framework, CRDs, and utilities for creating product-specific operators.
Key Features:
- GenericReconciler: Template Method Pattern-based reconciliation framework
- Extension System: Hook-based customization at cluster/role/role-group levels, with per-product registries
- Resource Builders: Fluent builders for StatefulSet, Service, ConfigMap, PDB, RBAC, ServiceAccount
- Role Declaration:
RoleProvider/RoleCatalog— a product states each role as data, once per reconcile, with the CR in hand - Config Folding:
FoldCommonConfig/FoldProductConfig— the framework folds its half of theconfigblock, the product folds its own, both under one rule set - Config Generation: Multi-format config file generation (XML, YAML, Properties, Env, INI)
- Logging Config: Framework-aware logging configuration generation (Log4j, Log4j2, Logback, Python)
- Health Checks: Business-level health check interface with composite checks
- Sidecar Management: Phase-ordered, validated sidecar injection framework with domain-specific providers
- Secret/Listener CSI Wiring:
SecretProvisionerandListenerProvisionerdeclare secret-operator / listener-operator CSI volumes and resolve their mount paths - CRD APIs: Common types for authentication, database, listeners, S3
Architecture Documentation (Authoritative Design Source)
IMPORTANT: The
docs/directory contains architecture documents that are the authoritative source of design constraints for this project. All implementations — including the SDK itself and any operators built with it — must follow the design defined in these documents. When code and documentation conflict, the documentation takes precedence. Consult these docs before making design decisions.
Scope of that rule.
docs/architecture.mdis authoritative about design intent: a conflict means the code should change, not that the doc should be quietly relaxed. TheAGENTS.mdfiles (this one and the per-package ones) are the opposite: they describe the API and behavior that exist today, and must be corrected whenever the code changes. Never treat a statement in anAGENTS.mdas a requirement the code has yet to meet — anything aspirational belongs indocs/architecture.mdand must be explicitly labelled as such.
Documentation Structure
| File | Description |
|---|---|
docs/architecture.md |
Core Technical Architecture — design philosophy, layered architecture, core module specifications, design patterns, key problem solutions. This is the primary reference for all SDK design decisions. |
docs/security.md |
Security Architecture — application security (SecretClass, CSI, AutoTLS, Kerberos) and infrastructure security (RBAC, ServiceAccounts, Pod security) |
docs/DOC_CHANGELOG.md |
Changelog tracking all documentation updates |
docs/examples/ |
CRD example YAMLs demonstrating the SDK's data model |
CRD Examples (docs/examples/)
| File | Description |
|---|---|
crd-base-example.yaml |
Base CRD template showing the generic structure all product CRDs follow |
crd-hdfs-example.yaml |
HDFS cluster CRD example (HA with NameNode, JournalNode, DataNode) |
crd-hive-example.yaml |
Hive Metastore CRD example (S3 integration, TLS, Kerberos) |
Key Architectural Principles (from docs/architecture.md)
- Interface-Driven Design (IDD): SDK core relies on interfaces, not concrete implementations. New products implement interfaces without modifying SDK core.
- Desired State Convergence: CR Spec is the desired state; reconciliation loop converges actual state. Bidirectional: also cleans orphaned resources.
- Separation of Common and Specific: SDK handles common logic (resource construction, config merging, webhook validation); products handle specific logic via extension interfaces.
- Type Safety and Idempotency: Go Generics for compile-time safety. All operations are idempotent.
- Strict Merge Strategy: Role/RoleGroup config merging follows defined rules — Deep Merge for maps,
SliceMergeStrategy(Replace by default) for slices, Strategic Merge Patch for PodTemplate. - Layered Architecture: Specific Product Layer → Abstract Interface Layer → Core Component Layer → Tools Layer → API Layer.
Development Environment
- Language: Go 1.25.3
- Dependency Management: Go Modules (
go.mod) - Testing: Ginkgo v2 + Gomega
- Tooling: Uses
Makefileto manage local binaries inbin/
Tool Versions
controller-gen: v0.19.0golangci-lint: v2.12.2kustomize: v5.7.1controller-runtime: v0.23.3k8s.io/api: v0.35.4
Common Commands
Run these from the project root:
| Command | Description |
|---|---|
make generate |
Generate DeepCopy methods via controller-gen |
make manifests |
Generate the test CRDs (config/crd/bases/) from pkg/testutil types |
make verify-generate |
Regenerate and fail if the committed generated files differ (both modules) |
make fmt |
Run go fmt against code |
make vet |
Run go vet against code |
make test |
Run unit tests with coverage (uses envtest for K8s integration); GOTESTFLAGS=-race for the race detector |
make lint |
Run golangci-lint |
make lint-fix |
Run golangci-lint with auto-fix |
make lint-config |
Verify golangci-lint configuration |
Directory Structure
Subdirectories with their own
AGENTS.mdprovide detailed file-level documentation. This section shows the top-level layout only.
operator-go/
├── pkg/ # Core SDK packages (see pkg/AGENTS.md)
│ ├── apis/ # Kubernetes API definitions — CRDs (see pkg/apis/AGENTS.md)
│ ├── builder/ # Fluent resource builders (see pkg/builder/AGENTS.md)
│ ├── common/ # Core interfaces, extensions, errors
│ ├── config/ # Config file generation and override merging (see pkg/config/AGENTS.md)
│ ├── constant/ # Kubedoop paths, labels, domains, restarter annotations, JMX agent
│ ├── listener/ # Listener provisioner (CSI volume registration)
│ ├── productlogging/ # Product logging config generation (Log4j, Log4j2, Logback, Python)
│ ├── reconciler/ # Reconciliation framework (see pkg/reconciler/AGENTS.md)
│ ├── s3/ # S3Connection/S3Bucket resolution, S3A properties, credential wiring
│ ├── security/ # Pod security defaults, SecretProvisioner (secret-operator CSI)
│ ├── sidecar/ # Sidecar injection framework (SidecarManager, SidecarProvider interface)
│ ├── vector/ # Vector sidecar implementation (config generation, discovery, provider)
│ ├── testutil/ # Testing utilities (envtest, mocks, matchers)
│ ├── util/ # K8s utilities, exec utilities
│ └── webhook/ # Webhook infrastructure (defaulter, validator)
├── docs/ # Architecture and design documentation (authoritative design source)
│ ├── architecture.md # Core Technical Architecture
│ ├── security.md # Security Architecture (SecretClass, CSI, RBAC, Pod security)
│ ├── DOC_CHANGELOG.md # Documentation changelog
│ └── examples/ # CRD example YAMLs (base, HDFS, Hive)
├── examples/ # Example operators (see examples/AGENTS.md)
│ └── trino-operator/ # Trino operator example (see examples/trino-operator/AGENTS.md)
├── hack/ # Scripts and boilerplate
└── bin/ # Local binaries (controller-gen, etc.)
Key Concepts
1. ClusterInterface and ClusterResource
All product CRs must implement ClusterInterface (defined in pkg/common/cluster_interface.go):
type ClusterInterface interface {
client.Object // sigs.k8s.io/controller-runtime/pkg/client
GetSpec() *v1alpha1.GenericClusterSpec
GetStatus() *v1alpha1.GenericClusterStatus
}
The embedded client.Object supplies name, namespace, UID, labels, annotations, generation and
GVK, and makes the CR usable directly wherever controller-runtime expects an object (client.Get,
Status().Update, SetControllerReference, event recording). A CR that embeds metav1.TypeMeta
and metav1.ObjectMeta and is registered with the manager's scheme already satisfies that half, so
GetSpec and GetStatus are the only two methods a product writes. There is no SetStatus:
GetStatus returns a pointer into the CR and the framework mutates the generic status through it,
which is why a product's own status fields survive a reconcile cycle.
The reconciler is parameterised by a companion constraint, not by ClusterInterface itself:
type ClusterResource[T ClusterInterface] interface {
ClusterInterface
DeepCopy() T
}
DeepCopy() T is what make generate (controller-gen) already emits for every root API type; the
reconciler needs it because new(T) is unavailable for a pointer type parameter, so it materialises
the object it reads into by copying a prototype. GenericReconciler, GenericReconcilerConfig and
NewGenericReconciler are declared [CR common.ClusterResource[CR]]. Everything else that is
generic over a CR — RoleGroupHandler[CR], ClusterExtension[CR], ExtensionRegistry[CR] — is
constrained by plain common.ClusterInterface.
Hold a CR as ClusterInterface; parameterise over one as ClusterResource[CR].
2. GenericReconciler (Template Method Pattern)
GenericReconciler[CR common.ClusterResource[CR]] provides a fixed reconciliation flow with
customizable extension points. It is built from a GenericReconcilerConfig[CR] through
NewGenericReconciler[CR]:
Reconciliation Flow:
- Fetch CR (NotFound ⇒ done; non-zero
deletionTimestamp⇒ done — see "Deletion" below) - Panic recovery: a recovered panic becomes a returned error plus a
ReconcilePanicWarning event; the status is left untouched - ClusterOperation gate:
reconciliationPausedreturns immediately;stoppedfalls through so every resource is still reconciled with replicas forced to 0 - Ensure the workload ServiceAccount (always; name derived from the CR by
ServiceAccountResourceName) and, whenWorkloadRBACRulesis set, its Role/RoleBinding (§11c) - PreReconcile Extensions (Hook)
- Validate declared dependencies (
GenericReconcilerConfig.Dependencies) 6b.RoleProvider.DeclareRoles— the product's role catalog for this CR, obtained once and validated againstspec.rolesbefore any role is reconciled (§3). A catalog error fails the pass. The catalog is threaded down the call chain rather than stored on the reconciler: one reconciler instance serves every cluster, so a field would leak one CR's declarations into another's pass - For Each Role (best effort — see below):
- Role PreReconcile Extensions
- For Each RoleGroup:
- RoleGroup PreReconcile Extensions
- Build RoleGroupBuildContext: fold the typed config (
FoldCommonConfig, §4b) → resolve the image and the Vector gates → derive (RoleGroupResolver, §10) → merge the overrides - Delegate to RoleGroupHandler.BuildResources()
- Apply Resources (CM -> HeadlessSvc -> Service -> Extras -> [sidecar Validate] -> STS -> per-group PDB -> MetricsSvc)
- RoleGroup PostReconcile Extensions
- Role-level PodDisruptionBudget
- Role PostReconcile Extensions
- Cleanup Orphaned Resources
- Update Health Status
- PostReconcile Extensions
- Final Status Update, then requeue
Each "Apply" is create-OR-UPDATE (issue #526): when the resource already exists, the live object is updated to the handler-built desired state every reconcile — labels are replaced wholesale, annotations are merged (foreign annotations survive, at both the object level and inside the StatefulSet's pod template), and spec/data is copied per kind while preserving Kubernetes immutable/allocated fields (StatefulSet selector/serviceName/volumeClaimTemplates/podManagementPolicy; Service clusterIP(s)/ipFamilies and allocated NodePorts). Arbitrary-GVK extras get a generic top-level field copy. See copyDesiredState in pkg/reconciler/apply.go. Changing an immutable field for an existing cluster requires a manual delete/recreate migration — and the framework now says so: when a handler's desired value for a preserved field differs from the live one, applyResource emits an ImmutableFieldIgnored Warning event on the CR naming the resource and the field paths. Preserving those fields silently is what let a storage resize be accepted, reported as ReconcileComplete=True, and never applied. Only a field the handler actually set is reported (an unset field is declining to have an opinion, not a change request), and among the Service's preserved fields only clusterIP is — the others are API-server allocations, so a difference there would be noise on every reconcile. The event is emitted before the write's own error is returned, so a rejected Update still says which of the user's changes the framework had already dropped.
volumeClaimTemplates is the one preserved field the pod template depends on, so the mounts follow
what was actually preserved. Preserving the claim templates alone left an incoherent StatefulSet in
both directions. Adding config.resources.storage to a live role group produced a template
mounting a claim the framework had just declined to create — volumeMounts[0].name: Not found: "data", a field the user never wrote. On Kubernetes 1.34+ the API server rejects that Update
outright, leaving the role group Degraded on every pass with no recovery short of deleting the
StatefulSet by hand; older servers accept it and reject every pod the StatefulSet controller then
creates, so the workload never progresses and the error never reaches the cluster's status at all. Removing it was worse, because it was accepted: the claim template stayed,
the mount did not, and the pods rolled into a product writing to the container's writable layer while
its bound PVCs sat mounted nowhere — no event, no condition, no log line. copyStatefulSetState
therefore drops a mount whose claim was not created (unless an ordinary volume of that name backs it)
and restores every mount a preserved claim had from the live template, so the path is read
rather than invented; the restore is keyed on the mount path, since Kubernetes lets one volume be
mounted several times and a path the desired template already uses wins. A rename hits both branches
and lands on the preserved claim. This makes the transition
converge and stay coherent — it does not make it happen, and spec.volumeClaimTemplates is still
reported as ignored. That report is also the one place an empty desired value counts: everywhere
else an unset field is the handler declining to have an opinion, but an empty volumeClaimTemplates
is the handler stating this role group has no storage, which is exactly the direction that used to be
applied in silence.
Inside a claim template, a value only the server filled in is not a change request. The live
template comes back carrying spec.volumeMode: Filesystem and a status block a handler-built one
has no way to state, so a whole-slice comparison was true on every pass: each role group with a
data PVC emitted ImmutableFieldIgnored forever while its StatefulSet's generation stayed at 1 and
no pod rolled (#627). That is not merely noisy — it is the same warning a genuine resize produces, so
the event stopped distinguishing "your resize was dropped" from the background. Those two fields are
therefore compared only when the handler states one; everything else — capacity, access modes, the
claim's name, storageClassName — still counts, and storageClassName in particular is not
defaulted into the template by the API server, so changing it is still reported.
A configOverrides change does not roll the pods by itself — the platform restarter does.
Editing configOverrides makes the framework rewrite the role group ConfigMap, and stop there: the
pod template is byte-identical, so the StatefulSet controller has no reason to roll anything, and
none of these products re-read their configuration files at runtime.
Delivering that change to the running processes is commons-operator's restarter, not this SDK.
Label the workload restarter.kubedoop.dev/enable=true (constant.LabelRestarterEnable /
LabelRestarterEnableValue) and, whenever a ConfigMap or Secret the pod references — as a volume
or through an env var's valueFrom — changes, the restarter writes
configmap.restarter.kubedoop.dev/<name> (or secret.restarter.kubedoop.dev/<name>) into the
workload's pod template; the StatefulSet controller then rolls the pods. The annotation's value is
<uid>/<resourceVersion>, so enabling the label always costs one rollout: the first pass stamps
a template that had no annotation. The SDK writes neither annotation — those prefixes exist in
pkg/constant/restarter.go to document the restarter's half of the contract, not for the framework
to emit. The same component also restarts pods whose secret-operator TLS/Kerberos secrets have
passed the expiry recorded in restarter.kubedoop.dev/expires-at.<...>.
The label goes on StatefulSet.metadata.labels, which is where the restarter's watch predicate
and its client.MatchingLabels list both read it — a pod-template label does not enable anything,
so podOverrides is not a way in. The framework's own channel is the cluster CR's labels: the
reconciler passes a writable clone of cr.GetLabels() as RoleGroupBuildContext.ClusterLabels, and
BaseRoleGroupHandler merges them into every built resource's metadata (and pod template), so
kubectl label <cluster-cr> restarter.kubedoop.dev/enable=true reaches the StatefulSet metadata the
restarter watches. Opting in is a deployment decision by whoever runs the cluster, not something
the operator's author hardcodes.
Three keys are withheld from that channel — metrics.kubedoop.dev/service, pdb.kubedoop.dev/role
and pdb.kubedoop.dev/role-group. They are the framework's slot markers: a reclaim selects an
object for deletion by their presence or value, and unlike the app.kubernetes.io/* set nothing
overwrites them afterwards, so a CR carrying one would stamp it on every built resource and make
each answer to a reclaim aimed at the slot. The filter is an enumerated set rather than a
kubedoop.dev prefix rule precisely because restarter.kubedoop.dev/enable shows that domain is
shared with the platform; every other CR label propagates unchanged.
Caveat (upstream bug). commons-operator's
getRefConfigMapRefsreturns after the first ConfigMap volume it finds, so a pod mounting several ConfigMaps only ever gets one of them watched — see zncdatadev/commons-operator#298. The secret path (getRefSecretRefs) is correct.
The framework already satisfies the restarter's precondition: the role group ConfigMap is mounted as
the config volume. What it does not do is set the label, and without it a configOverrides
change simply does not roll.
The apply path preserves the restarter's stamp. The pod template's annotations are MERGED, not
replaced — the same rule the object's own annotations follow, for the same reason: another controller
writes there. copyStatefulSetState assigns live.Spec = desired.Spec wholesale so that new mutable
fields converge by default, and the pod template lives inside that spec, so before this rule a
handler that never builds configmap.restarter.kubedoop.dev/<name> silently removed it on the next
reconcile. That Update woke the restarter — its predicate matches the label on every Update, not only
on Create — which re-stamped, which woke the reconciler through its own Owns(&appsv1.StatefulSet{})
watch. Neither side is failing, so the workqueue Forgets each pass and nothing backs off: the pods
rolled for as long as the label was set. Pod-template labels are still replaced wholesale, because
they must match the StatefulSet's immutable .spec.selector.
This applies to configOverrides alone. envOverrides and cliOverrides reach the container as
env vars and args through MergedConfig (StatefulSetBuilder.WithConfig), and podOverrides
patches the template directly — all three change the pod template, so they roll natively with no
restarter involved.
Role iteration is best-effort. A failing role does not stop the others, and a failing role group does not stop its siblings, the role-level PDB, or the role's PostReconcile hook. Steps 8-10 (orphan cleanup, health, cluster PostReconcile) run regardless. Roles and role groups are independent workloads: aborting at the first failure meant one unparsable value on the alphabetically-first role indefinitely blocked the deletion of an unrelated role group, the health of every other role, and the discovery ConfigMap a product publishes from PostReconcile. The per-role errors are combined with errors.Join and returned once, so the cluster still goes Degraded and the workqueue still backs off. Iteration stays sorted so that aggregated message is byte-stable across cycles — an unstable message would defeat the no-op guard in updateStatus and make the controller reschedule itself forever. The single exception is a 429: a *RateLimitError aborts the pass immediately, because pushing the remaining roles through would only deepen the backlog.
Requeue cadence. A successful reconcile returns ctrl.Result{RequeueAfter: HealthCheckInterval} (DefaultHealthCheckInterval = 120s; a negative value disables the periodic wakeup), or the cleaner's earliest pending wakeup when that is sooner — a remaining gray-delete deadline, or the drain poll interval of an orphan deletion in flight. Watches only cover the kinds the framework owns, so anything that changes without producing an event — a product ServiceHealthCheck probe, a grace period running out, a StatefulSet finishing its drain — depends on this timer.
Orphan cleanup discovers its work from the live cluster, not only from status.roleGroups. The cleaner unions two inventories: the role group ConfigMaps and StatefulSets this CR controller-owns that carry the framework's labels (instance + managed-by + component + role-group) and whose name is exactly what RoleGroupResourceName produces for those labels, plus the status.roleGroups ledger. The ledger alone is a record the operator must have successfully written, so losing it — a process death between applying a role group's resources and updating the CR, a backup tool restoring the CR without its status subresource, a kubectl replace — used to make those resources invisible to the cleaner permanently, holding their PVCs and pods until a human noticed. The name check is what keeps the live half safe: a discovery ConfigMap carries the same instance/managed-by pair and owner reference, and a product's ExtraResources may carry the handler's entire label set. An empty owner UID disables live discovery, as it does the role-PDB reclaim.
Orphan cleanup is a multi-pass state machine. A role group removed from the spec is retired over several reconciles: scale the StatefulSet to zero (under RetryOnConflict), wait for the controller's ordered drain (.status.replicas reaching 0), then delete PDB → [PVCs] → StatefulSet → [product extras] → ConfigMap → Service → headless → metrics, each step confirmed absent before the next is issued. The PVC step is opt-in (operator.zncdata.dev/delete-pvcs) and sits after the drain on purpose: deleting a role group is undoable right up until its data goes, so nothing irreversible happens while the pods are still running, and re-adding a group mid-teardown costs a restart rather than the data. It goes before the StatefulSet because the cleaner finds the PVCs through its selector — the other order would strand them — and the drain-timeout path falls through to it so a stuck pod cannot silently leak the volumes. A group's status entry is pruned only after a real deletion; a failure is isolated to its own group and the others still progress; the state machine's progress annotations (orphan.zncdata.dev/pending-deletion, orphan.zncdata.dev/drain-started) are reset on every role group that IS in the spec, so a group that was orphaned and then re-added starts its next teardown from scratch instead of inheriting the previous one's timestamps; a 429 becomes a *reconciler.RateLimitError that aborts the pass and backs off instead of marking the cluster Degraded. The cleaner also reclaims the role-level PDB of a role deleted from the spec outright, found by the pdb.kubedoop.dev/role label (reconciler.LabelRolePodDisruptionBudget) rather than by derived name.
Status writes are conditional. updateStatus skips the write entirely when the whole CR is apiequality.Semantic.DeepEqual to the object read at the start of the cycle — comparing the whole object, not just the embedded generic status, so a product's own status fields count too. Without that guard the controller's watch on its own CR would turn every reconcile into another reconcile. The write itself goes out from the in-memory object (a re-fetch would discard product-specific status fields); a 409 refreshes only the resourceVersion, preferring GenericReconcilerConfig.APIReader because the informer cache has not seen the competing write, and a NotFound is treated as success.
Deletion uses owner-reference garbage collection, not finalizers. The SDK registers no finalizer anywhere, so deleting a cluster CR runs no SDK teardown code. Everything the framework applies carries a controller owner reference and is reclaimed by Kubernetes GC. Reconcile detects deletion on two paths: background propagation (the kubectl delete default) removes the CR immediately and hits the IsNotFound branch, while foreground propagation (--cascade=foreground) and any product-registered finalizer leave the CR readable with a deletionTimestamp — checked right after the fetch, before the ClusterOperation gate and any mutating step. The second check is load-bearing: without it the pass re-creates every owned resource with BlockOwnerDeletion: true, which foreground deletion can never get past, producing a permanently Terminating CR in an un-backed-off recreate loop. The operator.zncdata.dev/delete-pvcs annotation (reconciler.AnnotationDeletePVCs) therefore only affects the orphan path — PVCs of a StatefulSet whose role group was removed or renamed in the spec. On cluster deletion those PVCs remain, because the SDK sets no persistentVolumeClaimRetentionPolicy and StatefulSet-managed PVCs carry no owner reference. Products with state outside owner-reference GC must clean it up themselves.
Validation failures are loud. Registered, enabled sidecar providers are validated before the StatefulSet is applied; a failure aborts the role group with *reconciler.ValidationError (NewValidationError / IsValidationError). A podOverrides layer that fails to decode is recorded on config.MergedConfig.PodOverrideErrors and re-emitted as a PodOverrideIgnored Warning event rather than being dropped silently. A fixed RoleGroupResources slot built under a name or namespace the framework does not own fails the same way, before anything is applied (§3).
A podOverrides volumeMount at a framework-owned mountPath replaces it — and that is now a build failure. Strategic merge patch keys volumeMounts by mountPath, not by name, so a mount declared at a path the framework already owns (/kubedoop/config, a CSI secret/listener path, the shared log volume) does not sit alongside the framework's — it rewrites the framework's entry to point at a different volume. When the override also declares that volume the resulting pod spec is completely valid, the API server accepts it, and the pods come up with the generated ConfigMap mounted nowhere: the product reads an empty config directory and crash-loops or silently runs on its built-in defaults. When it declares only the mount, the API server rejects the StatefulSet naming spec.template.spec.containers[0].volumeMounts[0].name — a field the user never wrote, with no mention of podOverrides. StatefulSetBuilder records both on PodOverrideViolations() and BaseRoleGroupHandler turns them into a *reconciler.ValidationError naming the mountPath, the displaced volume and podOverrides. Mounting at a new path is unaffected, which is what users normally mean.
3. RoleProvider, RoleDeclaration and RoleGroupHandler
A product declares its roles as DATA, once per reconcile, with the CR in hand.
RoleProvider is that seam:
type RoleProvider[CR common.ClusterInterface] interface {
DeclareRoles(ctx context.Context, c client.Client, cr CR) (RoleCatalog, error)
}
type RoleCatalog map[string]RoleDeclaration
Set it as GenericReconcilerConfig.RoleProvider. It is called once per pass, after the
dependency check and before any role is reconciled — so a product resolving a cluster-wide fact (an
authentication class that decides whether every role speaks TLS, an S3 connection) pays for that
lookup once rather than N times for N role groups. Returning a *common.RequeueAfterError reports
"not ready yet" without the cluster going Degraded (§5); any other error fails the pass.
RoleProviderFunc adapts a plain function.
RoleDeclaration is everything the product knows about one role:
return reconciler.RoleCatalog{
"coordinator": {
MainContainerName: "trino",
ContainerPorts: []corev1.ContainerPort{{Name: "http", ContainerPort: port}},
ServicePorts: []corev1.ServicePort{{Name: "http", Port: port}},
Command: []string{"/bin/bash", "-c", "launcher run"},
LogProducers: []productlogging.ContainerLogging{{Container: "trino", Framework: productlogging.LoggingFrameworkLogback}},
ConfigDefaults: &commonsv1alpha1.RoleGroupConfigSpec{Affinity: aff},
},
"worker": { /* … */ },
}
Full field set: Image, ContainerPorts, ServicePorts, MainContainerName, Command,
Lifecycle, ReadinessProbe/LivenessProbe/StartupProbe, DataVolume, ListenerClass,
PublishNotReadyAddresses, LogProducers, OwnsVectorConfig, LogVolumeSize, Env,
ConfigDefaults, Optional.
Almost none of it has a user layer above it, and that is the point. A role's ports, its primary container's name and command, whether it has a data PVC, which of its containers produce logs — the CRD offers the user no way to state any of these, so nothing merges and nothing is beaten. This replaced seven role-keyed maps on the handler, their handler-global fallbacks, and four per-call fields on the build context that outranked them: the declaration used to be shredded across three objects each with its own precedence rule, and the rule for one of them dropped sibling fields, which is the defect (#631) that started this.
Two fields are the exceptions, in opposite directions:
ConfigDefaultsis folded BENEATH the CR's role and role group levels byFoldCommonConfig(§4b). Anything the user states anywhere wins. An anti-affinity default belongs here and could not live on the old handler at all, because its selector names the cluster while the handler is a process-wide singleton shared by every cluster the operator serves.Envis emitted beneath the merged overrides. Kubernetes resolves a duplicate env name to the last entry, so a user'senvOverridesof the same name wins. It exists because the merged channel ismap[string]stringand cannot carry avalueFromat all; a value the product computes belongs inContribution.EnvVars(§10) instead.
Being produced per reconcile is what retires the per-call escape hatches. A port that moves because the CR enabled TLS is computed here, from this CR, rather than assigned into process-wide handler state that the next cluster inherits.
The catalog is validated once per pass, asymmetrically. A role the CR declares that the catalog
does not is a hard error naming the typo and listing what it could have meant — before, an
unrecognised name silently produced a workload with no ports, no image and no Service while the
reconcile reported success. A role the catalog declares that the CR does not use is a Warning
event (UnusedRoleDeclaration), since a product may support more roles than a given cluster runs;
marking it Optional: true silences even that. reconciler.ValidateCatalog is the exported check.
Images are resolved once per role, by the framework. GenericReconcilerConfig.ImageResolution
carries the product name and the operator's defaults:
ImageResolution: reconciler.ImageResolution{
ProductName: "trino", // app.kubernetes.io/name AND the repo path segment
Defaults: commonsv1alpha1.ImageSpec{
Repo: "quay.io/zncdatadev",
ProductVersion: "476",
KubedoopVersion: version.BuildVersion, // the operator's own build version
},
},
The layers fold per field — CR spec.image first, then RoleDeclaration.Image, then Defaults —
so a CR stating only productVersion still yields a valid …:476-kubedoop0.2.0 reference. The
result reaches the handler as RoleGroupBuildContext.ResolvedImage (Reference, PullPolicy,
PullSecretName, ProductVersion). RoleDeclaration.Image sits under the CR deliberately: a
product pinning one role to a different image must not silently beat a user who pinned
spec.image.custom to an air-gapped mirror.
Defaults is read every reconcile, which is what a webhook cannot do: webhook defaults are
persisted at admission and never recomputed, freezing kubedoopVersion at whatever operator version
first admitted the CR (§10 and docs/architecture.md §2.6).
An unresolvable spec.image fails the role group, naming the missing field, instead of silently
falling back to a static image and running a version nobody asked for. With ProductName empty the
framework resolves nothing from the CR beyond spec.image.custom — the shape a product uses when it
resolves images itself — and that path never errors. app.kubernetes.io/version follows the
resolved version, so it is present whenever the version came from Defaults too.
The assembled tag is validated, and an unusable one fails the role group. productVersion and
kubedoopVersion both land in the image tag, whose grammar is [A-Za-z0-9_][A-Za-z0-9._-]{0,127},
and nothing downstream would catch a value that breaks it: the API server does not validate
container.image at all, so an unparsable reference is accepted, stored, and surfaces only as
InvalidImageName on a pod while the reconcile reports success. The case this exists for is
KubedoopVersion: version.BuildVersion with the scaffold's dev default of "N/A" — the / makes
…:476-kubedoopN/A unparsable, so every development build on the structured image path produced
pods that could never start. The error names the offending field and says the value may have come
from the operator's own defaults rather than from the CR. spec.image.custom is deliberately not
validated: it is the user's verbatim reference, so a wrong one is their own visible mistake.
spec.image.pullSecretName names a docker-registry Secret added to every pod's
imagePullSecrets. It folds per field like the rest and is resolved independently of the
image: a pull secret is a property of where the image lives, so it applies on all three paths —
the assembled reference, custom, and a product that resolves its own images with ProductName
empty. Ten product CRDs already declared this field, and migrating to the commons ImageSpec used
to delete the behaviour silently — the CRD still accepted the value and no pod ever carried the
entry, so a private-registry install failed with ImagePullBackOff and nothing naming the cause.
One name rather than a list, matching those CRDs; it is applied before podOverrides, and strategic
merge patch keys imagePullSecrets by name, so an override adds a credential rather than
replacing this one.
3b. BaseRoleGroupHandler
RoleGroupHandler is what turns the declaration and the folded config into objects:
type RoleGroupHandler[CR common.ClusterInterface] interface {
BuildResources(ctx context.Context, k8sClient client.Client, cr CR, buildCtx *RoleGroupBuildContext) (*RoleGroupResources, error)
}
BaseRoleGroupHandler.BuildResources returns a ConfigMap, a headless Service, a StatefulSet, and a
client-facing Service when the role declares service ports. Construct it with the scheme alone —
everything else now arrives per reconcile:
handler := reconciler.NewBaseRoleGroupHandler[*v1alpha1.TrinoCluster](scheme)
The handler carries only reconcile-invariant settings a role cannot differ on:
ConfigGenerator, Scheme, ConfigMountPath, LabelDomain, the sidecar manager and the security
contexts. Everything role-shaped moved to RoleDeclaration; everything CR-shaped moved to
RoleProvider/RoleGroupResolver. That is what makes one handler instance safe for every cluster
the operator reconciles — the older idiom of assigning h.Image or calling h.SetRoleContainerPorts
inside BuildResources wrote per-cluster values into process-wide state, which races above
MaxConcurrentReconciles: 1 and leaks even at 1 (spark-k8s-operator shipped exactly that).
It does not return a PDB: the framework's PDB comes from roleConfig.podDisruptionBudget and is a
role-level resource built by BuildRolePodDisruptionBudget and applied once per role by the
reconciler (RoleGroupResources.PodDisruptionBudget remains an escape hatch for an extra per-group
PDB). BuildRolePodDisruptionBudget takes a single *RoleBuildContext — the role-scoped analogue
of RoleGroupBuildContext, carrying ClusterName, ClusterNamespace, ClusterLabels,
ClusterSpec, RoleName, RoleSpec, ProductName and ProductVersion — so a later role-level
input needs no new signature.
When building the StatefulSet, BaseRoleGroupHandler consumes the role group's folded config
(commons RoleGroupConfigSpec): resources (requests/limits, plus the data PVC when the role
declares a DataVolume; storage.storageClass is a *string because Kubernetes reads
storageClassName: "" as "bind only a pre-provisioned PV, never dynamically provision one" — a role
group that set it to "" to mean "inherit the role's" would get a PVC that stays Pending
forever), affinity (see below), and gracefulShutdownTimeout (a Go duration mapped to
terminationGracePeriodSeconds — unparsable or non-positive values fail the build). All of these
are applied before podOverrides, so user pod overrides keep precedence. The framework sets
affinity only when the config provides one, so products that post-process the built StatefulSet with
if podSpec.Affinity == nil {...} default guards remain correct.
config.affinity is decoded STRICTLY, and is replaced wholesale. The CRD carries it as a
schema-free RawExtension (type: object + x-kubernetes-preserve-unknown-fields), so the API
server neither validates nor prunes it. reconciler.DecodeAffinity therefore decodes with
DisallowUnknownFields and an unknown field fails the build, naming it. Before that, nodeAffinty
(one letter short) passed admission, decoded into an empty corev1.Affinity, and the pods were
scheduled anywhere — with no event, no log line and no status change, even though affinity is the
scheduling contract for these products (rack awareness, spreading a quorum, colocating a worker
with its data). The trade-off is deliberate: a field from a newer Kubernetes API than the SDK is
built against is now rejected rather than ignored, which is the honest answer, since the framework
cannot honor a field it does not know.
config.affinity is replaced wholesale by any layer that states one — the rule Kubernetes uses
for PodSpec.affinity — and an empty value clears. What a replacement discarded is reported as an
AffinityOverridden Warning event. See §4b for why the Kubernetes rule is kept and what it costs.
The affinity helpers are composable terms, not one canned policy.
PreferredAffinityTerm(weight, topologyKey, selector) builds one weighted term;
ClusterSelectorLabels(cluster) and RoleSelectorLabels(cluster, role) are the selectors, built
from the framework's own identity labels (app.kubernetes.io/instance +
app.kubernetes.io/component) — the part worth centralising, since a downstream operator re-typing
those keys as string literals gets a selector matching nothing the moment they change, and a
preferred term matching no pod is not an error. EncodeAffinity is DecodeAffinity's inverse, so a
product supplies a default with the typed Kubernetes API rather than a JSON literal:
aff, err := reconciler.EncodeAffinity(&corev1.Affinity{
// spread this role's pods across nodes
PodAntiAffinity: &corev1.PodAntiAffinity{PreferredDuringSchedulingIgnoredDuringExecution: []corev1.WeightedPodAffinityTerm{
reconciler.PreferredAffinityTerm(70, reconciler.TopologyKeyHostname,
reconciler.RoleSelectorLabels(cr.GetName(), "datanode")),
}},
// …while keeping the cluster together
PodAffinity: &corev1.PodAffinity{PreferredDuringSchedulingIgnoredDuringExecution: []corev1.WeightedPodAffinityTerm{
reconciler.PreferredAffinityTerm(20, reconciler.TopologyKeyHostname,
reconciler.ClusterSelectorLabels(cr.GetName())),
}},
})
if err != nil { return nil, err }
decl.ConfigDefaults = &commonsv1alpha1.RoleGroupConfigSpec{Affinity: aff}
This replaces a single-shot DefaultAntiAffinity helper that emitted exactly one anti-affinity
term — which is why hdfs-operator, the product with the most roles, could not use it at all: its
default is composite (a cluster-level pod affinity at weight 20 beside a role-level
anti-affinity at weight 70). Both operators that needed a composite hand-wrote it, and both got it
wrong — one commented the merge out entirely, the other applied its default unconditionally and
silently discarded the user's own config.affinity.
There is deliberately no Required constructor: a required spread turns a too-small cluster into
pods that never schedule (three nodes, five replicas, two Pending forever), with no way to say
"spread as far as you can". A product that genuinely requires it writes the corev1 term itself.
NormalizeAffinity is exported so a hand-assembled affinity produces the same bytes the fold does —
without it a member cleared with {} reaches the pod template as an empty struct, which differs
from absent in the serialized spec and shows up as a diff on every reconcile.
logging IS a supported config default now, folded through productlogging.MergeLoggingSpec
like every other field in the block. It used to be rejected with a *ValidationError, and the
rejection was correct for the code as it stood: the framework merged logging in the reconciler from
the CR's two levels only, and had already decided Vector enablement and rendered the config file
before a default was ever read — so one set here would have applied to neither. Moving the fold
ahead of both consumers is what made the field honourable, and the reconciler now reads Vector
enablement off the folded value. This is issue #631's item 2: the old godoc advertised logging
while the code hard-rejected it, and the fix was to make the code match the doc rather than the
other way round.
A data PVC is per role, declared not configured. RoleDeclaration.DataVolume{Name, MountPath}
opts the role in; nil means the role has none. That is a structural property of the role rather than
a consequence of what the user wrote — a product with one stateful and one stateless role could
previously only choose between "every role" and "none", so one migration deleted a
volumeClaimTemplate its predecessor rendered. Name defaults to builder.DefaultDataVolumeName
("data"); the size comes from the folded config.resources.storage, so the user still sets it.
RoleGroupHandlerFuncs is a function adapter for simple handlers that don't need a full struct.
The framework owns the NAME of every fixed slot; the handler owns its content. ConfigMap,
Service, StatefulSet and PodDisruptionBudget must be named RoleGroupBuildContext.ResourceName,
HeadlessService and MetricsService that name plus -headless / -metrics, and all six must sit
in the cluster's namespace. Both paths that remove a slot address it by that derived name — the
in-spec reclaim and the orphan teardown — so a slot filled under another name used to be applied,
owner-referenced and then reclaimed by nothing, surviving until the cluster CR itself was deleted (a
metrics Service left as a Prometheus target with no endpoints). It is now a *ValidationError
raised before any resource is applied, so a rejected declaration leaves nothing half-converged.
builder.MetricsServiceBuilder therefore exposes no name override, and ExtraResources is the
supported route for an object whose name the product chooses.
Besides the fixed fields (ConfigMap, Services, StatefulSet, PDB, MetricsService), RoleGroupResources.ExtraResources []client.Object lets products ship arbitrary per-role-group resources (e.g. a listeners.kubedoop.dev Listener CR) through the framework's apply path: same controller owner reference, applied BEFORE the StatefulSet because extras are typically pod-scheduling prerequisites. Extras of a removed role group are reclaimed too, provided the product registers their kinds through SetupWithManagerOptions.ExtraOwns and labels them with the role group's labels; the teardown deletes them right after the StatefulSet, mirroring the apply order. Unregistered or unlabelled extras keep the old behaviour and wait for owner-reference GC on cluster deletion (see §13 and the field's doc comment).
The slice type stays open; the entries are validated. Arbitrary GVKs are the point, so nothing
narrower than client.Object can be the type — but three properties are checked in the same
pre-apply gate the fixed slots use, each failing the role group with a *ValidationError naming the
index: every entry has a name (CreateOrUpdate addresses the object by name every reconcile, so
generateName would create a new object per pass instead of converging one), every entry sits in the
cluster's namespace (which also rejects a cluster-scoped object — Kubernetes honours no owner
reference from a namespaced CR to one, so the framework has no lifecycle to give it), and no two
entries — nor an entry and a fixed slot — address the same object. The first two only move a
failure that already existed inside applyResource earlier, to where nothing is half-applied yet.
The third catches what failed nowhere at all: two writers for one object in one pass mean the later
apply silently discards the earlier, and if their desired states differ the object is rewritten every
reconcile, each write waking the framework's own watch with nothing to back the loop off.
A batchv1.Job in that slice is create-once, and that is a typed rule rather than the generic
copy. The fallback assigns spec wholesale, and the API server generates spec.selector and
injects four UID-derived labels into spec.template at creation — neither of which a handler-built
desired object can carry. So the second reconcile of an unchanged Job was rejected
(spec.selector: Required value, spec.template: field is immutable) and the role group went
permanently Degraded quoting a field the user never wrote. copyJobState preserves the live
selector, template, completions, completionMode and manualSelector, and lets parallelism,
suspend, backoffLimit, activeDeadlineSeconds and ttlSecondsAfterFinished converge — the knobs
batch deliberately left mutable. Create-once is also the only semantics a Job has: its work is a
side effect that already happened, so a product needing it re-run changes the name, and a
differing template is reported through ImmutableFieldIgnored like any other preserved field.
4. RoleGroupBuildContext
Role and role group configuration reaches a handler through one struct, built per role group by the
reconciler and passed to BuildResources. There is no role-level interface a product implements:
pkg/common/role_interface.go (RoleInterface, RoleInfo, RoleGroupInfo) does not exist — the
reconciler iterates spec.Roles directly.
RoleGroupBuildContext carries ClusterName, ClusterNamespace, ClusterLabels, ClusterSpec,
RoleName, RoleSpec, RoleGroupName, RoleGroupSpec, MergedConfig (the folded
product-contribution/role/role-group overrides), ResourceName ({cluster}-{role}-{group},
truncated with a hash suffix by RoleGroupResourceName), ServiceAccountName (the SA the reconciler
derived and ensured — ServiceAccountResourceName(kind, cluster), never configured and never empty),
SidecarManager, VolumeProviders (see §16) and VectorAggregatorAddress.
Three fields are the framework's already-settled answers, written before BuildResources runs:
| field | what it is |
|---|---|
Declaration |
the RoleDeclaration this role group's product returned (§3) |
ResolvedImage |
Reference, PullPolicy, PullSecretName, ProductVersion — resolved once so the container and the sidecars cannot be told different things |
ProductName |
ImageResolution.ProductName, the app.kubernetes.io/name value |
Assigning Declaration from inside BuildResources is a half-honoured change and therefore
worse than one wholly ignored: the image, the config fold and the Vector gates are already settled
from it, while ports, container name, command, probes, data volume, listener class and log producers
are read during the build and would take effect. Declare the role in RoleProvider, where every
consumer sees the same answer.
The per-call escape hatches are gone. Image, ImagePullPolicy, ContainerPorts,
ServicePorts and MainContainerCustomizer on the build context no longer exist. Each was a
channel that outranked the handler's own state, and together they made the role's definition a
three-object precedence puzzle. A value that depends on the CR is computed in DeclareRoles, which
receives the CR:
func (h *MyHandler) DeclareRoles(ctx context.Context, c client.Client, cr *MyCluster) (reconciler.RoleCatalog, error) {
port := int32(8080)
if cr.Spec.Tls != nil { port = 8443 }
return reconciler.RoleCatalog{"server": {
ContainerPorts: []corev1.ContainerPort{{Name: "http", ContainerPort: port}},
Command: []string{"/bin/zkServer.sh"},
}}, nil
}
MainContainerCustomizer in particular is not replaced by a hook. Command, Lifecycle and the
three probes are declaration fields, applied to the primary container by identity — nobody indexes
Containers[0], an assumption a sidecar provider inserting a container earlier quietly breaks — and
applied before podOverrides are strategic-merged, which is why they cannot be a post-build
patch: a product editing the returned StatefulSet lands after the merge and silently beats the
user. Changing the image this way is impossible by construction now, rather than rejected by a
guard: the image is resolved once, propagated to the sidecars, and published read-only as
ResolvedImage.
listener.ServiceTypeFor is the shared class→type mapping restored from v0.12.6:
cluster-internal → ClusterIP, external-unstable → NodePort, external-stable →
LoadBalancer, anything else → ClusterIP. It lives in pkg/listener rather than pkg/builder
because pkg/listener already imports pkg/builder. A role's class is
RoleDeclaration.ListenerClass; a product that exposes the class as a user-settable field in its
own config block must set it from Contribution.ListenerClass (§10) instead, because such a value
arrives from the config fold per role group, after the declaration is fixed.
One handler instance serves every cluster — it is built once in main.go — which is why the
handler now carries nothing role- or CR-shaped at all (§3b).
sidecar.SidecarManager.CloneForBuild covers the framework's own instance of the same hazard, a
handler-registered manager whose configs SetProductImage writes into. See
docs/architecture.md §4.1.4.
Role and role group names are constrained by the CRD, not by convention. The keys of
spec.roles and spec.roles.<role>.roleGroups must be lowercase RFC 1123 labels — a CEL
x-kubernetes-validations rule on GenericClusterSpec.Roles and RoleSpec.RoleGroups rejects
anything else at kubectl apply. They are not free-form: each becomes a segment of
<cluster>-<role>-<group> and the value of an app.kubernetes.io/* label, so Coordinator,
my_role and a.b all yield resource names the API server refuses. Without the rule that refusal
surfaced mid-reconcile as a permanently Degraded role quoting a metadata.name the user never
wrote. Both maps also carry a maxProperties bound (64 roles, 256 role groups): it exists because
the CEL cost estimator has no other handle on a map — without it the API server rejects the rule at
CRD creation time — and is set far above any real deployment.
4b. Config Folding — Two Owners, One Rule Set
The config block is the one place a product extends a framework type. A product CRD embeds the
commons struct inline and adds its own fields beside it:
type ConfigSpec struct {
*commonsv1alpha1.RoleGroupConfigSpec `json:",inline"` // resources, affinity, logging, …
// A POINTER, not a bare string: the fold's "was this stated?" test is the zero value, so a
// bare string cannot tell `myProductSetting: ""` from an absent field.
MyProductSetting *string `json:"myProductSetting,omitempty"`
// A composite must say how it folds. `atomic` accepts wholesale replacement; the alternative
// is to flatten it into scalar pointers. Untagged, ValidateProductConfigType REFUSES it.
Tls *TlsSpec `json:"tls,omitempty" kubedoop:"atomic"`
}
Both annotations are load-bearing rather than stylistic — ValidateProductConfigType[ConfigSpec]()
rejects the shape without them, and FoldProductConfig calls that validator first, so the call below
would return an error rather than a config.
That is the sanctioned extension mechanism, so folding it has two owners, and each folds its own half with the framework's machinery:
common, err := reconciler.FoldCommonConfig(decl.ConfigDefaults, roleCfg.RoleGroupConfigSpec, groupCfg.RoleGroupConfigSpec)
product, err := reconciler.FoldProductConfig(productDefaults, roleCfg, groupCfg) // *ConfigSpec
FoldCommonConfig is what the framework calls on every pass, over three layers in this order:
RoleDeclaration.ConfigDefaults → the CR's role config → the CR's role group config. The result
is RoleGroupBuildContext.EffectiveConfig(). FoldProductConfig is generic and folds the product's
own fields the same way, skipping the embedded *RoleGroupConfigSpec so the two halves cannot fight
over it; a product calls it wherever it reads its own config.
Four concepts were conflated before this, and separating them is the whole fix (#631):
| where it sits | who states it | |
|---|---|---|
| Declaration | no user layer at all | product (RoleDeclaration) |
| Default | beneath the user's two levels | product (ConfigDefaults) |
| Derivation | computed from the folded result | product (RoleGroupResolver, §10) |
| Constraint | above the user — deliberately not implemented | — |
Constraint is absent on purpose. A framework that can overrule what a user wrote in their own CR
needs a way to tell them, and every candidate (silently clamping, failing the role group, a
Warning) is worse than letting the value through and letting the product validate it in a webhook,
where the user sees the rejection at kubectl apply.
Per-field rules. resources folds per leaf, so a default cpu.min survives a user who set
only cpu.max — a struct-level nil check would have discarded it. Scalars fold on presence.
affinity is replaced wholesale, and the replacement is reported:
roles:
server:
config:
affinity: # role: spread the quorum
podAntiAffinity: {…}
roleGroups:
default:
config:
affinity:
nodeAffinity: {…} # group: pin to an instance type
The group wins the whole field: the role's podAntiAffinity is gone. That is the rule Kubernetes
itself uses for PodSpec.affinity, and the rule a Helm value and a Kustomize patch use, and keeping
it is a deliberate choice about what a user has to learn. Per-member folding was implemented once
and reverted, and the objection that carried the revert still holds: it obliges a user to learn a
merge semantic for exactly one field of one CRD, and kubectl explain is the only place they could
learn it. resources can fold per leaf without that cost because resources.cpu.min is a knob, not
a Kubernetes type a user already knows.
The cost of keeping the Kubernetes rule is real, and it is paid to an event rather than to a
second semantic. FoldCommonConfig returns a second value naming the members each replacement
discarded, and the reconciler emits a AffinityOverridden Warning on the CR:
role "server" group "default": the role group's config replaces config.affinity wholesale,
discarding the podAntiAffinity declared beneath it. config.affinity follows the Kubernetes rule
and is not merged per member; restate the discarded member alongside your own to keep it
Without it the loss is invisible everywhere: the CR still reads as the user wrote it, the pod spec is
valid, and every status condition stays green while the quorum quietly stops being spread. An
empty value (affinity: {}) clears — the single-node development escape hatch — and reports
nothing, because clearing is exactly what that value asks for. That clearing rule works here and
not in resources because the schemas differ: affinity is
x-kubernetes-preserve-unknown-fields, so the API server never prunes inside it and a stored {} is
always something the user wrote, while resources is structural and cpu: {} may be a pruning
artifact.
A product field is folded presence-wins per top-level field, with no depth. A composite
(a struct or pointer-to-struct) is refused unless it carries the kubedoop:"atomic" struct tag,
which accepts wholesale replacement for a value that is a single policy rather than a set of knobs;
the alternative is to flatten it into scal
Truncated - read the full file at https://github.com/zncdatadev/operator-go/blob/70c571fed73e31d79ba1a0b0e41050b6c93fcbce/AGENTS.md.
