Skip to main content

Compliance Operator: Writing Your Own Checks with CEL and CustomRules

Five years ago I first tested and described how to achieve compliance for a Kubernetes cluster in my article Compliance Operator. Many years passed since then, but up to this day the Compliance Operator is still one of the most important tools I recommend to my customers and one of the first that is installed on a new cluster. Since the first release of the Operator, it has been constantly improved and extended, for example by adding your own guardrails to the cluster to make sure that the cluster is compliant with your organisation’s requirements.

Today I would like to show you how to write your own checks with the CEL scanner and the CustomRule resource.

Introduction

The Compliance Operator does a great job when it comes to industry benchmarks such as CIS, NIST 800-53, PCI-DSS or BSI. The idea described at Compliance Operator still works today. You define or select a profile and a time of scanning and the operator does the rest for you. Sooner or later, however, every auditor asks a question that is not part of any benchmark because it is individual for each organisation.

For example: "Who owns this namespace?", "Show me that every application is isolated on the network by default." or "Who exactly is cluster-admin, and why?" Until now, the answer usually was a shell script, an automation playbook, a screenshot or a check mark in a spreadsheet. With the CEL scanner and the CustomRule resource, the Compliance Operator can finally check such organisation-specific requirements itself — scheduled, repeatable and with results in the same place as all other compliance checks. All of this defined declaratively, so you can easily integrate it into your GitOps pipeline.

In this article I will build a small but real-world examples: three custom checks that act as platform guardrails. Two of them look at application namespaces, the third one audits who holds cluster-admin. This will extend the standard profile of the Compliance Operator and make it more useful for your organisation.

Once you have the rules in place, you can add them to a TailoredProfile so they become part of your compliance assessment.

The idea is to build the three rules and then add them to a TailoredProfile to make them part of your compliance assessment. If you just need one of the rules, simply remove the other two from the profile.

A Bit of History: CEL in the Compliance Operator

The Compliance Operator traditionally uses OpenSCAP to evaluate its rules. The content (XCCDF / OVAL) is generated by the ComplianceAsCode1 project and shipped with the operator. Writing your own rules in that format is possible, but let us be honest: nobody enjoys building their own content image just to check a label.

The Common Expression Language (CEL) changes this. CEL is a small, side-effect free expression language. If you have ever written a ValidatingAdmissionPolicy, you already know it. The integration into the Compliance Operator was finally globally available in version 1.10.0.

CEL does not replace the existing XCCDF profiles. It extends them. Your CIS or PCI-DSS scans continue to work exactly as before.

The Use Case: Platform Guardrails

Many customers have some kind of "platform contract". It is written down in a Confluence page, and it is usually not checked by anybody. Three requirements appear in nearly all of them:

  1. Ownership: every application namespace carries labels that tell who owns it and who pays for it. That is the asset inventory an auditor asks for.

  2. Network segmentation by default: every application namespace contains a default-deny NetworkPolicy for incoming traffic. Applications then explicitly allow what they need.

  3. Privileged access: only an approved list of identities is bound to the ClusterRole cluster-admin. The platform’s own service accounts are allowed, of course, but personal accounts or application service accounts are not.

Why are these good candidates for a CustomRule?

  • The label keys and the list of administrators are specific to each organisation. The operator cannot ship a generic rule for them.

  • The built-in CIS rule ocp4-configure-network-policies-namespaces only verifies that a namespace has at least one NetworkPolicy. A single policy that allows everything would pass. We want a real check: we want the deny-baseline.

  • All three checks only need information from the Kubernetes API. This is exactly what CustomRules are designed for.

CustomRules evaluate Kubernetes/OpenShift API resources only. They cannot inspect files or processes on the nodes. For node-level checks you still use the shipped node profiles, such as ocp4-cis-node.

Prerequisites

  • An OpenShift cluster with cluster-admin permissions.

  • The Compliance Operator 1.10.0 or newer installed in the namespace openshift-compliance. Version 1.8.0 and 1.9.x work as well, but the feature is Technology Preview there.

  • The command line tools oc and jq.

If the operator is not installed yet, have a look at my older articles:

Verify the installed version:

oc get csv -n openshift-compliance

How a CustomRule Works

Before we write our own rules, let us look at the anatomy of a CustomRule. Such YAML files can get large and messy, but the most important parts are expression and inputs:

CustomRule YAML
apiVersion: compliance.openshift.io/v1alpha1
kind: CustomRule
metadata:
  name: my-rule
  namespace: openshift-compliance (1)
spec:
  id: my_rule (2)
  title: A short title
  severity: medium (3)
  checkType: Platform (4)
  scannerType: CEL (5)
  inputs: (6)
    - name: pods
      kubernetesInputSpec:
        apiVersion: v1
        resource: pods
  expression: |- (7)
    pods.items.all(p, has(p.spec.securityContext))
  failureReason: Text shown when the check fails (8)
  instructions: How to find and fix the problem (9)
1CustomRules live in the namespace of the operator.
2The ID becomes part of the name of the ComplianceCheckResult (underscores are automatically converted to dashes).
3Can be one of info, low, medium, high (default is medium).
4CustomRules only support Platform.
5Must be CEL.
6The Kubernetes resources the rule needs. Each input becomes a variable in the expression. Without resourceName you get a list object, so the actual objects are found in .items. With resourceNamespace you can limit the query to one namespace.
7The CEL expression. It must return true when the cluster is compliant.
8The failureReason is added to the warnings of the ComplianceCheckResult if the check fails.
9instructions (as well as description and rationale) are copied into the ComplianceCheckResult. Use them to tell the reader how to find the offending objects.
The operator reads the inputs with the ServiceAccount api-resource-collector. This account is basically a cluster-reader, so namespaces, NetworkPolicies, pods, RBAC objects etc. can be read without any additional configuration. If your rule needs resources of an operator (for example HyperConverged of OpenShift Virtualization), you must grant read access to this ServiceAccount yourself.
A CustomRule produces one result for the whole rule, not one result per object. If one namespace out of 200 violates the policy, the rule fails. That is why the instructions field should always contain a command that lists the offending objects.

Step 1: The Ownership Rule

The first rule, that I would like to build, verifies that every application namespace has the labels example.com/owner and example.com/cost-centre with a non-empty value. Our auditor wants to know who owns the namespace and who pays for it.

Platform namespaces must be excluded, otherwise the rule would fail immediately as they do no not have such labels. A regular expression for this: everything that starts with openshift or kube-, plus the namespace default is ignored. In addition, I have also added the namespaces open-cluster-management, stackrox and rhacs- to the exclusion list, these are namespace for Advanced Cluster Management and Advanced Cluster Security respectively and at least on my own demo cluster they exist.

Do not worry about the following YAML. It looks huge, but most of it is just comments. The important parts are the expression field and the inputs section: they define what to check and where.
01-customrule-namespace-ownership-labels.yaml
apiVersion: compliance.openshift.io/v1alpha1
kind: CustomRule
metadata:
  name: namespace-ownership-labels
  namespace: openshift-compliance
  annotations:
    example.com/controls: "Nist Rule 1.2.3.4" (1)
spec:
  id: namespace_ownership_labels
  title: Application namespaces must have ownership labels set (2)
  severity: medium (3)
  checkType: Platform
  scannerType: CEL
  inputs:
    - name: namespaces (4)
      kubernetesInputSpec:
        apiVersion: v1
        resource: namespaces
  expression: |- (5)
    namespaces.items
      .filter(ns, !ns.metadata.name.matches('^(openshift.*|kube-.*|open-cluster-management.*|cert-manager.*|stackrox.*|rhacs-.*|default)$'))
      .all(ns,
        has(ns.metadata.labels) &&
        ['example.com/owner', 'example.com/cost-centre'].all(key,
          key in ns.metadata.labels && ns.metadata.labels[key] != ''
        )
      )
  description: |- (6)
    Every application namespace must have the labels 'example.com/owner'
    and 'example.com/cost-centre' with a non-empty value set. Platform namespaces (openshift*, kube-*, default)
    are excluded.
  rationale: |- (7)
    Auditors expect an up-to-date inventory of system components and a named
    owner for each of them.
  failureReason: |- (8)
    At least one application namespace is missing the label
    'example.com/owner' or 'example.com/cost-centre' (or the value is empty).
    Run the command from the instructions of this check to list the
    offending namespaces and label them, for example:
    oc label namespace <name> example.com/owner=<team> example.com/cost-centre=<id>
  instructions: |- (9)
    List all application namespaces without complete ownership labels:
    oc get namespaces -o json | jq -r '.items[]
      | select(.metadata.name | test("^(openshift.*|kube-.*|open-cluster-management.*|cert-manager.*|stackrox.*|rhacs-.*|default)$") | not)
      | select((.metadata.labels["example.com/owner"] // "") == ""
            or (.metadata.labels["example.com/cost-centre"] // "") == "")
      | .metadata.name'
1Optional: map the rule to your control framework.
2The title of the rule.
3The severity of the rule. Can be one of info, low, medium, high (default is medium).
4The input section defines the WHERE of the check. The resource namespaces contains the list of all namespaces.
5The expression field reads like a sentence:
  • filter(…​) removes all platform namespaces. matches() uses RE2 syntax. Note the anchors ^ and $: without them matches() searches for a substring.

  • all(…​) returns true only if the condition is true for every of the remaining namespaces.

  • has(ns.metadata.labels) protects against namespaces without any label. On a real cluster this does not happen, since Kubernetes always sets kubernetes.io/metadata.name, but it costs nothing.

  • The inner all() iterates over the list of mandatory label keys. key in map checks whether the key exists and that the value is not empty.

6Detailed description of the rule.
7Rationale: why are we checking this?
8The failureReason is added to the warnings of the ComplianceCheckResult if the check fails.
9The instructions are copied into the ComplianceCheckResult. Use it to tell the reader how to find the offending objects.
This is our first CustomRule object. We will define the other rules in the same way, then apply them all at once.

Step 2: The Default-Deny Rule

The second rule is more interesting because it combines two inputs:

  • the list of namespaces and

  • the list of all NetworkPolicies.

For every application namespace there must exist at least one NetworkPolicy that:

  • lives in that namespace,

  • selects all pods (empty podSelector),

  • covers the policy type Ingress and

  • does not define any ingress rule, which means: nothing is allowed.

In short, every application namespace must have a deny-all NetworkPolicy for ingress.

Create the file 02-customrule-namespace-default-deny-ingress.yaml:

02-customrule-namespace-default-deny-ingress.yaml
apiVersion: compliance.openshift.io/v1alpha1
kind: CustomRule
metadata:
  name: namespace-default-deny-ingress
  namespace: openshift-compliance
  annotations:
    example.com/controls: "Nist Rule 1.2.3.4"
spec:
  id: namespace_default_deny_ingress
  title: Application namespaces have a default-deny ingress NetworkPolicy
  description: |-
    Every application namespace must contain a NetworkPolicy that selects
    all pods (empty podSelector), covers the policy type 'Ingress' and
    defines no ingress rules.
    Additional allow policies are expected on top of it. Platform namespaces
    (openshift*, kube-*, default) are excluded.
  rationale: |-
    Without a default-deny baseline every pod accepts traffic from every
    other pod in the cluster.
  severity: high
  checkType: Platform
  scannerType: CEL
  inputs:
    - name: namespaces
      kubernetesInputSpec:
        apiVersion: v1
        resource: namespaces
    - name: networkpolicies (1)
      kubernetesInputSpec:
        apiVersion: networking.k8s.io/v1 (2)
        resource: networkpolicies
  expression: |-
    namespaces.items
      .filter(ns, !ns.metadata.name.matches('^(openshift.*|kube-.*|open-cluster-management.*|cert-manager.*|stackrox.*|rhacs-.*|default)$'))
      .all(ns,
        networkpolicies.items.exists(np,
          np.metadata.namespace == ns.metadata.name &&
          (!has(np.spec.podSelector.matchLabels) || size(np.spec.podSelector.matchLabels) == 0) &&
          (!has(np.spec.podSelector.matchExpressions) || size(np.spec.podSelector.matchExpressions) == 0) &&
          (!has(np.spec.policyTypes) || 'Ingress' in np.spec.policyTypes) &&
          (!has(np.spec.ingress) || size(np.spec.ingress) == 0)
        )
      )
  failureReason: |-
    At least one application namespace has no default-deny ingress
    NetworkPolicy (empty podSelector, policyType Ingress, no ingress rules).
    Run the command from the instructions of this check to list the
    offending namespaces and add a default-deny policy to each of them.
  instructions: |-
    Lists every application namespace that has no default-deny ingress policy.
    A policy counts as default-deny if it selects all pods, covers Ingress and
    allows nothing. Copy the lines below into a bash shell:

    for ns in $(oc get namespaces -o name | cut -d/ -f2 \
        | grep -Ev '^(openshift.*|kube-.*|open-cluster-management.*|cert-manager.*|stackrox.*|rhacs-.*|default)$'); do
      oc get networkpolicies -n "$ns" -o json | jq -e '.items[] | select(
          (.spec.podSelector.matchLabels      // {} | length) == 0 and
          (.spec.podSelector.matchExpressions // [] | length) == 0 and
          (.spec.policyTypes // ["Ingress"] | index("Ingress")) and
          (.spec.ingress // [] | length) == 0
        )' > /dev/null || echo "missing default-deny: $ns"
    done
1A second input. NetworkPolicies are namespace-scoped; since no resourceNamespace is set, the policies of all namespaces are fetched.
2For resources outside the core API group, group and version are written together, exactly like the apiVersion in a manifest.

Some details of the expression field are worth a second look:

  • filter(…​) removes all platform namespaces that matches()

  • all(…​) returns true only if the condition is true for every of the remaining namespaces.

  • a network policy object must exist that:

    • lives in that namespace,

    • has an empty podSelector,

    • has not policyTypes set to Ingress,

    • has no ingress rules,

  • Additional allow policies (for example to let the OpenShift router in) are fine. We only require that the deny baseline exists.

This is just an example that shows how to use two inputs in a CustomRule. If you enforce the baseline cluster-wide with a BaselineAdminNetworkPolicy or an AdminNetworkPolicy instead of per-namespace NetworkPolicies, this rule is not the right one for you. In that case write a rule that checks the (cluster-scoped) admin policy instead.

Step 3: The cluster-admin Allow-List

The third rule answers the classic auditor question "Who is cluster-admin?".

The CIS OpenShift Benchmark (control 5.1.1) recommends to restrict the use of cluster-admin, but it cannot know who is allowed in your organisation. That is a perfect job for a CustomRule.

A common mistake is to look only at the ClusterRoleBinding that is called cluster-admin. Any ClusterRoleBinding can reference the ClusterRole cluster-admin, and OpenShift itself ships several of them. Therefore, the rule looks at all ClusterRoleBindings whose roleRef points to cluster-admin, and checks every subject of these bindings.

Before you write the allow-list, look at the current state of your cluster:

oc get clusterrolebindings -o json | jq -r '.items[]
  | select(.roleRef.kind == "ClusterRole" and .roleRef.name == "cluster-admin")
  | .metadata.name as $crb | .subjects[]?
  | "\($crb)\t\(.kind)\t\(.namespace // "-")\t\(.name)"' | column -t

This selects all ClusterRoleBindings that reference the ClusterRole cluster-admin and lists the subjects of each binding. The output looks like this:

cluster-admin                    Group           -                                   system:masters
cluster-admin-0                  ServiceAccount  openshift-gitops                    openshift-gitops-argocd-application-controller
cluster-admin-crb                Group           -                                   cluster-admin
cluster-admins                   Group           -                                   system:cluster-admins
cluster-admins                   User            -                                   system:admin
cluster-network-operator         ServiceAccount  openshift-network-operator          cluster-network-operator
cluster-storage-operator-role    ServiceAccount  openshift-cluster-storage-operator  cluster-storage-operator
cluster-version-operator         ServiceAccount  openshift-cluster-version           default
...

You will find the platform-internal groups system:masters and system:cluster-admins, the user system:admin and a number of service accounts in openshift-* namespaces that belong to the platform operators. Everything else was added later by someone and deserves a closer look.

The allow-list — the subjects the rule treats as acceptable — has three parts:

Subject kindAllowed

Group

The platform groups system:masters and system:cluster-admins, plus the group of the platform team (here cluster-admin, for example synchronised from LDAP or Entra ID).

User

The platform user system:admin and one named break-glass account. Personal accounts are not allowed. The platform team gets its rights through the group.

ServiceAccount

Every service account in a platform namespace (openshift*, kube-*), plus an explicit list of namespace/name pairs for tools such as a cluster-wide Argo CD instance.

The following file is a full example of a CustomRule that checks the allow-list for cluster-admin. The callouts below walk through every part of the expression.

03-customrule-cluster-admin-allow-list.yaml
apiVersion: compliance.openshift.io/v1alpha1
kind: CustomRule
metadata:
  name: cluster-admin-allow-list
  namespace: openshift-compliance
  annotations:
    example.com/controls: "Nist Rule 1.2.3.4"
spec:
  id: cluster_admin_allow_list
  title: Only approved identities are allowed to hold cluster-admin privileges
  description: |-
    Every subject of every ClusterRoleBinding that grants cluster-admin
    privileges must be on the organisation's allow-list. A binding grants
    cluster-admin privileges if it references the ClusterRole 'cluster-admin'
    or any other ClusterRole with a wildcard rule ('*' API groups, '*'
    resources, '*' verbs), for example a renamed copy of cluster-admin.
    Allowed are: the platform-internal groups and users of OpenShift,
    service accounts in platform namespaces (openshift*, kube-*), an
    explicitly named break-glass account, explicitly listed service accounts
    and the group of platform administrators - but only if that group exists
    and every member of it is an approved administrator. Everything else, in
    particular personal user accounts, application service accounts and
    broad groups such as 'system:authenticated', fails the check.
  rationale: |-
    cluster-admin grants unrestricted access to every resource in the
    cluster. Privileged access must be limited to a small, documented and
    reviewed set of identities.
  severity: high
  checkType: Platform
  scannerType: CEL
  inputs:
    - name: crbs (1)
      kubernetesInputSpec:
        apiVersion: rbac.authorization.k8s.io/v1
        resource: clusterrolebindings
    - name: clusterroles (2)
      kubernetesInputSpec:
        apiVersion: rbac.authorization.k8s.io/v1
        resource: clusterroles
    - name: groups (3)
      kubernetesInputSpec:
        apiVersion: user.openshift.io/v1
        resource: groups
  expression: |-
    // ---- allow-list: might be exported into a ConfigMap ---- (4)
    // ---- every dynamic value is stored in a map ----
    [{
      'platformGroups':         ['system:masters', 'system:cluster-admins'],
      'adminGroups':            ['cluster-admin'],
      'approvedAdmins':         ['alice.admin@example.com', 'bob.admin@example.com'],
      'allowedUsers':           ['system:admin', 'breakglass-admin'],
      'allowedServiceAccounts': ['gitops-system/argocd-application-controller', 'rhacs-operator/rhacs-operator-controller-manager']
    }].all(cfg,
      crbs.items (5)
        .filter(crb,
          crb.roleRef.kind == 'ClusterRole' &&
          has(crb.subjects) && crb.subjects != null &&
          (crb.roleRef.name == 'cluster-admin' || (6)
           clusterroles.items.exists(cr,
             cr.metadata.name == crb.roleRef.name &&
             has(cr.rules) && cr.rules != null &&
             cr.rules.exists(r,
               has(r.apiGroups) && '*' in r.apiGroups &&
               has(r.resources) && '*' in r.resources &&
               '*' in r.verbs))))
        .all(crb, crb.subjects.all(s, (7)
          // groups: built-in groups, or an admin group that exists and has only approved members
          (s.kind == 'Group' && ( (8)
            s.name in cfg.platformGroups ||
            (s.name in cfg.adminGroups &&
             groups.items.exists(g, g.metadata.name == s.name) &&
             groups.items.filter(g, g.metadata.name == s.name).all(g,
               !has(g.users) || g.users == null ||
               g.users.all(u, u in cfg.approvedAdmins)))
          )) ||
          // users: allowed users, or service accounts of platform namespaces written as User
          (s.kind == 'User' && ( (9)
            s.name in cfg.allowedUsers ||
            s.name.matches('^system:serviceaccount:(openshift[^:]*|kube-[^:]*):[^:]+$')
          )) ||
          // service accounts: platform namespaces, or explicitly listed namespace/name pairs
          (s.kind == 'ServiceAccount' && has(s.namespace) && ( (10)
            s.namespace.matches('^(openshift.*|kube-.*)$') ||
            (s.namespace + '/' + s.name) in cfg.allowedServiceAccounts
          ))
        ))
    )
  failureReason: |-
    At least one subject that holds cluster-admin privileges is not on the
    allow-list, or the platform administrator group contains a member that
    is not an approved administrator.
  instructions: |- (11)
    List all subjects with cluster-admin privileges that are not on the
    allow-list:
    jq -rn \
      --slurpfile crbs <(oc get clusterrolebindings -o json) \
      --slurpfile roles <(oc get clusterroles -o json) \
      --slurpfile groups <(oc get groups -o json) '
      # ---- allow-list: edit here ----
      {
        platformGroups:         ["system:masters", "system:cluster-admins"],
        adminGroups:            ["cluster-admin"],
        approvedAdmins:         ["alice.admin@example.com", "bob.admin@example.com"],
        allowedUsers:           ["system:admin", "breakglass-admin"],
        allowedServiceAccounts: ["gitops-system/argocd-application-controller", "rhacs-operator/rhacs-operator-controller-manager"]
      } as $cfg
      | [$roles[0].items[]
          | select(.metadata.name == "cluster-admin" or any(.rules[]?;
              ((.apiGroups // []) | index("*")) and ((.resources // []) | index("*"))
              and ((.verbs // []) | index("*"))))
          | .metadata.name] as $adminRoles
      | $crbs[0].items[]
      | select(.roleRef.kind == "ClusterRole"
          and (.roleRef.name == "cluster-admin" or (.roleRef.name | IN($adminRoles[]))))
      | .metadata.name as $crb | .roleRef.name as $role | .subjects[]? | . as $s
      | select(
          ((.kind == "Group") and ((.name | IN($cfg.platformGroups[]))
             or ((.name | IN($cfg.adminGroups[]))
                 and ([$groups[0].items[] | select(.metadata.name == $s.name)] as $g
                      | ($g | length) > 0
                        and all($g[]; (.users // []) | all(IN($cfg.approvedAdmins[])))))))
          or ((.kind == "User") and ((.name | IN($cfg.allowedUsers[]))
                or (.name | test("^system:serviceaccount:(openshift[^:]*|kube-[^:]*):[^:]+$"))))
          or ((.kind == "ServiceAccount") and (((.namespace // "") | test("^(openshift.*|kube-.*)$"))
                or ((.namespace // "") + "/" + .name | IN($cfg.allowedServiceAccounts[]))))
          | not)
      | "\($crb) -> \($role): \(.kind) \(.namespace // "-")/\(.name)"
        + (if .kind == "Group" then
             ([$groups[0].items[] | select(.metadata.name == $s.name) | (.users // [])[]]
              | if length == 0 then " (no Group object or no members)"
                else " (members: " + join(", ") + ")" end)
           else "" end)'
1The list of all ClusterRoleBindings.
2The list of all ClusterRoles. It is used to find roles that are equivalent to cluster-admin.
3The list of all OpenShift Group objects. It is used to verify the members of the allowed group.
4A map for the alow list and any dynamic values. Easier to manage than a long list of strings.
5Selecting the relevant ClusterRoleBindings: kind == ClusterRole, subject exists
6Role reference name is cluster-admin or any other ClusterRole with a wildcard rule ('' API groups, '' resources, '*' verbs)
7Subject is a Group OR a User OR a ServiceAccount
8Group is either a platform group OR an admin group that exists and has only approved admins
9User is either an allowed user OR a service account of a platform namespace written as User
10ServiceAccount is either a platform namespace OR an explicitly listed namespace/name pair (here: the Argo CD application controller)
11Pretty decent command to list the offending bindings.

Some details are important here, and they are the reason why I did not simply allow everything that starts with system::

  • Broad system groups: system:authenticated, system:unauthenticated and system:serviceaccounts:<namespace> are system groups as well. Binding one of them to cluster-admin is about the worst thing you can do to a cluster. A prefix match on system: would happily accept them. Therefore the rule uses an explicit list of groups.

  • Service accounts as User subjects: A service account can also appear as a subject of kind User with the name system:serviceaccount:<namespace>:<name>. The rule treats such subjects like service accounts and only accepts them if the namespace is a platform namespace.

  • Exact namespace/name pairs: For service accounts outside the platform namespaces the rule compares namespace/name. A service account with the same name in another namespace does not pass.

The rule deliberately ignores RoleBindings. A RoleBinding to cluster-admin only grants full access within one namespace. ClusterRoles that grant broad privileges without using //* wildcards are also outside the scope of this example; extend the expression if you need to treat those as equivalent to cluster-admin.

Step 4: Create the Rules and Check Their Status

Apply the three rules:

oc apply -f 01-customrule-namespace-ownership-labels.yaml
oc apply -f 02-customrule-namespace-default-deny-ingress.yaml
oc apply -f 03-customrule-cluster-admin-allow-list.yaml

The operator validates the CEL expression immediately. Only rules in the state Ready can be used in a profile:

oc get customrules -n openshift-compliance
NAME                             STATUS   AGE
cluster-admin-allow-list         Ready    10s
namespace-default-deny-ingress   Ready    12s
namespace-ownership-labels       Ready    14s

If the status is Error, the reason can be found in the status of the object:

oc get customrule namespace-ownership-labels -n openshift-compliance -o jsonpath='{.status.errorMessage}{"\n"}'
Do not use the short name cr. In the Compliance Operator, cr stands for complianceremediations. I fell into that trap more than once.

Step 5: TailoredProfile and ScanSettingBinding

CustomRules are never scanned directly. They must be bundled into a TailoredProfile, which is then bound to a ScanSetting (when to scan) via a ScanSettingBinding (which profile to scan).

Create the file tailoredprofile.yaml:

tailoredprofile.yaml
apiVersion: compliance.openshift.io/v1alpha1
kind: TailoredProfile
metadata:
  name: platform-guardrails               (1)
  namespace: openshift-compliance
spec:
  title: Platform guardrails
  description: Organisation-specific checks for namespaces and privileged access
  enableRules:
    - kind: CustomRule                    (2)
      name: namespace-ownership-labels
      rationale: Asset inventory and accountability for every namespace
    - kind: CustomRule
      name: namespace-default-deny-ingress
      rationale: Network segmentation as the default for every namespace
    - kind: CustomRule
      name: cluster-admin-allow-list
      rationale: Privileged access limited to approved identities
1The name of the TailoredProfile becomes the name of the ComplianceScan and the prefix of all results.
2The kind must be set to CustomRule. Without it, the operator looks for a regular Rule with that name.
A TailoredProfile cannot mix CEL-based checks with the classic OpenSCAP-based Rule objects (such as the CIS rules). CustomRules may only be combined with other CEL rules. Keep your custom checks in their own profile, which is a good idea anyway.

Create the file scansettingbinding.yaml:

scansettingbinding.yaml
apiVersion: compliance.openshift.io/v1alpha1
kind: ScanSettingBinding
metadata:
  name: platform-guardrails
  namespace: openshift-compliance
profiles:
  - apiGroup: compliance.openshift.io/v1alpha1
    kind: TailoredProfile
    name: platform-guardrails
settingsRef:
  apiGroup: compliance.openshift.io/v1alpha1
  kind: ScanSetting
  name: default                           (1)
1The default ScanSetting created by the operator. It runs the scan daily at 01:00 (0 1 * * *). Use your own ScanSetting if you need a different schedule.

Step 6: Check the Results

Creating the binding starts the first scan immediately. Watch the ComplianceSuite until the phase is DONE:

oc get compliancesuite platform-guardrails -n openshift-compliance

This returns something like:

NAME                  PHASE   RESULT
platform-guardrails   DONE    NON-COMPLIANT

Why is it non-compliant? Let us look at the results.

To list the results of this scan, run:

oc get ccr -n openshift-compliance -l compliance.openshift.io/scan-name=platform-guardrails

On a cluster that has been running for a while, the output will most likely look like this (example output):

NAME                                                 STATUS   SEVERITY
platform-guardrails-cluster-admin-allow-list         FAIL     high
platform-guardrails-namespace-default-deny-ingress   FAIL     high
platform-guardrails-namespace-ownership-labels       FAIL     medium

Do not be surprised: this is the actual state of the cluster, and exactly what an auditor wants to know. The failure reason and the instructions are stored in the result:

oc get ccr platform-guardrails-namespace-ownership-labels -n openshift-compliance \
  -o jsonpath='{.warnings}{"\n"}{.instructions}{"\n"}'

Copy the command from the instructions to get the list of namespaces that must be fixed.


Step 7: Test the Rules

To be sure the rules do what they should, create objects that violate all three of them, verify that the checks fail, fix the findings and verify that the checks pass. On a lab cluster that is already compliant, that gives a clean before-and-after picture.

Create a test namespace without labels and without a NetworkPolicy, and bind cluster-admin to a personal account that is not on the allow-list:

oc create namespace cel-demo
oc create clusterrolebinding cel-demo-admin --clusterrole=cluster-admin --user=alice

Trigger a rescan so you do not have to wait until 01:00 — annotate the ComplianceScan:

oc annotate compliancescans/platform-guardrails -n openshift-compliance compliance.openshift.io/rescan=

Wait until the suite is DONE again and check the results: all three checks are FAIL. The helper commands from the instructions list the namespace cel-demo and the binding cel-demo-admin: User -/alice.

Now fix the findings. Remove the binding and label the namespace:

oc delete clusterrolebinding cel-demo-admin
oc label namespace cel-demo example.com/owner=platform-team example.com/cost-centre=CC-4711

cat <<'EOF' | oc apply -f -
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-ingress
  namespace: cel-demo
spec:
  podSelector: {}
  policyTypes:
  - Ingress
EOF

Rescan again:

oc annotate compliancescans/platform-guardrails -n openshift-compliance compliance.openshift.io/rescan=

As long as these were the only offenders, all three checks now report PASS. Finally, remove the test namespace:

oc delete namespace cel-demo

Summary

With CustomRule and the CEL scanner, the Compliance Operator finally covers the "last mile" of compliance: the requirements that are specific to your organisation. The three rules in this article are quite huge YAML files, most of it documentation, they run on the same schedule as your CIS or PCI-DSS scans, and their results end up in the same place, with the same labels, metrics and alerts.

Namespace ownership, default-deny NetworkPolicies and the cluster-admin allow-list are just a starting point. Other good candidates are: mandatory TLS settings on Routes, allowed image registries, or an audit that no application team runs its own database. If you can express it with oc get …​ -o json | jq, you can most likely express it in CEL.

Thank You

This article would not exist without Max Thüringer and Steffen Lützenkirchen. After listening to the talk "More than just Compliance" at OpenShift User Conference in Vienna, I wanted to see how far the CEL scanner of the Compliance Operator can go.

Thank you, Max Thüringer and Steffen Lützenkirchen, for sharing your ideas and for the inspiration.


Discussion

← Previous
Use arrow keys to navigate
Next →
←
→