Dedicated and Private Node Billing

How vBilling bills nodes that belong to one tenant cluster for their full capacity, whether they sit in the control plane cluster (dedicated nodes) or join the tenant cluster itself (vCluster Private Nodes, Auto Nodes, Standalone), without ever billing the workloads on them twice.

What dedicated nodes are

A dedicated node is a physical server, VM or bare-metal machine reserved for one tenant cluster. AI Clouds typically sell them by SKU (for example bm.gpu.h100.8), billing the whole machine regardless of utilization. Shared nodes, by contrast, are billed by what each pod allocates.

i
Dedicated nodes are control plane cluster nodes reserved through labels. With vCluster Private Nodes, nodes join the tenant cluster directly and the control plane cluster never sees them or their pods; see Private Nodes, Auto Nodes and Standalone.

How vBilling detects them

A node is dedicated to a tenant cluster when any of these labels match. The namespace is the tenant cluster's control plane cluster namespace; the name is the tenant cluster's name.

LabelMatches
vbilling.vcluster.com/tenant-clusterThe tenant cluster ID (vcluster-<namespace>-<name>) or name. Recommended.
vcluster.loft.sh/managed-byNamespace or name (also the name with a vcluster- or vc- namespace prefix stripped)
vcluster.loft.sh/clusterAs above
VCLUSTER_NODE_LABEL patternAny key=value where %s is replaced with the candidate names, e.g. tenant=%s

Taint dedicated nodes (for example dedicated=team-b:NoSchedule) and give the tenant's workloads a matching toleration, so other tenants cannot land on them.

kubectl label node bm-h100-01 vbilling.vcluster.com/tenant-cluster=vcluster-team-b-team-b kubectl label node bm-h100-01 vbilling.vcluster.com/sku=bm.gpu.h100.8 vbilling.vcluster.com/capacity-type=reserved kubectl taint node bm-h100-01 dedicated=team-b:NoSchedule

What gets billed

Each dedicated node produces its own events every window (resource_id = node name):

MetricQuantityDimensions
vcluster_private_node_hoursHours the node was healthysku (SKU label, instance type, or <n>x<GPU model>), capacity_type, instance_type, node
vcluster_gpu_hoursGPU capacity × hoursbilling_mode=dedicated_node, sku, gpu_type
vcluster_cpu_core_hoursCPU capacity × hoursbilling_mode=dedicated_node
vcluster_memory_gb_hoursMemory capacity (GiB) × hoursbilling_mode=dedicated_node

Price either the node-hours by SKU, or the capacity components by billing_mode, whichever matches how you sell. Set the other to zero.

No double billing

Pods that run on the tenant's own dedicated nodes are not metered again as shared usage: the node allocation already covers them. (v0.1 billed both the pods and the node.) A pod from a different tenant that lands on an untainted dedicated node is billed to its own tenant as shared usage.

Downtime

If a dedicated node is NotReady, or carries a taint listed in UNHEALTHY_NODE_TAINTS, its node-hours stop and its GPUs are recorded as vcluster_gpu_downtime_hours (informational, never invoiced), from the moment the condition or taint appeared.

Private Nodes, Auto Nodes and Standalone

With Private Nodes, vCluster syncs nothing to the control plane cluster: the nodes, their pods and their volumes exist only in the tenant cluster. vBilling meters these tenant clusters through their own API. Every node the tenant cluster's API reports that the control plane cluster does not have (and that is not one of vCluster's placeholder nodes, labeled vcluster.loft.sh/fake-node) is a private node, billed like a dedicated node with billing_mode=private_node. The tenant cluster's volumes and load balancers are billed as well. Shared tenant clusters are never billed twice, even when vBilling can read their API.

vBilling reads one Secret per tenant cluster, vbilling-kubeconfig in its namespace, and its RBAC is pinned to that name. Let vCluster write it with a read-only service account:

exportKubeConfig: additionalSecrets: - name: vbilling-kubeconfig serviceAccount: name: vbilling-reader namespace: kube-system clusterRole: vbilling-reader experimental: deploy: vcluster: manifests: |- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: vbilling-reader rules: - apiGroups: [""] resources: [nodes, pods, persistentvolumeclaims, services] verbs: [get, list, watch] - apiGroups: [resource.k8s.io] resources: [resourceclaims, resourceslices] verbs: [get, list] - apiGroups: [metrics.k8s.io] resources: [pods] verbs: [get, list]

vCluster writes kubeconfigs for localhost; vBilling connects to the tenant cluster's Service and verifies TLS against <name>.<namespace>. Tenant clusters outside the control plane cluster (vCluster Standalone, or any other cluster) are listed in TENANT_CLUSTERS_FILE with their own kubeconfig.

SettingEffect
PRIVATE_NODE_BILLING=node (default)Each private node is billed whole: node-hours, whole GPUs, CPU and memory capacity.
PRIVATE_NODE_BILLING=usageThe pods on private nodes are billed instead, like shared usage (namespaces in TENANT_EXCLUDE_NAMESPACES, default kube-system, are skipped).
vbilling.vcluster.com/billable=false node labelExempts a node, for example hardware the tenant owns. Works for dedicated nodes too.
TENANT_ADMIN_FALLBACK=trueAlso accepts the tenant cluster's admin kubeconfig (vc-<name>). Requires read access to Secrets; prefer the read-only export.

Outages. If a tenant cluster's API cannot be read, everything else is still metered and the gap is recorded under DATA_DIR. As soon as the API answers, exactly the missed windows are metered and appended as late events, deduplicated by ID, up to MAX_BACKFILL.

Scale-down. Nodes deleted between two windows (Auto Nodes scaling down, a machine removed) are billed until the second they were deleted. A deletion noticed only after a watch reconnects is dated to the last time vBilling saw the node, never later.

Whole GPUs, however they are carved

Dedicated and private nodes are billed for physical GPUs. MIG slices (mixed or single strategy, GKE GPU partitions) and time-slicing replicas (NVIDIA or GKE time-sharing) are converted back: a node exposing 56 MIG slices of 1g.10gb, or 32 time-sliced devices with 4 replicas, bills 8 GPUs. GPUs offered only through DRA ResourceSlices count too.

GPU model detection

The first label present wins: nvidia.com/gpu.product, nvidia.com/gpu.machine, amd.com/gpu.product-name, cloud.google.com/gke-accelerator, k8s.amazonaws.com/accelerator, accelerator. Spaces are replaced with dashes (NVIDIA H100 80GB HBM3 becomes NVIDIA-H100-80GB-HBM3). The -SHARED and -MIG-<profile> suffixes GPU Feature Discovery adds are removed, so the SKU is always the base GPU; DRA devices use their productName attribute.