Metrics Reference

Every metric vBilling v0.2 emits, how it is measured, and the dimensions billing backends can price on.

i
vBilling meters closed windows (60s by default) aligned to the wall clock. Allocation metrics come from object lifetimes, to the second, so missed windows can be backfilled exactly after a restart. Every event also carries tenant, region, project, resource_id, tenant_cluster, tenant_class and a deterministic id. Schema: usage-event.v1.json.

1. GPU Hours

vcluster_gpu_hours

Shared nodes Dedicated nodes
Source
Pod nvidia.com/gpu, amd.com/gpu, nvidia.com/mig-* and nvidia.com/gpu.shared requests, and GPUs allocated through DRA ResourceClaims; node labels (or DRA productName) for the GPU model
Unit
gpu-hours
Quantity
devices × GPU fraction × seconds running in the window / 3600

Billing starts at the first container start (image pulls and scheduling are free) and stops when the last container finishes. Pods that started or finished mid-window are billed for exactly the seconds they ran.

Fractional GPUs: a MIG slice counts as slices/7 of a GPU (A30: /4), e.g. mig-1g.10gb = 1/7. This works for the MIG mixed strategy (nvidia.com/mig-* resources), the single strategy (slices exposed as nvidia.com/gpu, product label ending in -MIG-<profile>) and GKE GPU partitions (cloud.google.com/gke-gpu-partition-size). On time-sliced nodes (nvidia.com/gpu.replicas > 1, or GKE cloud.google.com/gke-max-shared-clients-per-gpu with time-sharing or MPS) a device counts as 1/replicas. A DRA claim counts the devices allocated to it (a MIG device at its profile's fraction), split across the pods it is reserved for. The SKU becomes <model>-mig-<profile> or <model>-timeslice-<n>, always against the base GPU model (the -SHARED and -MIG- label suffixes are dropped).

Dedicated and private nodes: the node's GPU capacity is billed in whole GPUs with billing_mode=dedicated_node (control plane cluster nodes reserved for the tenant) or billing_mode=private_node (vCluster Private Nodes, Auto Nodes and Standalone, metered through the tenant cluster's own API): MIG slices and time-slicing replicas are converted back, so 8 GPUs shared 4 ways bill as 8, not 32. Pods on those nodes are not billed again. See Dedicated and Private Nodes.

Fair metering: GPU time on a node that is NotReady or carries an UNHEALTHY_NODE_TAINTS taint is emitted as vcluster_gpu_downtime_hours instead. So is a dedicated or private node's GPU time while its GPUs are installed (GPU Feature Discovery's nvidia.com/gpu.count) but not allocatable, for example while the driver rebuilds after a reboot or a MIG change.

Dimensions

KeyMeaning
skuGPU model (or the node's SKU label), plus the fractional profile
gpu_typeNormalized GPU model
gpu_profilefull, mig-<profile>, timeslice-<n>
capacity_typeon-demand, spot, preemptible, reserved
billing_modeshared, dedicated_node or private_node
zoneNode zone

2. CPU Core-Hours

vcluster_cpu_core_hours

Shared nodes Dedicated nodes
Source
CPU_MEMORY_BASIS: metrics-server usage, container requests, or the larger of both; node capacity for dedicated nodes
Unit
core-hours
Quantity
cores × seconds in the window / 3600

usage samples metrics-server for the live window only and is never extrapolated into backfilled windows. requests is computed from pod lifetimes, so it is backfilled exactly. max bills the larger of usage and requests per pod.

The tenant cluster's own control plane pods (app=vcluster) are covered by instance hours and excluded unless METER_CONTROL_PLANE=true.

Dimensions

KeyMeaning
capacity_typeCapacity type of the node
billing_modeshared, dedicated_node or private_node
namespaceNamespace inside the tenant cluster (with METER_BY_NAMESPACE)

3. Memory GiB-Hours

vcluster_memory_gb_hours

Shared nodes Dedicated nodes
Source
As CPU: usage, requests or max; node capacity for dedicated nodes
Unit
gib-hours
Quantity
GiB × seconds / 3600

1 GiB = 10243 bytes.

Dimensions

KeyMeaning
capacity_typeCapacity type of the node
billing_modeshared, dedicated_node or private_node

4. Dedicated Node Hours

vcluster_private_node_hours

Dedicated nodes
Source
Nodes dedicated to the tenant cluster (see Dedicated nodes)
Unit
node-hours
Quantity
seconds the node existed and was healthy in the window / 3600, one event per node

Downtime on an unhealthy node is not billed.

Dimensions

KeyMeaning
skuSKU label, instance type, or <n>x<GPU model>
capacity_typeCapacity type of the node
nodeNode name (also resource_id)
instance_typenode.kubernetes.io/instance-type

5. Storage GiB-Hours

vcluster_storage_gb_hours

Shared nodes
Source
Bound PVCs synced from the tenant cluster (provisioned capacity, else the request)
Unit
gib-hours
Quantity
GiB × seconds bound in the window / 3600

The tenant cluster's own data volume is excluded unless METER_CONTROL_PLANE=true.

Dimensions

KeyMeaning
skuStorage class

6. Tenant Cluster Hours

vcluster_instance_hours

Shared nodes Dedicated nodes
Source
Tenant cluster control plane readiness
Unit
hours
Quantity
seconds ready in the window / 3600

No charge while a tenant cluster is provisioning or asleep (no ready replica).

Dimensions

KeyMeaning
tenant_classTenant class

7. LoadBalancer Hours

vcluster_lb_hours

Shared nodes
Source
Services of type LoadBalancer with an assigned address
Unit
hours
Quantity
load balancers × seconds / 3600

A load balancer that is still pending (no IP) is not billed.

8. Network Egress GiB

vcluster_network_egress_gb

Prometheus / informational
Source
Prometheus: increase(container_network_transmit_bytes_total{namespace}) over the window, evaluated at the window end, or your EGRESS_QUERY
Unit
gib
Quantity
bytes / 10243

The default counts all transmitted bytes. To bill only internet egress, point EGRESS_QUERY at flow metrics from your CNI or fabric.

9. GPU Utilization (informational)

vcluster_gpu_utilization

Prometheus / informational
Source
DCGM exporter via Prometheus (DCGM_FI_DEV_GPU_UTIL; for GPUs in MIG mode, which DCGM reports per instance, DCGM_FI_PROF_GR_ENGINE_ACTIVE), matched on the workload's namespace or, where Prometheus renames it (kube-prometheus-stack), exported_namespace. GPU_UTIL_QUERY replaces the query for other exporters; none turns it off.
Unit
util-hours
Quantity
average utilization % × hours, per GPU model

Informational: delivered to the ledger, webhooks and exports, never to Stripe or Metronome (Lago keeps it for v0.1 plans).

Dimensions

KeyMeaning
gpu_typeGPU model

10. GPU Downtime Hours (informational)

vcluster_gpu_downtime_hours

Prometheus / informational
Source
Allocated GPU time on NotReady or unhealthy nodes
Unit
gpu-hours
Quantity
as GPU hours

An auditable record of the time tenants were not charged for, never sent to billing backends.

Dimensions

KeyMeaning
gpu_typeGPU model
nodeNode

Custom metrics

Declare operator metrics with CUSTOM_METRICS="code:unit[:key|key],..." (for example inference_output_tokens:tokens:model|region) and push closed windows to POST /api/v1/events. Billing backends create meters and billable metrics for them at startup.