Metrics Reference
Every metric vBilling v0.2 emits, how it is measured, and the dimensions billing backends can price on.
tenant, region, project, resource_id, tenant_cluster, tenant_class and a deterministic id. Schema: usage-event.v1.json.1. GPU Hours
vcluster_gpu_hours
- Source
- Pod
nvidia.com/gpu,amd.com/gpu,nvidia.com/mig-*andnvidia.com/gpu.sharedrequests, and GPUs allocated through DRA ResourceClaims; node labels (or DRAproductName) for the GPU model - Unit
- gpu-hours
- Quantity
- devices × GPU fraction × seconds running in the window / 3600
Billing starts at the first container start (image pulls and scheduling are free) and stops when the last container finishes. Pods that started or finished mid-window are billed for exactly the seconds they ran.
Fractional GPUs: a MIG slice counts as slices/7 of a GPU (A30: /4), e.g. mig-1g.10gb = 1/7. This works for the MIG mixed strategy (nvidia.com/mig-* resources), the single strategy (slices exposed as nvidia.com/gpu, product label ending in -MIG-<profile>) and GKE GPU partitions (cloud.google.com/gke-gpu-partition-size). On time-sliced nodes (nvidia.com/gpu.replicas > 1, or GKE cloud.google.com/gke-max-shared-clients-per-gpu with time-sharing or MPS) a device counts as 1/replicas. A DRA claim counts the devices allocated to it (a MIG device at its profile's fraction), split across the pods it is reserved for. The SKU becomes <model>-mig-<profile> or <model>-timeslice-<n>, always against the base GPU model (the -SHARED and -MIG- label suffixes are dropped).
Dedicated and private nodes: the node's GPU capacity is billed in whole GPUs with billing_mode=dedicated_node (control plane cluster nodes reserved for the tenant) or billing_mode=private_node (vCluster Private Nodes, Auto Nodes and Standalone, metered through the tenant cluster's own API): MIG slices and time-slicing replicas are converted back, so 8 GPUs shared 4 ways bill as 8, not 32. Pods on those nodes are not billed again. See Dedicated and Private Nodes.
Fair metering: GPU time on a node that is NotReady or carries an UNHEALTHY_NODE_TAINTS taint is emitted as vcluster_gpu_downtime_hours instead. So is a dedicated or private node's GPU time while its GPUs are installed (GPU Feature Discovery's nvidia.com/gpu.count) but not allocatable, for example while the driver rebuilds after a reboot or a MIG change.
Dimensions
| Key | Meaning |
|---|---|
sku | GPU model (or the node's SKU label), plus the fractional profile |
gpu_type | Normalized GPU model |
gpu_profile | full, mig-<profile>, timeslice-<n> |
capacity_type | on-demand, spot, preemptible, reserved |
billing_mode | shared, dedicated_node or private_node |
zone | Node zone |
2. CPU Core-Hours
vcluster_cpu_core_hours
- Source
CPU_MEMORY_BASIS: metrics-server usage, container requests, or the larger of both; node capacity for dedicated nodes- Unit
- core-hours
- Quantity
- cores × seconds in the window / 3600
usage samples metrics-server for the live window only and is never extrapolated into backfilled windows. requests is computed from pod lifetimes, so it is backfilled exactly. max bills the larger of usage and requests per pod.
The tenant cluster's own control plane pods (app=vcluster) are covered by instance hours and excluded unless METER_CONTROL_PLANE=true.
Dimensions
| Key | Meaning |
|---|---|
capacity_type | Capacity type of the node |
billing_mode | shared, dedicated_node or private_node |
namespace | Namespace inside the tenant cluster (with METER_BY_NAMESPACE) |
3. Memory GiB-Hours
vcluster_memory_gb_hours
- Source
- As CPU: usage, requests or max; node capacity for dedicated nodes
- Unit
- gib-hours
- Quantity
- GiB × seconds / 3600
1 GiB = 10243 bytes.
Dimensions
| Key | Meaning |
|---|---|
capacity_type | Capacity type of the node |
billing_mode | shared, dedicated_node or private_node |
4. Dedicated Node Hours
vcluster_private_node_hours
- Source
- Nodes dedicated to the tenant cluster (see Dedicated nodes)
- Unit
- node-hours
- Quantity
- seconds the node existed and was healthy in the window / 3600, one event per node
Downtime on an unhealthy node is not billed.
Dimensions
| Key | Meaning |
|---|---|
sku | SKU label, instance type, or <n>x<GPU model> |
capacity_type | Capacity type of the node |
node | Node name (also resource_id) |
instance_type | node.kubernetes.io/instance-type |
5. Storage GiB-Hours
vcluster_storage_gb_hours
- Source
- Bound PVCs synced from the tenant cluster (provisioned capacity, else the request)
- Unit
- gib-hours
- Quantity
- GiB × seconds bound in the window / 3600
The tenant cluster's own data volume is excluded unless METER_CONTROL_PLANE=true.
Dimensions
| Key | Meaning |
|---|---|
sku | Storage class |
6. Tenant Cluster Hours
vcluster_instance_hours
- Source
- Tenant cluster control plane readiness
- Unit
- hours
- Quantity
- seconds ready in the window / 3600
No charge while a tenant cluster is provisioning or asleep (no ready replica).
Dimensions
| Key | Meaning |
|---|---|
tenant_class | Tenant class |
7. LoadBalancer Hours
vcluster_lb_hours
- Source
- Services of type LoadBalancer with an assigned address
- Unit
- hours
- Quantity
- load balancers × seconds / 3600
A load balancer that is still pending (no IP) is not billed.
8. Network Egress GiB
vcluster_network_egress_gb
- Source
- Prometheus:
increase(container_network_transmit_bytes_total{namespace})over the window, evaluated at the window end, or yourEGRESS_QUERY - Unit
- gib
- Quantity
- bytes / 10243
The default counts all transmitted bytes. To bill only internet egress, point EGRESS_QUERY at flow metrics from your CNI or fabric.
9. GPU Utilization (informational)
vcluster_gpu_utilization
- Source
- DCGM exporter via Prometheus (
DCGM_FI_DEV_GPU_UTIL; for GPUs in MIG mode, which DCGM reports per instance,DCGM_FI_PROF_GR_ENGINE_ACTIVE), matched on the workload'snamespaceor, where Prometheus renames it (kube-prometheus-stack),exported_namespace.GPU_UTIL_QUERYreplaces the query for other exporters;noneturns it off. - Unit
- util-hours
- Quantity
- average utilization % × hours, per GPU model
Informational: delivered to the ledger, webhooks and exports, never to Stripe or Metronome (Lago keeps it for v0.1 plans).
Dimensions
| Key | Meaning |
|---|---|
gpu_type | GPU model |
10. GPU Downtime Hours (informational)
vcluster_gpu_downtime_hours
- Source
- Allocated GPU time on NotReady or unhealthy nodes
- Unit
- gpu-hours
- Quantity
- as GPU hours
An auditable record of the time tenants were not charged for, never sent to billing backends.
Dimensions
| Key | Meaning |
|---|---|
gpu_type | GPU model |
node | Node |
Custom metrics
Declare operator metrics with CUSTOM_METRICS="code:unit[:key|key],..." (for example inference_output_tokens:tokens:model|region) and push closed windows to POST /api/v1/events. Billing backends create meters and billable metrics for them at startup.