Meter tenant clusters to the second, keep a durable, verifiable ledger, and ship reconcilable usage to Stripe, Metronome, Lago, or your own pipeline.
New in v0.2: Stripe and Metronome, a verifiable ledger, vCluster Private Nodes, and DCGM through your own Prometheus. See it on GKE with shared and private GPUs →
Adapter-agnostic
vBilling is the pipe, not the billing engine. Meter tenant clusters once and route events to whichever billing adapter you run.
How it works
vBilling collects usage and ships it to your billing adapter. Your adapter handles pricing, plans, and invoicing. You stay in control of what anything costs.
01 · Choose
Pick the billing backend you already run: Stripe, Metronome (with Stripe for payments), Lago, or your own pipeline through a signed webhook. Send to several at once.
02 · Configure
Define metrics, plans, and per-unit rates in your billing platform. vBilling never decides what anything costs. Your platform team owns the price sheet.
03 · Connect
Deploy vBilling and point it at your backend. Tenant clusters are auto-discovered, and usage starts flowing from the next closed window.
Features
vBilling meters shared and dedicated GPU capacity, commits it to a durable ledger, and delivers it to every billing backend you run, exactly once.
Each window is fsynced to a hash-chained ledger before delivery. Restarts and outages are backfilled: no usage lost, nothing billed twice.
Fan out to several destinations at once. Each has its own cursor, so a billing API outage never blocks your data platform feed.
Recompute any day from the ledger and compare it with what Stripe or Metronome recorded, dead letters accounted for.
Detects H100, L40S, A100 and more from node labels. Stripe gets a meter per SKU; Metronome and Lago price on the sku dimension.
MIG slices and time-sliced GPUs are metered at their fraction, so shared GPUs can be sold by the slice.
No charges for image pulls, provisioning, sleeping clusters, or GPU time on unhealthy nodes. Downtime is recorded as an audit trail.
Bill whole bare-metal nodes by SKU or instance type, without double-billing the pods that run on them.
GPU machines that join a tenant cluster directly are read through its own API, with a read-only kubeconfig vCluster exports, and billed whole. Auto Nodes and Standalone too.
DCGM GPU utilization from any Prometheus, Thanos or Mimir, MIG instances included, or from vCluster Platform fleet observability.
Define your own metrics in PromQL, such as inference tokens or GPU energy in kWh, and bill them like built-in usage.
Every event carries on-demand, spot, preemptible or reserved, so your billing platform applies the discount and quantities stay physical.
Stripe dunning and Metronome spend alerts label, annotate, or soft-suspend tenant clusters through an admission policy.
Push inference tokens, Slurm accounting, or storage usage into the same durable pipeline as collector events.
One vBilling per control plane cluster: every event, ledger record and billing state stays in its region.
Finds tenant clusters and vCluster Platform projects. Customers, meters and metrics are created in your backend automatically.
Architecture
vBilling runs as a controller in your Control Plane Cluster. It watches tenant clusters, collects capacity and usage metrics, and streams events to the billing adapter you choose.
Quick start
Pick your billing adapter below. Each tab walks through deploy, install, and pricing for that platform.
Spin up the open-source billing engine with Docker Compose. UI at :8081, API at :3000 (vBilling's own dashboard uses :8080).
# Clone vBilling and start Lago
git clone https://github.com/vClusterLabs-Experiments/vbilling.git
cd vbilling/deploy/lago
openssl genrsa 2048 > lago_rsa.key
echo "LAGO_RSA_PRIVATE_KEY=$(base64 -i lago_rsa.key | tr -d '\n')" > .env
docker compose --env-file .env up -d
Deploy the controller via Helm. Point it at your Lago instance; the API key stays in a Secret.
kubectl create namespace vbilling-system
kubectl -n vbilling-system create secret generic vbilling-lago --from-file=api-key=./lago-api-key
helm upgrade --install vbilling deploy/helm/vbilling --namespace vbilling-system \
--set adapters='{lago}' \
--set lago.apiURL=http://lago-api:3000 \
--set lago.existingSecret=vbilling-lago
vBilling creates a skeleton plan with $0 pricing. Set your rates in the Lago UI or API.
# In Lago UI → Plans → vCluster Standard
# Set prices per metric:
CPU Core-Hours $0.065
Memory GB-Hours $0.009
GPU Hours (H100) $4.50
Storage GB-Hours $0.0002
Network Egress GB $0.09
Node Hours $25.00
LoadBalancer Hours $0.025
Use a test-mode key first. vBilling creates one Billing Meter per metric, and one per SKU as it sees them.
kubectl create namespace vbilling-system
kubectl -n vbilling-system create secret generic vbilling-stripe \
--from-file=api-key=./stripe-test-key
helm upgrade --install vbilling deploy/helm/vbilling -n vbilling-system \
--set adapters='{stripe}' \
--set stripe.existingSecret=vbilling-stripe \
--set region=ap-southeast-2
Add metered prices to the meters you charge for. Tag them metadata[vbilling_plan]=vcluster-standard and set stripe.autoSubscribe=true to subscribe new tenants automatically.
# Meters created by vBilling
vcluster_gpu_hours__nvidia_h100_80gb_hbm3
vcluster_gpu_hours__nvidia_l40s
vcluster_private_node_hours__bm_gpu_h100_8
vcluster_cpu_core_hours vcluster_instance_hours ...
Point a Stripe webhook at /webhooks/stripe. Failed payments mark tenants delinquent; unpaid subscriptions can soft-suspend them.
# add key webhook-secret (whsec_...) to the vbilling-stripe Secret, then
--set enforcement.mode=annotate
Metronome rates usage; vBilling creates each tenant's Stripe customer and links it, so Stripe collects payment, invoices and tax.
kubectl -n vbilling-system create secret generic vbilling-metronome --from-file=api-token=./metronome-token
kubectl -n vbilling-system create secret generic vbilling-stripe --from-file=api-key=./stripe-key
helm upgrade --install vbilling deploy/helm/vbilling -n vbilling-system \
--set adapters='{metronome}' \
--set metronome.existingSecret=vbilling-metronome --set metronome.rateCard=payg-aud \
--set metronome.stripeLink=true --set stripe.existingSecret=vbilling-stripe
vBilling creates SUM billable metrics grouped on region, sku, capacity_type, billing_mode and tenant_cluster. Add rate card products with pricing group keys.
# Metronome product on vcluster_gpu_hours
pricing_group_key: [region, sku, capacity_type]
presentation_group_key: [tenant_cluster]
Compare the ledger with what Metronome recorded for the previous UTC day.
curl -s localhost:8080/api/v1/reconcile | jq '.results[] | {destination, ok}'
vBilling is structured so any HTTP billing API can plug in. Implement the Destination interface in a small Go package, rebuild the image, and point vBilling at your endpoint.
Or skip Go entirely: the webhook destination streams signed CloudEvents batches to any HTTP endpoint, and the ingest API accepts usage from your own services.
Metrics
Collected automatically from the Kubernetes API, metrics-server, and Prometheus. Custom sources and transforms extend the defaults without forking vBilling.
| Metric | Source | Granularity |
|---|---|---|
| Node hours | Node watch | Per dedicated node |
| CPU core-hours | Node capacity | Full node capacity |
| Memory GB-hours | Node capacity | Full node capacity |
| GPU hours (by SKU) H100 / A100 / T4 | Node labels | Per GPU SKU |
| GPU utilization | DCGM via Prometheus | Per GPU % |
| Storage GB-hours | PVC sizes | Per PVC |
| Network egress GB | CNI / Prometheus | Per tenant |
| LoadBalancer hours | Service count | Per LB service |
| Control plane hours | tenant cluster watch | 1 per cluster |
For AI Clouds
If you run an AI cloud on Kubernetes, vBilling gives you the billing pipe, without forcing a billing backend on you.
AI Clouds give each customer a tenant cluster with dedicated GPU nodes. vBilling detects the hardware, meters node capacity by GPU SKU, and streams events into whichever billing adapter you run.