Troubleshooting
Common issues you might encounter when running vBilling, with step-by-step solutions.
No Tenant Clusters Discovered
"[discovery] found 0 vCluster(s)"
Symptoms
vBilling starts but reports zero tenant clusters. No customers or events appear in Lago. The controller log shows [discovery] found 0 vCluster(s) on every reconcile cycle.
Cause
vBilling discovers tenant clusters by scanning for StatefulSets and Deployments with the label app=vcluster. If your tenant clusters use different labels, or if RBAC prevents listing these resources, discovery will find nothing.
Fix Check that tenant clusters have the expected labels
# List StatefulSets with the vcluster label
kubectl get statefulsets --all-namespaces -l app=vcluster
# Should list your tenant clusters, e.g.:
# NAMESPACE NAME READY AGE
# vcluster-team-alpha my-vcluster 1/1 3d
# Also check Deployments (some vCluster configs use Deployments)
kubectl get deployments --all-namespaces -l app=vclusterFix Check RBAC permissions
The vBilling ServiceAccount needs permission to list StatefulSets, Deployments, Pods, Services, PVCs, Nodes, and PodMetrics across all namespaces (or the specific namespaces configured in WATCH_NAMESPACES).
# Verify the ServiceAccount can list StatefulSets
kubectl auth can-i list statefulsets --as=system:serviceaccount:vbilling-system:vbilling --all-namespaces
# Should return: yes
# If RBAC is the issue, check the ClusterRole
kubectl get clusterrolebinding | grep vbillingFix Verify Platform API access (optional)
If using vCluster Platform, vBilling tries the Platform API first (management.loft.sh/v1 VirtualClusterInstance resources). If this CRD does not exist, vBilling falls back to workload scanning. Check the logs for:
[discovery] Platform API not available, falling back to StatefulSet scanning: ...This is normal if you are not running vCluster Platform -- the fallback to StatefulSet scanning is the expected behavior.
Fix Check namespace filter
# If WATCH_NAMESPACES is set, only those namespaces are scanned
# Make sure your vCluster namespace is included
kubectl get configmap -n vbilling-system -o yaml | grep WATCH_NAMESPACESMetrics Unavailable
"[metrics] warning: cannot get pod metrics for <namespace>"
Symptoms
vBilling discovers tenant clusters but CPU and memory metrics are always zero. The log shows warnings about failing to get pod metrics.
Cause
vBilling requires the Kubernetes metrics-server to be installed and functioning. Without it, the PodMetrics API is unavailable.
Fix Install metrics-server
# Check if metrics-server is running
kubectl get pods -n kube-system -l k8s-app=metrics-server
# If not installed, install it:
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# For development clusters (minikube, kind) with self-signed certs:
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
kubectl patch deployment metrics-server -n kube-system \
--type=json \
-p '[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'Fix Verify the metrics API works
# Test the metrics API directly
kubectl top pods -n <vcluster-namespace>
# Should show CPU and memory usage:
# NAME CPU(cores) MEMORY(bytes)
# my-vcluster-0 25m 128Mi
# If this fails, metrics-server is not workingGPU Type Empty
GPU events show gpu_type "unknown"
Symptoms
GPU hours are being reported, but the gpu_type property in Lago events shows "unknown" instead of the actual GPU model.
Cause
vBilling reads the GPU type from node labels. If none of the expected labels are present, it falls back to "unknown".
Fix Check which GPU labels exist on your nodes
# Check all GPU-related labels on your GPU nodes
kubectl get nodes -l nvidia.com/gpu -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.metadata.labels}{"\n"}{end}'
# vBilling checks these labels in order:
# 1. nvidia.com/gpu.product (NVIDIA GPU Operator)
# 2. nvidia.com/gpu.machine (alternative)
# 3. accelerator (GKE generic)
# 4. cloud.google.com/gke-accelerator (GKE specific)
# 5. k8s.amazonaws.com/accelerator (EKS)
# 6. node.kubernetes.io/instance-type (fallback)Fix Install the NVIDIA GPU Operator
The easiest way to get GPU labels is to install the NVIDIA GPU Operator, which automatically labels nodes with nvidia.com/gpu.product:
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm install gpu-operator nvidia/gpu-operator --namespace gpu-operator --create-namespaceFix Manually label your GPU nodes
# If you know the GPU model, label manually
kubectl label node <gpu-node> nvidia.com/gpu.product="NVIDIA-A100-SXM4-80GB"Dedicated Nodes Not Detected
Dedicated nodes labeled but not appearing in billing events
Symptoms
You have labeled control plane cluster nodes with vcluster.loft.sh/managed-by but vBilling does not bill them as dedicated nodes. No vcluster_private_node_hours events with billing_mode=dedicated_node appear. (For nodes that join a tenant cluster through vCluster Private Nodes, see Private Nodes Not Metered.)
Cause
The most common cause is a label value mismatch between the node label and the vCluster namespace name. vBilling tries multiple label values, but they must correspond to the namespace.
Fix Understand the label matching logic
For a vCluster in namespace vcluster-team-gpu, vBilling searches for nodes with these selectors:
# Selector 1: full namespace name
vcluster.loft.sh/managed-by=vcluster-team-gpu
vcluster.loft.sh/cluster=vcluster-team-gpu
# Selector 2: namespace with "vcluster-" prefix stripped
vcluster.loft.sh/managed-by=team-gpu
vcluster.loft.sh/cluster=team-gpu
# Selector 3: custom VCLUSTER_NODE_LABEL (if set)Fix Verify the node labels
# Check what value your nodes are actually labeled with
kubectl get nodes -l vcluster.loft.sh/managed-by -o custom-columns="NAME:.metadata.name,MANAGED-BY:.metadata.labels.vcluster\.loft\.sh/managed-by"
# NAME MANAGED-BY
# gpu-node-001 team-gpu
# Make sure the label value matches one of the selectors above
# For namespace "vcluster-team-gpu", either of these will work:
kubectl label node gpu-node-001 vcluster.loft.sh/managed-by=vcluster-team-gpu
# OR
kubectl label node gpu-node-001 vcluster.loft.sh/managed-by=team-gpuFix Test the exact selector vBilling uses
# Simulate what vBilling does internally:
kubectl get nodes -l vcluster.loft.sh/managed-by=team-gpu
# If this returns your nodes, vBilling will find them
# If empty, the label value doesn't matchPrivate Nodes Not Metered
Nodes joined with vCluster Private Nodes produce no events
Symptoms
A tenant cluster has private nodes, but no vcluster_private_node_hours events with billing_mode=private_node appear, or GET /api/v1/status lists the tenant cluster under failed tenant APIs.
Cause
vBilling reads private nodes through the tenant cluster's own API, with the kubeconfig in the Secret vbilling-kubeconfig in the tenant cluster's namespace. Without that Secret, or with the tenant cluster's control plane down, there is nothing to read.
Fix Export the read-only kubeconfig
Add exportKubeConfig.additionalSecrets with the vbilling-reader service account to the tenant cluster's vcluster.yaml (see Private Nodes), then check the Secret exists:
kubectl -n <tenant-namespace> get secret vbilling-kubeconfigFix The tenant cluster's control plane exits
you are trying to use a vCluster pro feature 'Private Nodes' (vcp-distro-private-nodes) that is not available means the tenant cluster has no license: register it with vCluster Platform (vcluster platform add vcluster <name> -n <namespace>). denied by loft access key scope means the CLI's access key is scoped and cannot register tenant clusters; log in with a key that can. vCluster 0.36.2 and later need Platform 4.11.3 or later.
Fix GPU-hours missing right after a node reboot
While a driver container rebuilds (after a reboot, a spot restart or a MIG change), a Ready node advertises no GPUs. vBilling bills the node and records its installed GPUs (GPU Feature Discovery's nvidia.com/gpu.count) as vcluster_gpu_downtime_hours with the reason "GPUs installed but not allocatable" until they are allocatable again.
GPU Utilization or Energy Missing
No vcluster_gpu_utilization events, or a DCGM metric reads zero or too high
Cause
DCGM behaves differently from what a query might assume:
- For GPUs in MIG mode DCGM reports no
DCGM_FI_DEV_GPU_UTIL. The default query falls back to each instance'sDCGM_FI_PROF_GR_ENGINE_ACTIVE; a customGPU_UTIL_QUERYneeds the same fallback. - Device-level fields (power, energy, framebuffer) repeat on every MIG instance: summing them bills the GPU once per slice. Keep one series per GPU with
max by (modelName, UUID, ...)first. - The exporter refreshes every 30 seconds by default (
DCGM_EXPORTER_INTERVAL), soincrease()over a 30 second window can read zero. Integrate a gauge instead:avg_over_time(...[{{window}}]) * {{window_seconds}}. - DCGM on private nodes runs inside the tenant cluster. Ship it to your Prometheus with an agent that adds an external label naming the tenant cluster, and select on it with
{{vcluster}}.
The GKE example has working queries for all four cases.
Lago Connection Failed
"lago API error" or connection refused
Symptoms
vBilling logs show errors like HTTP POST /api/v1/events/batch: connection refused or lago API error 401. The bootstrap step may fail with a warning.
Fix Check the Lago URL
# If running in-cluster, verify the service exists
kubectl get svc -n lago-system
# Test connectivity from inside the cluster
kubectl run test-curl --rm -it --image=curlimages/curl -- \
curl -s http://lago-api.lago-system.svc.cluster.local:3000/api/v1/billable_metrics \
-H "Authorization: Bearer $LAGO_API_KEY"
# If running locally, check that Lago is running
curl -s http://localhost:3000/api/v1/billable_metrics \
-H "Authorization: Bearer $LAGO_API_KEY"Fix Check the API key
# Verify the API key is correct
curl -s -o /dev/null -w "%{http_code}" \
http://localhost:3000/api/v1/billable_metrics \
-H "Authorization: Bearer $LAGO_API_KEY"
# 200 = API key is valid
# 401 = API key is wrong
# 000 = Connection failed (Lago not running or wrong URL)Fix Check the Kubernetes secret
# If using lago.existingSecret, verify the secret exists and has the right key
kubectl get secret vbilling-lago-secret -n vbilling-system -o jsonpath='{.data.api-key}' | base64 -dEvents Sent But $0.00 in Lago
Usage events appear in Lago but invoices show $0.00
Symptoms
vBilling is successfully sending events (you can see them in the Lago UI under the customer's events tab), but the current usage and invoices all show $0.00.
Cause
When vBilling bootstraps the vcluster-standard plan, all charge amounts default to $0.00 per unit. This is by design -- vBilling creates the plan structure, but you must configure the actual pricing.
Fix Configure pricing in Lago
- Open the Lago UI (e.g.,
http://localhost:80) - Navigate to Plans in the left sidebar
- Click on vCluster Standard
- For each charge, click the edit icon and set a non-zero price per unit
- Save the plan
See Configuration: Lago Pricing for recommended pricing for each metric.
Duplicate Subscriptions
Duplicate or missing subscriptions after controller restart
Symptoms
After restarting the vBilling controller, you see warnings about subscription creation failing, or duplicate subscriptions appear in Lago.
Cause
vBilling is designed to be idempotent on restart. When it discovers a vCluster, it first probes for an existing subscription using the GetCurrentUsage API call. If the subscription already exists, Lago returns usage data and vBilling reuses the subscription. If the probe fails (404), vBilling creates a new subscription.
How it works
# vBilling's subscription detection logic:
# 1. Try GetCurrentUsage(customerID, subscriptionID)
# 2. If success: subscription exists, reuse it
# 3. If error: subscription does not exist, create it
[controller] subscription sub-vcluster-ns-name already exists, reusing
# OR
[controller] created subscription sub-vcluster-ns-name -> plan vcluster-standardFix This is normally self-healing
If you see "subscription already exists" warnings, the controller is correctly detecting and reusing existing subscriptions. No action needed.
Fix Clean up duplicates in Lago
If actual duplicates exist (multiple subscriptions for the same vCluster), you can terminate the extra ones via the Lago UI or API:
# List subscriptions for a customer
curl -s http://localhost:3000/api/v1/subscriptions\?external_customer_id=vcluster-ns-name \
-H "Authorization: Bearer $LAGO_API_KEY" | jq '.subscriptions[].external_id'
# Terminate a duplicate subscription
curl -X DELETE http://localhost:3000/api/v1/subscriptions/<duplicate-external-id> \
-H "Authorization: Bearer $LAGO_API_KEY"CORS Issues with Dashboard
Browser console shows CORS errors when accessing Lago UI or vBilling dashboard
Symptoms
When accessing the Lago UI or any vBilling dashboard from a browser, you see errors like: Access to fetch at 'http://localhost:3000/...' from origin 'http://localhost:8081' has been blocked by CORS policy.
Cause
The browser enforces Cross-Origin Resource Sharing (CORS) restrictions when the frontend and API are served from different origins (different ports or domains).
Fix Configure Lago CORS settings
In the Lago Docker Compose setup, ensure the LAGO_FRONT_URL environment variable matches your frontend URL:
# In lago/.env or docker-compose.yml
LAGO_FRONT_URL=http://localhost
LAGO_API_URL=http://localhost:3000Fix Use a reverse proxy
For production, put both the API and frontend behind the same domain using a reverse proxy (nginx, Traefik, etc.) so CORS is not an issue.
Fix Port-forward with matching origins
# If accessing in-cluster Lago from your machine:
kubectl port-forward svc/lago-front -n lago-system 80:80 &
kubectl port-forward svc/lago-api -n lago-system 3000:3000 &
# Then access the frontend at http://localhostGeneral Debugging Tips
[component] prefix convention. Key prefixes to filter:
[discovery]-- vCluster discovery events[controller]-- Reconciliation and event sending[metrics]-- Resource collection (CPU, memory, GPU, etc.)[bootstrap]-- Lago metric and plan setup[lago]-- Raw API calls to Lago
Useful commands
# Stream vBilling logs
kubectl logs -n vbilling-system -l app=vbilling -f
# Filter for errors only
kubectl logs -n vbilling-system -l app=vbilling | grep -i "error\|warning\|fatal"
# Check vBilling pod status
kubectl get pods -n vbilling-system
# Describe pod for events and conditions
kubectl describe pod -n vbilling-system -l app=vbilling
# Check cluster-wide vCluster StatefulSets
kubectl get statefulsets --all-namespaces -l app=vcluster
# Check Lago health
curl -s http://localhost:3000/api/v1/billable_metrics \
-H "Authorization: Bearer $LAGO_API_KEY" | jq '.billable_metrics | length'
# Should return: 9
# Check which customers exist in Lago
curl -s http://localhost:3000/api/v1/customers \
-H "Authorization: Bearer $LAGO_API_KEY" | jq '.customers[] | {external_id, name}'