Troubleshooting

Common issues you might encounter when running vBilling, with step-by-step solutions.

No Tenant Clusters Discovered

"[discovery] found 0 vCluster(s)"

Symptoms

vBilling starts but reports zero tenant clusters. No customers or events appear in Lago. The controller log shows [discovery] found 0 vCluster(s) on every reconcile cycle.

Cause

vBilling discovers tenant clusters by scanning for StatefulSets and Deployments with the label app=vcluster. If your tenant clusters use different labels, or if RBAC prevents listing these resources, discovery will find nothing.

Fix Check that tenant clusters have the expected labels

# List StatefulSets with the vcluster label kubectl get statefulsets --all-namespaces -l app=vcluster # Should list your tenant clusters, e.g.: # NAMESPACE NAME READY AGE # vcluster-team-alpha my-vcluster 1/1 3d # Also check Deployments (some vCluster configs use Deployments) kubectl get deployments --all-namespaces -l app=vcluster

Fix Check RBAC permissions

The vBilling ServiceAccount needs permission to list StatefulSets, Deployments, Pods, Services, PVCs, Nodes, and PodMetrics across all namespaces (or the specific namespaces configured in WATCH_NAMESPACES).

# Verify the ServiceAccount can list StatefulSets kubectl auth can-i list statefulsets --as=system:serviceaccount:vbilling-system:vbilling --all-namespaces # Should return: yes # If RBAC is the issue, check the ClusterRole kubectl get clusterrolebinding | grep vbilling

Fix Verify Platform API access (optional)

If using vCluster Platform, vBilling tries the Platform API first (management.loft.sh/v1 VirtualClusterInstance resources). If this CRD does not exist, vBilling falls back to workload scanning. Check the logs for:

[discovery] Platform API not available, falling back to StatefulSet scanning: ...

This is normal if you are not running vCluster Platform -- the fallback to StatefulSet scanning is the expected behavior.

Fix Check namespace filter

# If WATCH_NAMESPACES is set, only those namespaces are scanned # Make sure your vCluster namespace is included kubectl get configmap -n vbilling-system -o yaml | grep WATCH_NAMESPACES

Metrics Unavailable

"[metrics] warning: cannot get pod metrics for <namespace>"

Symptoms

vBilling discovers tenant clusters but CPU and memory metrics are always zero. The log shows warnings about failing to get pod metrics.

Cause

vBilling requires the Kubernetes metrics-server to be installed and functioning. Without it, the PodMetrics API is unavailable.

Fix Install metrics-server

# Check if metrics-server is running kubectl get pods -n kube-system -l k8s-app=metrics-server # If not installed, install it: kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml # For development clusters (minikube, kind) with self-signed certs: kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml kubectl patch deployment metrics-server -n kube-system \ --type=json \ -p '[{"op":"add","path":"/spec/template/spec/containers/0/args/-","value":"--kubelet-insecure-tls"}]'

Fix Verify the metrics API works

# Test the metrics API directly kubectl top pods -n <vcluster-namespace> # Should show CPU and memory usage: # NAME CPU(cores) MEMORY(bytes) # my-vcluster-0 25m 128Mi # If this fails, metrics-server is not working

GPU Type Empty

GPU events show gpu_type "unknown"

Symptoms

GPU hours are being reported, but the gpu_type property in Lago events shows "unknown" instead of the actual GPU model.

Cause

vBilling reads the GPU type from node labels. If none of the expected labels are present, it falls back to "unknown".

Fix Check which GPU labels exist on your nodes

# Check all GPU-related labels on your GPU nodes kubectl get nodes -l nvidia.com/gpu -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.metadata.labels}{"\n"}{end}' # vBilling checks these labels in order: # 1. nvidia.com/gpu.product (NVIDIA GPU Operator) # 2. nvidia.com/gpu.machine (alternative) # 3. accelerator (GKE generic) # 4. cloud.google.com/gke-accelerator (GKE specific) # 5. k8s.amazonaws.com/accelerator (EKS) # 6. node.kubernetes.io/instance-type (fallback)

Fix Install the NVIDIA GPU Operator

The easiest way to get GPU labels is to install the NVIDIA GPU Operator, which automatically labels nodes with nvidia.com/gpu.product:

helm repo add nvidia https://helm.ngc.nvidia.com/nvidia helm install gpu-operator nvidia/gpu-operator --namespace gpu-operator --create-namespace

Fix Manually label your GPU nodes

# If you know the GPU model, label manually kubectl label node <gpu-node> nvidia.com/gpu.product="NVIDIA-A100-SXM4-80GB"

Dedicated Nodes Not Detected

Dedicated nodes labeled but not appearing in billing events

Symptoms

You have labeled control plane cluster nodes with vcluster.loft.sh/managed-by but vBilling does not bill them as dedicated nodes. No vcluster_private_node_hours events with billing_mode=dedicated_node appear. (For nodes that join a tenant cluster through vCluster Private Nodes, see Private Nodes Not Metered.)

Cause

The most common cause is a label value mismatch between the node label and the vCluster namespace name. vBilling tries multiple label values, but they must correspond to the namespace.

Fix Understand the label matching logic

For a vCluster in namespace vcluster-team-gpu, vBilling searches for nodes with these selectors:

# Selector 1: full namespace name vcluster.loft.sh/managed-by=vcluster-team-gpu vcluster.loft.sh/cluster=vcluster-team-gpu # Selector 2: namespace with "vcluster-" prefix stripped vcluster.loft.sh/managed-by=team-gpu vcluster.loft.sh/cluster=team-gpu # Selector 3: custom VCLUSTER_NODE_LABEL (if set)

Fix Verify the node labels

# Check what value your nodes are actually labeled with kubectl get nodes -l vcluster.loft.sh/managed-by -o custom-columns="NAME:.metadata.name,MANAGED-BY:.metadata.labels.vcluster\.loft\.sh/managed-by" # NAME MANAGED-BY # gpu-node-001 team-gpu # Make sure the label value matches one of the selectors above # For namespace "vcluster-team-gpu", either of these will work: kubectl label node gpu-node-001 vcluster.loft.sh/managed-by=vcluster-team-gpu # OR kubectl label node gpu-node-001 vcluster.loft.sh/managed-by=team-gpu

Fix Test the exact selector vBilling uses

# Simulate what vBilling does internally: kubectl get nodes -l vcluster.loft.sh/managed-by=team-gpu # If this returns your nodes, vBilling will find them # If empty, the label value doesn't match

Private Nodes Not Metered

Nodes joined with vCluster Private Nodes produce no events

Symptoms

A tenant cluster has private nodes, but no vcluster_private_node_hours events with billing_mode=private_node appear, or GET /api/v1/status lists the tenant cluster under failed tenant APIs.

Cause

vBilling reads private nodes through the tenant cluster's own API, with the kubeconfig in the Secret vbilling-kubeconfig in the tenant cluster's namespace. Without that Secret, or with the tenant cluster's control plane down, there is nothing to read.

Fix Export the read-only kubeconfig

Add exportKubeConfig.additionalSecrets with the vbilling-reader service account to the tenant cluster's vcluster.yaml (see Private Nodes), then check the Secret exists:

kubectl -n <tenant-namespace> get secret vbilling-kubeconfig

Fix The tenant cluster's control plane exits

you are trying to use a vCluster pro feature 'Private Nodes' (vcp-distro-private-nodes) that is not available means the tenant cluster has no license: register it with vCluster Platform (vcluster platform add vcluster <name> -n <namespace>). denied by loft access key scope means the CLI's access key is scoped and cannot register tenant clusters; log in with a key that can. vCluster 0.36.2 and later need Platform 4.11.3 or later.

Fix GPU-hours missing right after a node reboot

While a driver container rebuilds (after a reboot, a spot restart or a MIG change), a Ready node advertises no GPUs. vBilling bills the node and records its installed GPUs (GPU Feature Discovery's nvidia.com/gpu.count) as vcluster_gpu_downtime_hours with the reason "GPUs installed but not allocatable" until they are allocatable again.

GPU Utilization or Energy Missing

No vcluster_gpu_utilization events, or a DCGM metric reads zero or too high

Cause

DCGM behaves differently from what a query might assume:

  • For GPUs in MIG mode DCGM reports no DCGM_FI_DEV_GPU_UTIL. The default query falls back to each instance's DCGM_FI_PROF_GR_ENGINE_ACTIVE; a custom GPU_UTIL_QUERY needs the same fallback.
  • Device-level fields (power, energy, framebuffer) repeat on every MIG instance: summing them bills the GPU once per slice. Keep one series per GPU with max by (modelName, UUID, ...) first.
  • The exporter refreshes every 30 seconds by default (DCGM_EXPORTER_INTERVAL), so increase() over a 30 second window can read zero. Integrate a gauge instead: avg_over_time(...[{{window}}]) * {{window_seconds}}.
  • DCGM on private nodes runs inside the tenant cluster. Ship it to your Prometheus with an agent that adds an external label naming the tenant cluster, and select on it with {{vcluster}}.

The GKE example has working queries for all four cases.

Lago Connection Failed

"lago API error" or connection refused

Symptoms

vBilling logs show errors like HTTP POST /api/v1/events/batch: connection refused or lago API error 401. The bootstrap step may fail with a warning.

Fix Check the Lago URL

# If running in-cluster, verify the service exists kubectl get svc -n lago-system # Test connectivity from inside the cluster kubectl run test-curl --rm -it --image=curlimages/curl -- \ curl -s http://lago-api.lago-system.svc.cluster.local:3000/api/v1/billable_metrics \ -H "Authorization: Bearer $LAGO_API_KEY" # If running locally, check that Lago is running curl -s http://localhost:3000/api/v1/billable_metrics \ -H "Authorization: Bearer $LAGO_API_KEY"

Fix Check the API key

# Verify the API key is correct curl -s -o /dev/null -w "%{http_code}" \ http://localhost:3000/api/v1/billable_metrics \ -H "Authorization: Bearer $LAGO_API_KEY" # 200 = API key is valid # 401 = API key is wrong # 000 = Connection failed (Lago not running or wrong URL)

Fix Check the Kubernetes secret

# If using lago.existingSecret, verify the secret exists and has the right key kubectl get secret vbilling-lago-secret -n vbilling-system -o jsonpath='{.data.api-key}' | base64 -d

Events Sent But $0.00 in Lago

Usage events appear in Lago but invoices show $0.00

Symptoms

vBilling is successfully sending events (you can see them in the Lago UI under the customer's events tab), but the current usage and invoices all show $0.00.

Cause

When vBilling bootstraps the vcluster-standard plan, all charge amounts default to $0.00 per unit. This is by design -- vBilling creates the plan structure, but you must configure the actual pricing.

Fix Configure pricing in Lago

  1. Open the Lago UI (e.g., http://localhost:80)
  2. Navigate to Plans in the left sidebar
  3. Click on vCluster Standard
  4. For each charge, click the edit icon and set a non-zero price per unit
  5. Save the plan

See Configuration: Lago Pricing for recommended pricing for each metric.

i
After updating pricing, new events will be billed at the new rate. Existing events in the current billing period will also be recalculated by Lago.

Duplicate Subscriptions

Duplicate or missing subscriptions after controller restart

Symptoms

After restarting the vBilling controller, you see warnings about subscription creation failing, or duplicate subscriptions appear in Lago.

Cause

vBilling is designed to be idempotent on restart. When it discovers a vCluster, it first probes for an existing subscription using the GetCurrentUsage API call. If the subscription already exists, Lago returns usage data and vBilling reuses the subscription. If the probe fails (404), vBilling creates a new subscription.

How it works

# vBilling's subscription detection logic: # 1. Try GetCurrentUsage(customerID, subscriptionID) # 2. If success: subscription exists, reuse it # 3. If error: subscription does not exist, create it [controller] subscription sub-vcluster-ns-name already exists, reusing # OR [controller] created subscription sub-vcluster-ns-name -> plan vcluster-standard

Fix This is normally self-healing

If you see "subscription already exists" warnings, the controller is correctly detecting and reusing existing subscriptions. No action needed.

Fix Clean up duplicates in Lago

If actual duplicates exist (multiple subscriptions for the same vCluster), you can terminate the extra ones via the Lago UI or API:

# List subscriptions for a customer curl -s http://localhost:3000/api/v1/subscriptions\?external_customer_id=vcluster-ns-name \ -H "Authorization: Bearer $LAGO_API_KEY" | jq '.subscriptions[].external_id' # Terminate a duplicate subscription curl -X DELETE http://localhost:3000/api/v1/subscriptions/<duplicate-external-id> \ -H "Authorization: Bearer $LAGO_API_KEY"

CORS Issues with Dashboard

Browser console shows CORS errors when accessing Lago UI or vBilling dashboard

Symptoms

When accessing the Lago UI or any vBilling dashboard from a browser, you see errors like: Access to fetch at 'http://localhost:3000/...' from origin 'http://localhost:8081' has been blocked by CORS policy.

Cause

The browser enforces Cross-Origin Resource Sharing (CORS) restrictions when the frontend and API are served from different origins (different ports or domains).

Fix Configure Lago CORS settings

In the Lago Docker Compose setup, ensure the LAGO_FRONT_URL environment variable matches your frontend URL:

# In lago/.env or docker-compose.yml LAGO_FRONT_URL=http://localhost LAGO_API_URL=http://localhost:3000

Fix Use a reverse proxy

For production, put both the API and frontend behind the same domain using a reverse proxy (nginx, Traefik, etc.) so CORS is not an issue.

Fix Port-forward with matching origins

# If accessing in-cluster Lago from your machine: kubectl port-forward svc/lago-front -n lago-system 80:80 & kubectl port-forward svc/lago-api -n lago-system 3000:3000 & # Then access the frontend at http://localhost

General Debugging Tips

i
Enable verbose logging: vBilling logs to stdout with the [component] prefix convention. Key prefixes to filter:
  • [discovery] -- vCluster discovery events
  • [controller] -- Reconciliation and event sending
  • [metrics] -- Resource collection (CPU, memory, GPU, etc.)
  • [bootstrap] -- Lago metric and plan setup
  • [lago] -- Raw API calls to Lago

Useful commands

# Stream vBilling logs kubectl logs -n vbilling-system -l app=vbilling -f # Filter for errors only kubectl logs -n vbilling-system -l app=vbilling | grep -i "error\|warning\|fatal" # Check vBilling pod status kubectl get pods -n vbilling-system # Describe pod for events and conditions kubectl describe pod -n vbilling-system -l app=vbilling # Check cluster-wide vCluster StatefulSets kubectl get statefulsets --all-namespaces -l app=vcluster # Check Lago health curl -s http://localhost:3000/api/v1/billable_metrics \ -H "Authorization: Bearer $LAGO_API_KEY" | jq '.billable_metrics | length' # Should return: 9 # Check which customers exist in Lago curl -s http://localhost:3000/api/v1/customers \ -H "Authorization: Bearer $LAGO_API_KEY" | jq '.customers[] | {external_id, name}'
✓
If you encounter an issue not covered here, please open a GitHub issue with your vBilling logs and cluster details.