Helm Install
Complete a Helm install of CubeSandbox on an existing Kubernetes cluster.
Install process
① Cluster ready
↓
② Label nodes (and role taints)
↓
③ Prepare values
↓
④ helm upgrade --install
↓
⑤ Verify1. Prerequisites
| Item | Requirement |
|---|---|
| Kubernetes | ≥ v1.24+ |
| Tools | kubectl, Helm v3.10+ |
| Storage | A usable StorageClass in the cluster |
| Image pull | Nodes can reach the Internet to pull images |
Node roles
CubeSandbox uses labels to distinguish two kinds of nodes:
| Role | Suggested size | Purpose |
|---|---|---|
| Control-plane node | ≥ 1; production suggests 3+; ≥ 4C8G | Runs Cube's control-plane components |
| Compute node | ≥ 1; suggest 16C32G+ | Runs sandboxes |
Recommendation
Use dedicated machines for compute nodes.
What is PVM? (Technical background)
PVM (Pagetable-based Virtual Machine) is a pagetable-based nested virtualization framework built on top of KVM. Unlike traditional nested virtualization, PVM does not rely on the host hypervisor exposing Intel VT-x / AMD-V hardware virtualization extensions to the guest. Instead, it performs privilege-level switching and memory virtualization inside the guest kernel layer through shared-memory regions and shadow page tables, and is fully transparent to the host hypervisor.
PVM was originally proposed in the paper PVM: Efficient Shadow Paging for Deploying Secure Containers in Cloud-native Environment. Tencent Cloud has since done substantial feature and performance work on top of it, fixed many bugs, and open-sourced the upstreamed kernel changes in OpenCloudOS Kernel for the community. Over the years we have deployed a large fleet of PVM instances in Tencent Cloud production, and its reliability has been validated in production.
2. Confirm the cluster is ready
kubectl get nodes -o wide
kubectl get storageclass # Must not be empty
helm version --short # ≥ v3.10Note the node names you will use later. Confirm at least one node for the control plane and another for the compute plane.
3. Label nodes and roles
3.1 Multi-node (control / compute split, recommended)
Under normal circumstances, add at least 2 machines to Kubernetes and deploy with the control plane and compute plane separated:
Run the following commands to label the relevant nodes:
# Control-plane nodes
kubectl label nodes <control-node> cube.tencent.com/cube-control=true --overwrite
# Compute nodes (run sandboxes)
kubectl label nodes <compute-node> cube.tencent.com/cube-node=true --overwrite
# If the compute node is not a bare-metal server, you also need to apply this label, and PVM will be installed automatically
kubectl label nodes <pvm-compute-node> cube.tencent.com/allow-pvm-bootstrap=true --overwriteSet role taints (to keep unrelated workloads off these nodes):
# Optional
# kubectl taint nodes <control-node> cube.tencent.com/control=true:NoSchedule --overwrite
# Required for compute nodes
kubectl taint nodes <compute-node> cube.tencent.com/compute=true:NoSchedule --overwrite3.2 Single-node (not recommended)
Hint
For a single-node setup, we recommend deploying via Quick Start instead of Kubernetes.
Single-node labels and taints
If your Kubernetes cluster has only one machine, you can try the following labels and taints (not recommended):
export NODE=<your-node-name>
kubectl label nodes "$NODE" \
cube.tencent.com/cube-control=true \
cube.tencent.com/cube-node=true \
cube.tencent.com/allow-pvm-bootstrap=true \
--overwrite
kubectl taint nodes "$NODE" cube.tencent.com/control=true:NoSchedule --overwrite
kubectl taint nodes "$NODE" cube.tencent.com/compute=true:NoSchedule --overwrite4. Prepare the values file
Create the configuration file:
cp deploy/kubernetes/chart/runtime-values.example.yaml runtime-values.yamlEdit it to fit your needs. At minimum, fill in the following:
cubeProxy:
advertiseIP: "10.0.1.10"
domain: "cube.app"
tls:
mode: selfSigned # Trial; production: existingSecret / certManager
# Required
mysql:
host: "" # Empty = use built-in MySQL
password: "replace-me-mysql-password"
rootPassword: "replace-me-mysql-root-password"
redis:
host: ""
password: "replace-me-redis-password"More configuration
Control-plane PVC configuration
See 8.1 Control-plane PVC configuration.
Compute node data disk
Cube writes data to the host at /data/cubelet, which must be an XFS filesystem. For a smoother first-time experience, we create a 25 GB loopback image file and mount it by default.
For production deployments, or to adjust related configuration, see 8.2 Compute node data disk configuration.
Compute node networking
Decide before you deploy
The default host network keeps sandbox devices in the host netns across cube-node Pod recreation; drain remains the supported path until cubelet-restart survival is live-validated. On the Pod network (cubeNode.hostNetwork: false) the Pod must not be recreated while sandboxes run — all sandbox networking on that node breaks, with no self-healing. Details and trade-offs (including NetworkPolicy): 8.3 cube-node networking and Pod recreation.
The host-network default is new and not yet validated against every CNI / admission configuration. A hostNetwork Pod bypasses CNI IPAM and is not selectable by NetworkPolicy; on an eBPF CNI (e.g. Cilium) the CNI's own eBPF programs may conflict with the cubevs hooks on the host interface, and a cluster enforcing the restricted Pod Security Admission standard rejects hostNetwork Pods. Validate the default in a test cluster first before relying on it in production; if your cluster cannot allow hostNetwork, set cubeNode.hostNetwork: false (see 8.3).
5. Helm install
Users in Mainland China
Add -f deploy/kubernetes/chart/values-cn.yaml to whichever install command you use below to pull images from the China registry.
Run the following commands from the repository root to complete the install:
Because this process downloads many resources and runs initialization actions, it can take some time depending on your machine performance and network conditions.
Standard multi-node deployment:
# Standard multi-node
helm upgrade --install cube ./deploy/kubernetes/chart \
-n cube-system \
--create-namespace \
-f runtime-values.yaml \
--wait \
--timeout 90mTencent Cloud TKE:
helm upgrade --install cube ./deploy/kubernetes/chart \
-n cube-system \
--create-namespace \
-f deploy/kubernetes/chart/values-tke.yaml \
-f runtime-values.yaml \
--wait \
--timeout 90mSingle-node trial:
helm upgrade --install cube ./deploy/kubernetes/chart \
-n cube-system \
--create-namespace \
-f deploy/kubernetes/chart/values-single-node.yaml \
-f runtime-values.yaml \
--wait \
--timeout 90m6. Verify the deployment
# 1) Are Pods Ready?
kubectl get pods -n cube-system -o wide
# 2) Have compute nodes registered with CubeOps?
kubectl exec -n cube-system deploy/cube-cubemastercli -- \
sh -lc 'cubeopscli --address "$CUBEOPSCLI_ADDRESS" --port "$CUBEOPSCLI_PORT" node list'
# 3) Built-in end-to-end tests (a few minutes)
helm test cube -n cube-system --timeout 20m --logsExpect:
cube-nodeReady count = number of nodes labeledcube-node=true- All services in the
cube-systemnamespace are Ready helm testpasses
Access the Cube WebUI:
kubectl -n cube-system port-forward svc/cube-webui 12088:12088
# Open http://127.0.0.1:12088 in your browser7. Uninstall
helm uninstall cube -n cube-system
kubectl delete namespace cube-systemUninstall does not automatically clean up:
- Cube labels / role taints on nodes
- Compute-node hostPath data
- PVM host kernel changes
8. Advanced configuration
8.1 Control-plane PVC configuration
- Default: CubeMaster / MySQL / Redis / MinIO use the cluster default StorageClass
- To specify an SC: set
persistence.storageClassName: <name>inruntime-values.yaml - Single-node / no CSI: switch to hostPath (see comments in
runtime-values.example.yaml)
8.2 Compute node data disk configuration
For production deployments, we recommend provisioning a dedicated data disk for each compute node, formatting it as XFS, and mounting it under /data/cubelet. This gives the best performance and stability. Then configure runtime-values.yaml:
bootstrap:
nodeInit:
dataCubelet:
loopback:
enabled: falseIf your nodes have no dedicated data disk, by default we mount a 25 GB loop device at /data/cubelet during compute-node initialization.
Only affects the "first" creation
By default, the disk image file is created automatically only when /data/cubelet-xfs.img does not exist. Subsequent config changes do not auto-expand. To resize, manual steps in a maintenance window are required.
To customize the data disk configuration before install:
| values path | Default | Description |
|---|---|---|
bootstrap.nodeInit.dataCubelet.loopback.enabled | true | When true, bootstrap creates and mounts a loopback XFS on the node |
bootstrap.nodeInit.dataCubelet.loopback.imagePath | /data/cubelet-xfs.img | Path to the image file |
bootstrap.nodeInit.dataCubelet.loopback.size | 25G | Size when the image is first created (truncate -s) |
bootstrap:
nodeInit:
dataCubelet:
loopback:
enabled: true
size: 200G # Size to capacity plan; must be less than free space on the filesystem holding the image8.3 cube-node networking and Pod recreation
In one sentence
cube-node runs on the host network by default, so recreating the Big Pod no longer changes the network namespace sandbox networking lives in. Drain remains the supported compute-plane path until cubelet-restart survival is live-validated — see Upgrade. On the Pod network (cubeNode.hostNetwork: false) it must not be recreated, because recreation breaks network connectivity for all sandboxes on the node (inbound and outbound) with no self-healing.
Why the network mode decides this
Sandbox network devices (TAP devices) and the cubevs hooks live in the network namespace of the cube-node Pod, so the Pod's network mode decides what a recreation does to them:
| Item | Description |
|---|---|
| Trigger | Anything that recreates the Pod: DaemonSet template change, image bump, manual kubectl delete pod, … |
| Host network (default) | The netns is the host's and does not change across recreation, so devices stay. Survival through a cubelet restart is not live-validated — drain until it is (see Upgrade) |
Pod network (hostNetwork: false) | The Pod netns is destroyed: all sandboxes on the node lose network connectivity — both inbound and outbound — with no self-healing. The only recovery is to destroy and recreate the affected sandboxes |
Compute-plane upgrades are affected for the same reason; see Upgrade.
Default: host network
cubeNode.hostNetwork defaults to true; a fresh install needs nothing. What to check when you keep it:
| Item | Description |
|---|---|
| DNS | dnsPolicy becomes ClusterFirstWithHostNet automatically, so in-cluster DNS keeps resolving |
| Port conflicts | cubelet (9998 / 9999 / 9966) and CubeS3lvol (when enabled) bind on the host; cube-node-init fails fast if another process holds one (bootstrap.nodeInit.checkHostPorts). CubeEgress ports are not in this check — verify them yourself if it is enabled. Cubelet ports are unauthenticated — firewall them like kubelet's 10250 |
| NetworkPolicy | Cannot govern sandbox traffic; see below |
| Monitoring / firewalls | Re-point anything keyed on the Pod IP / Pod CIDR at the node IP instead |
| CNI / admission | A hostNetwork Pod bypasses CNI IPAM and is not selectable by NetworkPolicy; an eBPF CNI (e.g. Cilium) may conflict with the cubevs hooks on the host interface; restricted Pod Security Admission rejects hostNetwork Pods. This default is new — validate it in a test cluster first, or set cubeNode.hostNetwork: false |
Opting into the Pod network
Set cubeNode.hostNetwork: false when Kubernetes NetworkPolicy (e.g. "can sandboxes access this Service") must govern sandbox traffic. Everything else about the Pod network is a caveat:
- a
cube-nodePod recreation breaks networking for every sandbox on the node; - the guest MTU has to fit the CNI overhead —
cubeNode.network.mtu: autolowers it to the detected interface.
Changing the value on an existing release recreates every Big Pod, so a pre-upgrade Hook gates it: see Upgrade · changing the network mode.
Further work: PR #1189 (not merged; for reference only)
For NetworkPolicy over sandbox traffic to in-cluster Services / Pods while staying on the host network, and for a fully in-place cube-node replacement, see PR #1189:
- Only traffic destined for the cluster CIDRs is forwarded through a node-local EgressProxy Pod and SNAT'd to the Proxy Pod IP, so it is governed by your own NetworkPolicies; all other traffic keeps the normal route.
FAQ quick reference
| Symptom | What to do |
|---|---|
| Control-plane Pods stay Pending | By default nodes need cube-control=true (or your custom nodeSelector); also check taints / tolerations |
cube-node DESIRED=0 | Are compute nodes labeled cube-node=true? |
Helm reports CHANGE_ME_* / validation failure | Check that passwords and placement in runtime-values.yaml are complete |
| First install is slow / nodes reboot | Expected when PVM is enabled; wait for fingerprints to be ready and gates cleared, then check bootstrap / Big Pod |
More detail in the other docs in this directory: