DATACENTER AS AN APPLICATION™

User Guide

datacenters

Virtual Datacenters

On top right of the Mission Planner interface, select Datacenter. Create or select a virtual datacenter. You have can have many datacenters that are virtually isolated from each other. 

placeholder image

Data Layers

Left pane is a drawer that lets you select what you want to show on the Mission Planner canvas. The screenshot here shows no layers selected so its cleaner but without contextual details. 

placeholder image

System Arsenal

System Arsenal pull-out drawer on the right side of the datacenter canvas lists all virtual system templates. Switches, firewalls, routers, virtual machines. Deploy to the canvas by clicking one.

placeholder image

Selection Details

Click a node on the canvas and the right pane shows selection details. Runtime state, interfaces, template used, worker execution host, capacity, GPU details.

placeholder image

Right Click Menu

Right click a node on the canvas gives a radial menu for editing and controlling the node. Open the node’s console, edit system, create network or serial link, suspend/resume, hibernate, snapshot management, and move/migrate to another STRATUM worker.

placeholder image

Snapshots

From the right-click radial menu you can manage snapshots. You can include snapshotting memory runtime as well, not just disk. Due to the nature of massive scaling capabilities (10s of thousands of stratum workers), the current version of STRATUM will not let you migrate/move a VM to another host while snapshots are enabled. To migrate/move you will need to press the button ‘Consolidate to Shared Storage’. 

placeholder image

Migrate VM

From the right-click radial menu you can move/migrate VMs to specific hosts when there is a specific need to do so. STRATUM shows you what target hosts are compatible with resources, GPUs, ARM vs x86 architecture, etc… 

placeholder image

Node Editor (runtime)

From the right-click radial menu you can edit a node. Name, template used, payload image used, ethernet interfaces, and runtime details enabling or disabiling: acceleration, 3D acceleration (OpenGL shared with host), encrypted runtime (memory of VM), and high availability (HA). HA is active-passive, meaning if a VM crashes or its host goes offline, the VM will be rescheduled to run on another worker. Memory state is not preserved with HA in the current STRATUM version. 

placeholder image

Node Editor (GPU)

Scrolling down in the node editor gives more settings, removable media and STRATUM’s AI/GPU Accelerator Fabric. You can attach local-to-host PCIe passthrough, but STRATUM also can attach remote accelerators. Any GPU in your datacenter can be shared and are accessible as paravirtual GPU PCIe devices in your VM. You do not need to install any GPU-over-IP software and maintain GPU-over-IP software, just enable the STRATUM GPU device and it enables a paravirtual GPU as a PCIe device. 
The current version of STRATUM’s special GPU PCI device only supports NVIDIA CUDA runtime. 

placeholder image

VM Console

From the right-click radial menu you can open the console. Depending on your VM setup, this could be the display console or a serial port. Serial ports are often used with switches, routers, and firewalls. 

placeholder image

Removable Media (runtime)

From the console UI, on the top right is a button with a CDrom icon. That opens up the removable media interface to add or swap ISO images, or add STRATUM’s guest tools image. Guest tools are needed for STRATUM’s paravirtual PCIe GPU device enabling access to remote GPUs as if local. Windows-based VMs should also install STRATUM Guest Tools to enable VirtIO device drivers for performance. Your modern Linux platforms already come with VirtIO device drivers. 

placeholder image

All Nodes UI

The All Nodes view provides a centralized table for managing virtual machines and network devices at scale. View CPU, memory, ethernet, GPU, image, and runtime details, edit configurable settings, and manage each node individually or perform lifecycle actions across all selected nodes at once.

placeholder image

STRATUM Switch

While you can use any virtual switch appliance/VM, from Cisco, Arista, Palo Alto, etc.. STRATUM does provide an integrated, fast, and powerful software L2 and L3 switch that uses a familiar interface. Setup VLANs, setup port mirroring, everything you would do in a physical switch.  
On the right selection panel, you can get more switch and individual port details.

placeholder image

Link Impairment Profile

Selecting a network link on the canvas shows its details in the selection pane on the right. This is useful for testing labs to similiate latency, jitter, and packet loss. Some of our customers use STRATUM as a digital twin, complete with network realism of satelitte communications and large geographical distances. 

placeholder image

Edit Link Impairment

Editing a link Impairment profile lets the user enter custom latency, jitter, or packet loss settings or choose from a set of prepared mission effects profiles: clean, high-latency backhaul, degraded RF, or contested link. 

placeholder image

Interfaces

Selected Data Layers show in-line details such as Interfaces. A fast operational way to see the digital landscape.

placeholder image

Execution

Selelcted Data Layers show in-line details such as execution/worker host context. With 100s or thousands of STRATUM worker hosts, you can clearly see impact and mission risk quickly.

placeholder image

Storage

Selected Data Layers show in-line details operationally such as storage. STRATUM uses a storage protocol and data is stored on the Controller. STRATUM is also unique in that template VM data is immutable, the individual VMs store deltas but still using the same storage protocl, However snapshots, hibernation, and paused VMs (memory runtime) store that specific metadata locally on the worker node. 

placeholder image

Arsenal Forge: Templates

Everything in STRATUM is a template. You deploy a template to the canvas and run it. Think of templates like how Linux container images are to applications. You can deploy 100 application using the same image. Arsenal Forge lets you create, manage, import, and export templates. 

placeholder image

Arsenal Forge: Image Create

After you create a template, you need to create a VM image folder using that same template name and a version# (template name = test, VM folder = test-1.0.0). Then click the create VM image button to add a VM disk image into that folder. 

placeholder image

Arsenal Forge: VM Images

Your VM images should be in folders that are the same name as its related template, with a version number after it. ISO images can be added in Arsenal Forge too. 

placeholder image

Arsenal Forge: ISOs

The STRATUM Arsenal Forge uploads, versions, and manages ISO images used across your virtual datacenters. Administrators can organize, inspect, download, rename, and track image history from one central library.

placeholder image

STRATUM Ops

Operations interface is a datacenter visualization in real-time and is accessible from a button on the left rail. Operational layers include all assets, physical, worker scheduler, workloads, networks, GPU, Storage, and assets requriing operational attention. If you have 100s of nodes, they will show up as their own drill-down node that can be expanded or collapsed. Relationships are key here, and identifying the blast radius to a failure, lineage, or resource contention. 

Administrator Guide

placeholder image

Compute Fabric

Cluster-wide (all datacenters, clouds, labs, edge) partitions, jobs, nodes, and compute fabric status. Every VM, network switch, firewall,or router are submitted as jobs, similiar to how High Performance Computing (HPC) delivers massive execution scale. 

placeholder image

AI Accelerator Fabric

Cluster-wide remote GPU capacity, placement, and VM attachment readiness. Users select AI/GPU capability only and STRATUM owns GPU placement, VM transport, and guest CUDA/NVML exposure.

placeholder image

Network Fabric

The STRATUM Continuum is encrypted mesh management paths and distributed private datacenter networks. All STRATUM worker nodes connect through this network fabric even if they are remote, Cloud, or edge. It provides a virtual datacenter network substrate across everything to form 1 platform. Workers initiate outbound HTTPS enrollment and connect into the distributed Continuum mesh autonomously. Their private tunnel keys never leave the worker. For manual enrollment, you download Token and Controller CA from this interface. 

placeholder image

License Management

License add, replace, download, and details

placeholder image

Users

User administration. Create, delete, edit, select role, select groups. Users are able to use CAC or YukiKey but those require a valid TLS certificate for STRATUM. The user sets those up from the login window. 

placeholder image

Roles

Manage User roles such as platform administration, enabling users to edit mission topology where you can have users view-only and operate VMs and the network but cannot make changes.

placeholder image

Groups

Manage User groups.

placeholder image

Sessions

List users logged in and if active. 

placeholder image

Audit

Administative action audit showing time, actor, action, and target. 

placeholder image

Execution Hosts

STRATUM host execution enrollment is automated, but you are able to add worker nodes manually, edit, and delete them. Not usually necessary but it is here in case manual worker host configuration is needed. 

placeholder image

Storage Fabric

The STRATUM Storage Fabric distributes controller-authoritative VM image replicas across dedicated STRATUM Storage nodes to reduce centralized storage load and place VM data closer to where it runs. Administrators assign templates to specific storage nodes, review the desired placement, and ‘Bake’ the Storage to replicate those images to their selected destinations. Storage nodes can operate across edge locations, cloud environments, multiple datacenters, or even disconnected platforms that have selective needs, while remaining connected through direct Continuum iSCSI paths. The controller retains the authoritative copy, giving STRATUM a consistent source of truth while allowing new VM disks to be served from the most appropriate storage location.

DevSecOps Guide

placeholder image

Workload Layering

Your infrastructure, workloads, and deployment logic are layered into one portable Datacenter-as-an-Application. Package virtual machines directly into a custom STRATUM OCI image by copying their disks into the image and creating or reusing a template definition under /stratum/html/templates/intel/ within the container image.

STRATUM automatically maps each disk according to its naming convention: sata* disks are distributed across AHCI controllers, sd* disks use VirtIO SCSI, vd* and virtio* disks use VirtIO Block, and hd* disks use IDE. Raw disk formats such as .raw, .img, and .bin are presented through the iSCSI URI provided by STRATUM Store rather than accessed through a local filesystem path.

The resulting OCI image contains STRATUM, the VM workloads, and their deployment definitions as one versioned artifact. Distribute it through an existing container registry, deploy it to any supported Linux or Kubernetes environment, import a Virtual Datacenter, and start the complete environment with its required virtual machines already embedded.

docker load -i stratum-container.tar
cat > Dockerfile.workload <<‘EOF’
FROM stratum:1.0.0
COPY templates/esxi.yaml /stratum/html/templates/intel/customvm/
COPY disks/sataa.qcow2 /stratum/addons/qemu/customvm/
COPY disks/satab.qcow2 /stratum/addons/qemu/customvm/
EOF
docker build -f Dockerfile.workload -t stratum:1.0.0-customvm .
docker save -o stratum-container-customvm.tar stratum:1.0.0-customvm
Distribute stratum-container-customvm.tar

placeholder image

Move from VMware to STRATUM

Bring existing VMware workloads forward without rebuilding them from scratch. STRATUM provides free, open-source migration tools on GitHub that convert VMware-exported OVA and OVF packages into portable STRATUM Arsenal Bundles, ready to import, version, and deploy.

For enterprise migrations, STRATUM can also use an optional virt-v2v backend to inspect and prepare Windows or Linux guests for KVM, including installing the VirtIO drivers required for STRATUM. It supports VMware-exported OVA files and unpacked OVA directories, providing a practical path from legacy virtualization to infrastructure delivered as a versioned application.

See our GitHub page for migration tools, documentation, and examples.

00 / SYSTEM REQUIREMENTS

System Requirements

STRATUM / SYSTEM REQUIREMENTS

Size the host for STRATUM—and the workloads it will run.

STRATUM requires a 64-bit Linux host with a working container runtime and hardware-assisted virtualization. The platform minimums below cover STRATUM itself; your host must also provide enough CPU, memory, networking, GPU, and storage capacity for all simultaneously running workloads.

01 HostMinimum Requirements Minimum host requirements The operating system, runtime, virtualization, memory, storage, and permissions required to run STRATUM.
RequirementMinimum or guidance
Operating system64-bit Linux distribution on x86_64 or Arm64. S390x Linux, Microsoft Windows, and MacOS will be released soon if we have demand for them.
Container runtimeDocker Engine or Podman, installed, configured, and running.
Processor64-bit CPU with hardware virtualization enabled.
Virtualization accessLinux KVM available to STRATUM through /dev/kvm.
MemoryAt least 1 GB of available RAM for the STRATUM platform.
Platform storageAt least 10 GB of available storage for STRATUM, excluding workload capacity.
PermissionsRoot or equivalent administrative access for installation and host configuration.
Storage mediaSSD or NVMe recommended. RAID-backed HDD storage is supported when performance is sufficient.
Workload capacity is additional.

The host must also provide sufficient CPU, RAM, network bandwidth, GPU resources, and storage performance for every workload that may run concurrently.

02 ComputeKVM Processor and virtualization Hardware-assisted virtualization must be enabled and exposed to Linux through KVM.

The host processor must support 64-bit operation and hardware-assisted virtualization. Virtualization must be enabled in the system firmware and available to Linux through /dev/kvm.

When STRATUM itself runs inside a virtual machine or cloud instance, the outer platform must expose nested virtualization to that guest.

KVM access is a host requirement.

A container runtime alone is not sufficient. STRATUM must be able to access the Linux KVM device in order to run hardware-accelerated virtual machines.

03 CapacityMemory Plan host memory Add Linux, STRATUM, active workloads, and operational headroom—not just configured guest RAM.

The 1 GB minimum is for the STRATUM platform only. It does not include memory assigned to running virtual machines, virtual switches, routers, firewalls, or other workload components.

Linux host+STRATUM+simultaneous workloads+operational headroom

For example, four virtual machines configured with 4 GB of RAM each require capacity for 16 GB of guest memory, in addition to the memory required by Linux, STRATUM, and normal operating headroom.

Plan for concurrency.

Size memory for the workloads that can run at the same time, not merely the number of templates or systems stored in the environment.

04 CapacityStorage Plan platform and workload storage The 10 GB platform minimum does not include templates, VM disks, snapshots, media, data, or backups.

The 10 GB minimum covers the STRATUM platform only. Budget additional capacity for:

VM templates and imported disk images
Writable virtual-machine disks
Container images and writable layers
Snapshots and checkpoints
ISO images and installation media
Persistent application data
Logs and temporary files
Exports, updates, and backups

A virtual machine created from a STRATUM template does not initially require another complete copy of the template. The template data is shared, while each VM stores its own changes and deltas.

Leave recovery and growth headroom.

As a practical planning baseline, provide roughly twice the storage you expect active VMs and platform data to consume. This leaves room for deltas, snapshots, updates, temporary operations, and unexpected growth.

05 Container RuntimeStorage Path Check the actual container storage filesystem Free space on / does not help when Docker or Podman stores its data on a smaller /var filesystem.

Docker and Podman store downloaded images, container layers, metadata, and logs in the container runtime's storage directory. Common default locations are:

Docker/var/lib/docker
Podman/var/lib/containers/storage

On many Linux systems, /var is a separate and relatively small filesystem. In that layout, the root filesystem may have ample free space while the container storage filesystem becomes full.

Check Docker storage
df -h /var/lib/docker
Check Podman storage
df -h /var/lib/containers/storage
Measure the filesystem that actually holds runtime data.

It must have at least 10 GB available for STRATUM, plus sufficient room for image updates, logs, writable layers, and normal runtime growth.

06 Persistent DataVM Storage Provide capacity for STRATUM persistent storage VM disks and persistent platform data normally live below /persistent, but the location is configurable.

STRATUM's default host location for VM and persistent platform storage is below /persistent. This path can be changed through the stratum-host configuration.

Ensure the selected filesystem has enough capacity for VM disks, templates, snapshots, persistent databases, exports, and expected growth. For larger deployments, place persistent data on a dedicated high-capacity SSD, NVMe device, storage array, or RAID-backed volume.

Separate runtime and workload planning.

The container runtime storage path and STRATUM persistent storage path may reside on different filesystems. Verify and size both.

07 StoragePerformance Match storage performance to the workload Capacity alone is not enough when several active virtual machines generate random I/O.
SSD or NVMe

Strongly recommended for hosts running multiple active virtual machines, frequent snapshots, or I/O-intensive applications.

RAID-backed HDD

Supported when the array provides adequate random I/O performance, resilience, and capacity for the intended workload.

Single HDD

Generally appropriate only for light evaluation, archival use, or workloads with limited disk activity.

Evaluate latency and random I/O.

Large sequential capacity does not guarantee acceptable VM performance. Size storage for the number and behavior of active workloads.

08 KubernetesPersistent Volumes Kubernetes storage requirements Kubernetes deployments require durable PersistentVolume capacity sized for both STRATUM and its workloads.

Kubernetes deployments require ample persistent volume capacity for STRATUM platform data, templates, virtual disks, snapshots, and workload growth.

The detailed PersistentVolume, StorageClass, access-mode, and deployment-specific requirements are included with the STRATUM Helm chart download package.

The same sizing rules still apply.

A PVC changes how storage is presented to STRATUM; it does not remove the need to plan capacity, performance, redundancy, snapshots, and backups.

01 / HOST SETUP

Setup host

Setting up a few things on your Linux host

run ./stratum-host command.
It opens up a console-based menu interface. Scroll to Host Action Status. Tab over to right side. ‘P’ for Plan to see what commands would be used to setup. Also select which GPU and PCI devices to have available to STRATUM (default will automatically load all GPU but in some cases you may want to have more control).
Go to Dashboard, press A to apply host setup.
CTRL-S to save, Q to quit

Command-line capability:
run sudo ./stratum-host apply
Compare the saved desired state with the actual host:
run sudo ./stratum-host status

For automation, the same configuration (/etc/stratum/host.yaml) can be copied to many hosts and used without the menu interface:
run sudo ./stratum-host apply —yes

NOTE: You do not need to use Network Continuum, which is an encrypted virtual network layer, if you are running with Kubernetes that already has a virtual network or you are on a protected LAN. Unless you are using a 3rd party Kubernetes service, then you should keep Continuum for data protection.

The UI is the interactive editor and operator interface; /etc/stratum/host.yaml is the persistent source of truth.
STRATUM Host Manager r005-r001 | /etc/stratum/host.yaml
────────────────────────────┬──────────────────────────────────────────────────────────
Dashboard │ Host Action Status
Runtime │
Certificates │ [x] Container engine ok docker is usable
Network │ [x] Create STRATUM directories warning missing /etc/s…
Slurm & GPU │ [x] Persistent data permissions ok persistent da…
Host Setup │ [x] Large-host sysctl warning not applied
Controller Services │ [x] Network sysctl warning not applied
Advanced │ [x] irqbalance ok enabled
• Host Action Status │ [x] fstrim timer ok enabled
PCI & GPU Inventory │ [x] NVMe tuning ok no NVMe devices
Deployment Roles │ ▶ [x] pnet bridge ok pnet0 healthy with…
Help │ [x] pnet NAT ok persistent masquer…
Quit │
│ Selected: pnet bridge
│ Status: ok - pnet0 healthy with uplink ens33
│ IPv4 192.168.1.110/24
│ ↑/↓ select | Enter inspect | Space enable/disable | P
│ plan | A apply
↑↓ select | Enter inspect | Space enable | P plan | A apply | Tab menu | Ctrl-S save
STRATUM Host Manager r005-r001 | /etc/stratum/host.yaml
────────────────────────────┬──────────────────────────────────────────────────────────
Dashboard │ PCI & GPU Inventory
Runtime │
Certificates │ 0000:81:00.0 10de:2331 driver=nvidia group=65
Network │
Slurm & GPU │ Selected: 0000:81:00.0
Host Setup │ Class=030200 Driver=nvidia IOMMU=65 BootVGA=false
Controller Services │ 0000:81:00.0 3D controller [0302]: NVIDIA Corporation
Advanced │ GH100 [H100 PCIe] [10de:2331] (rev a1)
Host Action Status │
▶ PCI & GPU Inventory │ Vendor: NVIDIA Model: H100 PCIe VRAM: 80 GB HBM3
Deployment Roles │ Driver: 535.183.01 CUDA: 12.2 MIG: supported
Help │ PCIe: Gen5 x16 NUMA: 0 Serial: 0324519087123
Quit │
│ ↑/↓ select
Tab switch pane | ↑↓/j/k navigate | Enter edit/select | Ctrl-S save | Q quit

02 / DEPLOY

Deploy container

Everything STRATUM is deployed as a single container

Simply deploy the container using a container manager such as docker or podman. You can also deploy from a container register server if you have one. This is how STRATUM can deploy and scale to thousands of nodes quickly, using the same mechanism as cloud-scale applications use.
Run sudo ./stratum-host and configure the STRATUM controller certificates (not needed on worker nodes). The default will create self-signed certificates, however you should create valid certificates for production use.

Container-based application deliveries have changed the world already. Now STRATUM lets it change infrastructure too. 

$ sudo docker load -i stratum.tar

$ sudo docker images
IMAGE            ID             DISK USAGE   CONTENT SIZE 
stratum:latest   38b731499409         5GB             0B   

$ sudo ./stratum-host
STRATUM Host Manager r005-r001 | /etc/stratum/host.yaml
────────────────────────────┬──────────────────────────────────────────────────────────
  Dashboard                 │ Certificates
  Runtime                   │
• Certificates            ▶ Certificate mode             self-signed
  Network                   │ Self-signed is the demo-friendly default. External
  Slurm & GPU               │ never overwrites certificate files.
  Host Setup                │   Certificate directory        /etc/stratum/certs
  Controller Services       │   Demo PKCS#12 password        1234
  Advanced                  │   Expiration warning days      30
  Host Action Status        │
  PCI & GPU Inventory       │ Enter edits/cycles the selected value. Left/right cycles
  Deployment Roles          │ booleans and choices.
  Help                      │
  Quit                      │
                            │
                            │
                            │
                            │
                            │
                            │
                            │
Tab switch pane | ↑↓/j/k navigate | Enter edit/select | Ctrl-S save | Q quit

03 / START

Start container

Simple

Run sudo ./stratum-host
Tab over in Dashboard. Press S to start. X to stop.
Workers call home to the Controller and auto-join the STRATUM Continuum mesh which is a secure network overlay on top of your existing network that spans datacenters, clouds, and edge.

Command-line capability:
run sudo ./stratum-host start

STRATUM Host Manager r005-r001 | /etc/stratum/host.yaml
────────────────────────────┬──────────────────────────────────────────────────────────
▶ Dashboard                 │ Dashboard
  Runtime                   │
  Certificates              │ Role:        controller
  Network                   │ Node:        ctrl1.stratum.lab
  Slurm & GPU               │ Controller:  ctrl1.stratum.lab
  Host Setup                │ Public UI:   https://ctrl1.stratum.lab:8443/stratum/
  Controller Services       │ Engine:      docker
  Advanced                  │ Container:   running running=true
  Host Action Status        │ Image:       stratum:latest
  PCI & GPU Inventory       │ Network:     ok - pnet0 healthy with uplink ens33
  Deployment Roles          │ Firewall:    warning - UFW missing required rules
  Help                      │ IOMMU:       0 groups
  Quit                      │ GPUs:        0 matching
                            │ Certificate: ok - certificate valid for 728 more days
                            │
                            │ Actions:
                            │   S start   X stop   C reconcile
                            │   P plan    A apply host setup
                            │   V verify  R refresh
                            │
Tab switch pane | ↑↓/j/k navigate | Enter edit/select | Ctrl-S save | Q quit

CLI alternative: sudo ./stratum-host start

ALT: Kubernetes

Alternative: K8s deploy

STRATUM is a container, deploy it to Kubernetes

Many datacenters are buying and deploying full racks, not indivudual servers these days. They tend to come with Kubernetes to scale out to thousands of nodes. Deploy STRATUM to Kubernetes and scale out your virtualized datacenter massively. 

The chart design is simple: the controller uses a one-replica StatefulSet with a Persistent Volume (PVC) and image-based init container; workers use a node-labelled DaemonSet. Kubernetes values generate both the mounted host.yaml and the runtime environment, so there is still one configuration source.

STRATUM / KUBERNETES DEPLOYMENT

Prepare the cluster, then deploy with Helm.

Ensure that a Kubernetes environment is already installed and operational. Confirm that you can deploy applications to the cluster using Helm.

01 Container ImageRegistry Import the STRATUM image Load the STRATUM OCI container image into a registry accessible to the cluster.

Import the STRATUM container image into the private or enterprise registry used by your Kubernetes environment.

The repository and tag you publish here will be referenced later in values.yaml or supplied directly to Helm.

Registry access

Ensure every Kubernetes node that may run a STRATUM pod can resolve, authenticate to, and pull from the selected registry.

02 Helm Configurationvalues.yaml Configure the deployment Set the image location, persistent storage, certificates, and optional worker deployment.

Edit values.yaml for your environment, including the image repository location, PVC names, and certificate settings.

Enable STRATUM workers
workers:
  enabled: true

Workers are disabled by default. When workers are enabled, label the Kubernetes nodes intended to provide STRATUM worker or storage roles.

Example node labels
kubectl label node worker001 stratum.io/role=worker
kubectl label node worker002 stratum.io/role=worker
kubectl label node storage01 stratum.io/role=storage
03 KubernetesNamespace Create the STRATUM namespace Create the deployment namespace and apply the required privileged Pod Security labels.
Create namespace
kubectl create namespace stratum

STRATUM requires privileged pod access for host-integrated virtualization, networking, storage, and hardware operations.

Apply Pod Security labels
kubectl label namespace stratum \
  pod-security.kubernetes.io/enforce=privileged \
  pod-security.kubernetes.io/audit=privileged \
  pod-security.kubernetes.io/warn=privileged
04 Node PlacementController Assign the controller node Label the Kubernetes node that will host the STRATUM controller role.

Label the intended controller node before installing the Helm release.

Controller label
kubectl label node ctrl1 stratum.io/role=controller

Replace ctrl1 with the actual Kubernetes node name selected for the STRATUM controller.

05 HelmInstall or Upgrade Deploy STRATUM Install the chart—or upgrade an existing release—with your registry and public-host settings.

Install the STRATUM Helm chart into the prepared namespace. The same command upgrades the release when it already exists.

Helm deployment
helm upgrade --install stratum \
  ./stratum-1.0.0.tgz \
  --namespace stratum \
  --set image.repository=registry.stratum.local/stratum \
  --set image.tag=latest \
  --set stratum.publicHost=ctrl1.stratum.lab

Replace the example repository, image tag, chart package, and public hostname with the values for your environment.

ALT: Appliances

Alternative: STRATUM Cloud

STRATUM appliances that boot STRATUM container

STRATUM appliances are customized Linux kernels that boot, have nothing else except enough to boot a Linux container. And that container is STRATUM. So you can deploy a VMware OVF image, an Amazon AWS AMI, etc… and boot STRATUM that way. Its still a container, just wrapped in a Linux kernel. 

STRATUM appliances will auto-enlarge it’s filesystem if the host/hypervisor expands it. STRATUM appliances also make use of pre-labeled disks, identifying specific labels as VM storage. 

01 VirtualizationOVA Appliance VMware Deploy the STRATUM OVA appliance to VMware vSphere or ESXi.

Deploy the STRATUM OVA appliance through the VMware Deploy OVF Template workflow.

  1. Select the downloaded stratum.ova file.
  2. Choose the destination host, datastore, and network.
  3. Review the virtual hardware settings and complete deployment.
  4. Increase virtual disk size as needed
  5. Power on the appliance and continue setup from the STRATUM console.
02 Amazon Web ServicesAMI Amazon EC2 Deploy the STRATUM AMI with nested virtualization enabled.

Launch the STRATUM AMI on an EC2 instance type that supports nested virtualization.

AWS CLI
aws ec2 run-instances \
  --region us-east-1 \
  --image-id "$STRATUM_AMI_ID" \
  --instance-type c8i-flex.large \
  --cpu-options "NestedVirtualization=enabled" \
  --key-name "$KEY_NAME" \
  --security-group-ids "$SECURITY_GROUP_ID" \
  --subnet-id "$SUBNET_ID"
GPU deployment options

GPU families such as G4, G5, G6, P4, and P5 provide GPUs to the outer EC2 instance, but their standard virtual instance sizes are not currently listed by AWS as nested-KVM instance types.

  1. Use a GPU bare-metal instance, such as g4dn.metal, when STRATUM must run local nested VMs and use the attached GPU on the same host.
  2. Use a lower-cost virtual GPU instance as a STRATUM GPU Fabric node. The host contributes its GPUs to the STRATUM virtual datacenter (GPU-over-IP), while KVM-capable STRATUM nodes run the VMs that consume those GPU resources.

STRATUM can also combine on-premises and cloud resources, allowing cloud GPUs to serve on-site workloads—or on-premises GPUs to serve cloud-hosted workloads—through the same virtual datacenter.

03 Google CloudGCE Image Google Compute Engine Deploy the STRATUM Google Cloud image with nested virtualization enabled.

Create a Google Compute Engine instance from the STRATUM Google Cloud image.

Google Cloud CLI
gcloud compute instances create STRATUM \
  --enable-nested-virtualization \
  --zone=ZONE \
  --min-cpu-platform="Intel Haswell"

Replace ZONE with the deployment zone and add the STRATUM image, machine type, network, disk, and access options required by your environment.

04 Microsoft AzureVHD Image Microsoft Azure Deploy the STRATUM Azure VHD as a managed image or virtual machine.

Upload the STRATUM Azure VHD, create an Azure managed disk or image from it, and deploy a new virtual machine.

  1. Upload the VHD to an Azure Storage account.
  2. Create a managed disk or reusable image from the uploaded VHD.
  3. Select an Azure VM size that supports nested virtualization.
  4. Configure networking, storage, access, and security settings, then start the VM.
05 Oracle Cloud InfrastructureCloud Image Oracle Cloud Infrastructure Deploy the STRATUM Oracle Cloud image to an OCI compute instance.

Import the STRATUM Oracle Cloud image as a custom image, then use it to launch an Oracle Cloud Infrastructure compute instance.

  1. Upload the STRATUM image to an OCI Object Storage bucket.
  2. Import it as a custom compute image.
  3. Select the target shape, virtual cloud network, subnet, storage, and access settings.
  4. Launch the instance and continue setup from the STRATUM console.

Stratum-host command (TUI & CLI)

STRATUM / HOST MANAGER

Configure the host. Deploy the datacenter.

stratum-host is the CLI and TUI for preparing Linux, configuring STRATUM, operating the OCI container, enrolling workers, managing pnet0 and nat0, assigning GPUs, validating storage, and automating fleets.

18 sections shown
01 GuideGetting StartedStart here What is stratum-host? The TUI for people; the CLI for repeatable operations.

stratum-host is the host-side control surface for STRATUM. It prepares Linux, manages the STRATUM OCI container, validates configuration, handles certificates and networking, exposes hardware inventory, and keeps the running container reconciled with /etc/stratum/host.yaml.

One command, two layers. The TUI edits desired configuration. The CLI makes that configuration repeatable through scripts, golden images, configuration management, and fleet provisioning.

Open the TUI

sudo stratum-host
# or
sudo stratum-host tui
KeyAction
TabSwitch between the left menu and the current pane.
↑ ↓ or j kMove through menus and fields.
EnterEdit, select, inspect, or run the highlighted operation.
Ctrl-SSave settings to /etc/stratum/host.yaml.
QQuit the TUI.

Dashboard shortcuts

KeyActionWhat it does
SStartStarts the configured controller, worker, or storage container.
XStopStops the container but retains it and its persistent data.
CReconcileRecreates a drifted container while preserving persistent state.
PPlanShows host changes without applying them.
AApplyApplies enabled host setup actions. Root is required.
VVerifyRuns checks for enabled features without changing the host.
RRefreshReloads live status.
stratum-host Dashboard menu
Runtime configuration in stratum-host.
02 RunbookGetting StartedWorker Install a STRATUM worker Plan first, apply host changes deliberately, install the trusted bundle, then reconcile.

A worker bootstrap bundle normally contains these five files:

host.yaml
stratum-ca.crt
stratum-container.tar
stratum-enroll
stratum-host

Copy both the stratum-host binary and stratum-container.tar to every host. The binary prepares and operates the host; the tarball supplies the STRATUM OCI image that Host Setup loads.

1. Inspect and apply host preparation

chmod +x ./stratum-host
sudo ./stratum-host plan
sudo ./stratum-host apply --yes
Yes, apply --yes may interrupt the network. Default physical-bridge mode moves the active uplink under pnet0. Use an out-of-band console, or select cloud-safe routed mode before applying when preserving SSH is critical.

2. Install enrollment material and configuration

sudo install -m 0755 stratum-host /usr/local/sbin/stratum-host
sudo install -d -m 0700 /etc/stratum/secure
sudo install -m 0600 stratum-enroll /etc/stratum/secure/stratum-enroll
sudo install -m 0644 stratum-ca.crt /etc/stratum/secure/controller-ca.crt
sudo install -m 0600 host.yaml /etc/stratum/host.yaml

3. Start and reconcile the worker

sudo stratum-host start --reconcile
sudo stratum-host status --deep
sudo stratum-host logs --follow

The worker must have a unique FQDN, and it must be able to resolve and reach the controller FQDN. The controller CA and enrollment secret are required when the Continuum fabric is enabled.

Repeatable by design. Run plan during change review, then run apply --yes from Ansible, Puppet, Chef, Terraform-driven provisioning, a golden-image pipeline, or your own scripting framework.
03 ScaleAutomationFleet pattern From one worker to 1,000 Standardize the host, distribute a small trusted bundle, and let workers identify themselves.

For a large fleet, treat stratum-host as an idempotent node-preparation step rather than a manual installer.

Fleet requirementGuidance
Consistent host preparationApply the same enabled Host Setup actions through golden images or stratum-host apply --yes.
Bootstrap filesDistribute stratum-host, stratum-container.tar, host.yaml, stratum-enroll, and stratum-ca.crt.
EntitlementReuse the approved entitlement or enrollment code generated in the STRATUM Admin UI where your license permits it.
IdentityEvery physical worker needs a unique, stable FQDN. identity.nodeName: auto uses the detected host FQDN.
Controller discoveryEvery worker must resolve the controller FQDN through DNS or /etc/hosts.
Portable GPU configurationKeep fleet YAML on gpu.policy: auto. Avoid host-specific PCI BDFs in a common image.
Controlled startupThe default Continuum startup jitter spreads enrollment and heartbeat traffic instead of creating a synchronized burst.

Sites divide one Continuum into regions

identity.site is a placement label. It can be default, us-east, boston, taiwan, tokyo, or any naming convention meaningful to your organization.

sudo stratum-host config set identity.site boston
sudo stratum-host start --reconcile

A virtual datacenter can set a default site. VMs deployed into that VDC then prefer workers in that site unless a more specific placement rule overrides it. This creates regions inside one enterprise-wide STRATUM Continuum without operating a separate control plane for every facility.

Site is scheduling metadata, not automatic data sovereignty. Storage, backups, logs, keys, identity systems, and external services must also follow the required regional policy.
04 TUI MenuRuntimeCore settings Runtime Role, identity, controller discovery, container lifecycle, and persistent state.

The Runtime menu defines what this host is, how it handles VM storage I/O, which container it runs, where it finds the controller, and where mutable state lives.

SettingPlain-English meaning
rolecontroller, worker, or dedicated storage. Fabric-gateway and failover-controller are visible but reserved.
execution.slurmSchedulingEnabled sends VMs and STRATUM Switch jobs through managed workers. Disabled uses the direct/local legacy path.
vm.ioProcessingControls IOThread assignment for virtio-blk and virtio-scsi devices: compatibility, auto, or dedicated. The selected mode applies to newly started VMs.
identity.nodeNameStable Continuum identity. auto uses controller host, worker FQDN, or a storage-specific name.
identity.sitePlacement region or facility label used by scheduling.
controller.hostController FQDN or IP. Workers should normally use a resolvable FQDN.
controller.ipController underlay IP. auto resolves the controller host or detects the controller IPv4 address.
public.hostBrowser-facing hostname used by the UI, authentication layer, and certificate SANs.
public.httpsPortHTTPS port. Current host-networked runtime uses 8443.
runtime.engineauto prefers a responsive Docker daemon, then Podman.
runtime.imageOCI image to run. Use a versioned release tag in production rather than an unpinned moving tag.
runtime.containerNameDocker or Podman container name. Use a separate name for a second storage container on the same host.
runtime.restartPolicyContainer restart behavior: unless-stopped, always, on-failure, or none.
runtime.socketContainer API socket: auto, none, or an explicit Docker-compatible path.
storage.rootHost path for mutable STRATUM state. Default: /persistent/stratum.
storage.persistControllerDataKeeps datacenters, data, Arsenal content, and databases outside the replaceable image.
storage.seedControllerDataCopies image defaults into missing persistent paths on first deployment.
storage.repairPermissionsAligns persistent path ownership with the numeric stratum UID/GID inside the selected image.
New deployment rule: leave automatic identity and engine detection enabled unless the environment has a concrete reason to override them. Explicit settings are most useful for fixed controller addresses, multiple containers on one host, or strict release pinning.
stratum-host Runtime menu
Runtime configuration in stratum-host.
05 TUI SettingRuntimeVM storage I/O I/O Processing Choose how STRATUM assigns IOThreads for virtio storage so you can balance compatibility, latency, and host thread count.

I/O Processing controls how STRATUM assigns QEMU IOThreads to VM storage devices. An IOThread is a dedicated event-loop thread that one or more devices can use for disk I/O instead of sending all storage work through the main VM loop. Moving storage activity off that primary loop can reduce lock contention, improve I/O latency, and cut down guest-visible jitter on busy systems.

STRATUM applies this feature to virtio-blk and virtio-scsi storage, which are the paravirtualized disk paths most likely to benefit. The chosen mode affects newly started VMs; already running VMs keep the threading model they were launched with.

Recommended default: keep Compatibility if you want strict continuity with historic STRATUM behavior. Use Auto when you want a practical middle ground. Reserve Dedicated for testing, benchmarking, or specialized high-I/O workloads that justify the extra host threads.

Implemented modes

ModeVM behavior
CompatibilityPreserves STRATUM's established behavior and does not introduce the newer IOThread assignment model. This is the safest choice when consistency matters more than I/O experimentation.
AutoLazily creates one shared IOThread per VM when that VM has virtio-blk or virtio-scsi storage. It improves separation from the main VM loop without multiplying threads aggressively.
DedicatedCreates one IOThread per virtio-blk disk and one shared IOThread for the VM's virtio-scsi controller. This gives the most explicit separation, but it also increases host thread count the most.

What changes and what does not

AreaBehavior
Affectedvirtio-blk and virtio-scsi devices participate in the selected IOThread policy.
UnchangedSATA/AHCI, IDE, CD-ROM, floppy, kernel, and initrd handling remain unchanged. This setting is specifically about the virtio storage paths.
When it appliesThe change applies to newly started VMs. If you change the setting, restart affected VMs to launch them with the new I/O processing model.
Why Auto is usually the best next step

It introduces an IOThread only when the VM actually uses supported virtio storage, and it keeps the design simple by sharing one thread per VM instead of creating many threads.

Why Dedicated should stay specialized

On a dense host, a dedicated-thread model can substantially increase the total number of virtualization threads. That may be useful for targeted workloads, but it should be validated instead of enabled broadly by default.

Operator guidance

Use Compatibility when you need repeatable behavior across an existing estate. Use Auto when you want a low-friction improvement path for virtio-backed storage. Use Dedicated for focused testing, for known high-I/O guests, or when benchmarking shows a measurable benefit worth the additional host overhead.

Dedicated is not a free upgrade. More IOThreads can reduce contention inside a VM, but they also increase scheduling overhead on the host. On high-density systems, that can become a meaningful tradeoff.
stratum-host Runtime page showing the VM I/O processing setting
The Runtime menu exposes the VM I/O processing mode so operators can keep legacy behavior or enable IOThread-based virtio storage handling for newly started VMs.
06 TUI MenuRuntimeTLS Certificates Create demo certificates or validate externally managed PKI without overwriting it.

The Certificates menu controls the TLS material used by the controller and workers.

SettingMeaning
certificates.modeself-signed creates and preserves the demo-friendly CA/server/client bundle. external validates supplied files and never overwrites them.
certificates.directoryHost certificate directory mounted into the container. Default: /etc/stratum/certs.
certificates.demoP12PasswordPassword for the current demo PKCS#12 client bundle. Change only when the consuming workflow is also updated.
certificates.warningDaysHow early expiration health becomes a warning.

Useful certificate checks

stratum-host cert status
stratum-host cert trust
sudo stratum-host cert ensure

On a worker, status and trust perform a live controller TLS verification with /etc/stratum/secure/controller-ca.crt. Keep connect.insecureTLS disabled outside a disposable demo.

stratum-host Certificates menu
Certificates configuration in stratum-host.
07 TUI MenuNetworkingHost networking Network Continuum transport, pnet0 provider access, rollback safety, and nat0 DHCP/NAT.

The Network menu configures the Continuum underlay, the physical provider network, and the simple host-local NAT network.

SettingMeaning
connect.enabledStarts the STRATUM Continuum agent. Disable only for an intentional LAN-only worker design.
connect.interfaceContinuum interface name, normally stratum0.
connect.portContinuum UDP port, normally 51820.
storage.connectHostPortPublished Continuum UDP port for an isolated storage container. It must be unique when multiple storage containers share a host.
connect.payloadMTUPayload MTU carried through the Continuum. Default 1350 avoids common overlay fragmentation.
connect.pollSecondsWorker report or heartbeat interval.
connect.meshPollSecondsConditional mesh-state check interval.
connect.startupJitterSecondsRandom startup delay that spreads fleet enrollment and polling load.
connect.insecureTLSDevelopment escape hatch that weakens controller certificate verification. Leave false.
network.pnetModebridge gives direct physical-LAN attachment. routed is cloud/SSH-safe and leaves the management NIC untouched.
network.pnetBridgeNameProvider bridge name, normally pnet0.
network.pnetUplinkPhysical uplink. auto chooses the default-route interface.
network.pnetCIDR / pnetGatewayPrivate subnet and gateway used only in routed pnet0 mode.
network.pnetRollbackSecondsBridge mode safety timer. If validation does not complete, the old uplink configuration is restored.
network.pnetValidationTargetsOptional hosts, IPs, or host:port targets that must be reachable before bridge rollback is cancelled.
network.natEnabledCreates the host-local nat0 bridge plus DHCP and NAT egress.
network.natBridgeNameHost-local NAT bridge name, normally nat0.
network.natCIDR / natGatewayPrivate NAT subnet and guest gateway.
network.natDHCPStart / natDHCPEndGuest DHCP lease range.
network.natDNSDNS servers advertised to nat0 guests.

Physical bridge versus cloud-safe pnet0

Physical bridge — default
VM / tap → pnet0 bridge → physical NIC

Best when you control Layer 2 and want VMs directly on the LAN. Applying it can interrupt SSH while the uplink moves under the bridge.

Cloud-safe routed mode
VM / tap → pnet0 → host routing / NAT → uplink

Keeps the cloud- or provider-managed NIC intact. It prioritizes preserving access over extracting the last increment of direct Layer 2 performance during initial setup.

sudo stratum-host config set network.pnetMode routed
sudo stratum-host plan
sudo stratum-host apply --yes

Bridge mode arms a crash-resistant rollback before changing the management uplink. Inspect or control it with:

stratum-host network status
sudo stratum-host network confirm
sudo stratum-host network rollback
stratum-host Network menu
Network configuration in stratum-host.
08 ReferenceNetworkingCanvas networks STRATUM network types Bridge, Internal, Private, pnet0, and nat0 are intentionally different scopes.

These are network objects used inside the STRATUM canvas. They choose the Layer 2 domain to which a VM or virtual switch attaches.

Network objectScope and purpose
BridgeA normal, distinct Layer 2 network object inside the current virtual datacenter or lab. Each Bridge object creates its own broadcast domain. It has no DHCP or internet access unless you attach a service that provides them.
InternalA shared internal backbone for the current tenant and running lab/session. Multiple Internal attachments in that same session converge on the same segment. It remains separate from other lab sessions.
PrivateA tenant-scoped private backbone that can be reused across that tenant’s labs or virtual datacenters. It stays isolated from other tenants and is useful for common management or service networks.
Management (Cloud0) / pnet0pnet = Physical Network. Connects the virtual topology to the host’s provider network. In physical-bridge mode this reaches the external LAN directly; in cloud-safe mode it uses host routing/NAT.
NAT / DHCP (nat0)A host-local network with automatic DHCP and outbound NAT. It is never stretched across workers, and inbound access requires an explicit routing or forwarding design.
Simple distinction: Bridge is one explicit network object. Internal is shared inside one active vDatacenter/session. Private is shared for the tenant. pnet0 reaches the provider network. nat0 gives easy host-local DHCP and outbound access.
09 TUI MenuGpu PciScheduling Workers & GPU Dynamic registration, job accounting, admission, GRES, and advanced GPU controls.

This menu controls worker scheduling, admission, accounting, and expert GPU overrides.

SettingMeaning
slurm.accountingRecords worker job history for Compute Fabric reporting.
slurm.accountingPortTCP port used by worker job accounting.
slurm.jobAcctGatherFrequencyResource-sampling interval. Zero keeps job records but disables periodic samples.
slurm.dynamicNodestrue: workers register themselves dynamically. false: use legacy static worker definitions; the operator must manage worker host details manually in the Admin GUI worker configuration.
slurm.maxNodeCountController capacity ceiling for worker records. The default is intentionally large.
slurm.waitForFabricPrevents execution service admission until the Continuum is ready.
slurm.admissionPollSecondsHow frequently a worker rechecks Continuum execution admission.
slurm.gresRaw Slurm Generic RESources override. Normally leave blank so the GPU menu and inventory generate it.
slurm.gpuAutoDetectOptional Slurm discovery backend: nvidia, nvml, rsmi, or oneapi.
slurm.dynamicConfAdvanced raw dynamic-node configuration string.
gpu.policyFleet policy normally maintained by the GPU menu: auto, host-all, vfio-all, mixed, or disabled.
gpu.vendorLimit discovery to NVIDIA, AMD, Intel, or any supported vendor.
gpu.includeCompanionFunctionsIncludes audio/USB functions in the same PCI slot or safe IOMMU group when preparing passthrough.
gpu.allowBootDisplayGPUAllows passthrough of the GPU currently serving the host console. Dangerous unless another console path exists.
gpu.allowSharedIOMMUGroupsAllows passthrough when unrelated devices share an IOMMU isolation group. Advanced and normally unsafe.
gpu.runtimeDiscoveryControls runtime discovery. vfio-all resolves it off because devices are detached from vendor drivers.
gpu.vfioDevicesHost-specific PCI BDFs reserved for direct passthrough.
gpu.hostDevicesHost-specific GPUs intentionally excluded from STRATUM.

Dynamic Worker nodes

Keep slurm.dynamicNodes: true for normal scale-out. A new worker advertises its own CPU, memory, architecture, site, and GPU capabilities. Set it false only for an advanced static-node workflow where worker definitions are manually maintained.

What does GRES mean?

GRES means Generic RESources. STRATUM's scheduler service, Slurm, uses it for schedulable devices or capacity that are not ordinary CPU cores or RAM. STRATUM primarily uses it for GPUs.

gpu:a100:8

This means “the worker has eight schedulable GPUs of type A100.” A workload requesting A100 capacity will only be placed on an eligible worker with enough matching resources. Leave the override blank unless automatic inventory cannot represent a qualified custom resource.

stratum-host Workers and GPU menu
Workers & GPU configuration in stratum-host.
10 TUI PagesGpu PciHardware GPU assignment and PCI inventory Understand GPU Fabric, VFIO passthrough, Host Only, BDFs, and reboot requirements.

The GPU page is operational: it assigns each detected GPU. The PCI page is inventory-only in the current release; generic PCI passthrough, other than for GPUs, is intentionally not enabled.

GPU assignmentWhat happens
GPU FabricDefault. The GPU stays on its vendor driver and is served through stratum and the virtual STRATUM GPU hardware device.
PassthroughThe device is bound to vfio-pci and reserved for direct attachment to a VM on that same worker. A reboot is required when bindings change.
Host OnlyThe GPU remains on the host driver but is excluded from STRATUM inventory and scheduling.

VFIO, boot display, and BDFs

VFIO is the Linux framework used to hand a physical PCI device directly to a VM. Binding a GPU to vfio-pci removes it from the normal host graphics or compute driver.

A boot-display GPU is the adapter the host firmware or Linux console is using. Passing it through may blank the local screen and remove your emergency console. Enable gpu.allowBootDisplayGPU only when the host has another GPU, serial console, BMC/iDRAC/iLO, or a verified cloud console.

A PCI BDF is the device address in domain:bus:device.function form:

0000:65:00.0

Find the correct values with:

stratum-host gpu list
stratum-host pci list

Do not guess BDFs. They are host-specific and may change after hardware, firmware, or topology changes. The GPU page is the preferred way to maintain gpu.vfioDevices and gpu.hostDevices.

vfio-all is broad. It assigns all matching GPUs—and eligible companion functions—to passthrough. Keep the boot-GPU and IOMMU-group safety checks enabled unless the hardware has been qualified.

Existing NVIDIA MIG instances are detected as GPU Fabric scheduling units. STRATUM does not create or repartition MIG layouts in this release.

stratum-host GPU and PCI menu
GPU configuration in stratum-host.
11 TUI MenuHost SetupHost reconciliation Host Setup Container engine, images, networking, KVM, KSM, storage tuning, and firewall preparation.

Host Setup is the reconciler for Linux capabilities STRATUM needs. Every action has a check, a dry-run plan, and an apply path.

ActionMeaning
hostSetup.containerEngineInstalls Docker on Debian/Ubuntu or Podman on RHEL-family systems when no usable engine exists.
hostSetup.loadImagesLoads stratum-container.tar.
hostSetup.sharedDirectoriesCreates /etc/stratum and persistent storage roots.
hostSetup.sysctlLargeHostRaises memory-map, queue, and file-descriptor limits for large virtualization hosts.
hostSetup.sysctlNetworkApplies high-throughput TCP and network settings.
hostSetup.irqbalanceInstalls and enables IRQ balancing across CPUs.
hostSetup.fstrimEnables periodic SSD/NVMe trim where supported.
hostSetup.nvmeTuneUses the low-overhead none NVMe scheduler where supported.
hostSetup.pnetBridgeBuilds pnet0 in physical-bridge or cloud-safe routed mode.
hostSetup.pnetL2FilterInstalls Layer 2 bridge filtering rules.
hostSetup.pnetDot1xAllows required 802.1X link-local forwarding through the bridge.
hostSetup.kvmTuneEnables nested virtualization and safe KVM host defaults.
hostSetup.ksmEnables Kernel Same-page Merging memory deduplication.
hostSetup.ksmPagesNumber of memory pages KSM examines per scan batch.
hostSetup.ksmSleepMillisecondsDelay between KSM scan batches; lower is more aggressive and uses more CPU.
hostSetup.tunedInstalls and selects an appropriate host performance profile.
hostSetup.firewallApplies role-aware UFW or firewalld rules.

KSM is memory deduplication

Kernel Same-page Merging scans eligible anonymous memory, finds identical pages, and replaces them with one shared copy-on-write page. It can improve density when many VMs come from the same template. Spin up 10 STRATUM VM templates yet consume the memory footprint of only 2, maybe 3, depending on workloads.

KSM is not free RAM: it consumes CPU for scanning, saves little on dissimilar or encrypted pages, and may be disabled across hostile multi-tenant boundaries because memory deduplication can increase side-channel risk. Tune pages_to_scan and the sleep interval from measured workload behavior.

Open Memory Deduplication in the TUI to see the live savings estimate, scan progress, and the kernel settings currently controlling KSM.

sudo stratum-host plan
sudo stratum-host apply --yes
sudo stratum-host verify
stratum-host Host Setup menu
Host Setup configuration in stratum-host.
12 TUI PageHost SetupLive KSM telemetry Memory Deduplication See how much physical memory KSM is avoiding, how aggressively it is scanning, and whether the current workload is benefiting.

The Memory Deduplication page turns Linux Kernel Same-page Merging statistics into practical host-capacity information. KSM looks for identical eligible memory pages—often created when several VMs run the same operating system or originate from the same template—and keeps one shared copy until a workload writes to it.

Start with “Net physical memory avoided.” This is the closest single answer to “How much host RAM is KSM saving right now?” STRATUM uses the kernel's net-profit value when the running kernel exposes it, rather than presenting only a raw page count.

How to read the savings summary

FieldWhat it tells you
Net physical memory avoidedThe estimated physical RAM no longer required after duplicate pages are merged and KSM's own bookkeeping cost is considered. Use this as the primary savings figure.
Estimated workload reductionSTRATUM's estimate of how much less physical memory the KSM-eligible workload is consuming because of deduplication. It describes workload efficiency, not additional RAM assigned to the VMs.
Host RAM reclaimedNet memory avoided as a percentage of the host's total physical RAM. This shows whether deduplication is materially changing whole-host capacity.
Gross deduplicated pagesThe raw duplicate memory represented by merged page copies before net KSM overhead is deducted. Gross savings can therefore be higher than net physical memory avoided.
KSM stateRunning means the kernel scanner is actively looking for mergeable pages. The dashboard headline is shown only while KSM is running.
Completed scansThe number of complete passes KSM has made through its registered memory areas. This normally increases over time and confirms the scanner is progressing.

Kernel counters and scan behavior

FieldMeaning
Net calculationIdentifies the source used for the net estimate. kernel_general_profit means the host kernel supplied its KSM general-profit calculation.
Shared physical pagesThe number of physical pages currently retained as shared KSM pages.
Eliminated page copiesThe duplicate page copies currently replaced by those shared pages. This is the main raw indication that consolidation is occurring.
Zero pages eliminatedPages consolidated through the kernel's shared zero-page behavior. A value of zero is expected when use_zero_pages is disabled or no eligible zero pages are present.
Pages per scan batchThe current pages_to_scan setting. A larger value searches memory faster but can consume more CPU.
Sleep between batchesThe current sleep_millisecs delay. A shorter delay makes KSM more aggressive; a longer delay reduces scanner overhead.
Merge across NUMA nodesWhen enabled, identical pages may be merged across NUMA nodes for greater savings. NUMA-sensitive systems should balance that gain against memory locality.
Use kernel zero pagesShows whether the kernel may merge eligible empty pages into the shared zero page.
Workloads that usually benefit most

Multiple VMs built from the same image, similar operating systems, repeated application stacks, VDI pools, and other environments with substantial identical anonymous memory.

Workloads that may show little savings

Encrypted or compressed memory, highly diverse guests, rapidly changing working sets, and applications whose pages contain mostly unique data.

Enable, tune, and observe

Enable hostSetup.ksm in Host Setup, then plan and apply the host change. hostSetup.ksmPages controls the number of pages examined in each batch, and hostSetup.ksmSleepMilliseconds controls the pause between batches.

sudo stratum-host config set hostSetup.ksm true
sudo stratum-host plan
sudo stratum-host apply --yes
sudo stratum-host verify

Return to this page after workloads have been running long enough for KSM to complete several scans. Press R to refresh. Savings can rise or fall as VMs start, stop, modify shared pages, or change their working sets.

Measure before increasing scan aggressiveness. KSM exchanges CPU time for memory density. It can be valuable on trusted, similar VM fleets, but it should be evaluated carefully for latency-sensitive systems and across mutually untrusted security boundaries.
stratum-host Memory Deduplication page showing KSM savings, scan progress, and tuning values
The Memory Deduplication page separates practical host-memory savings from the underlying KSM counters and scan settings.
13 TUI MenusHost SetupAdvanced Controller Services, Demo Data & Advanced Know which switches are current, simulated, reserved, or development-only.

Controller Services

These controls describe the intended controller service boundary, but they are reserved placeholders in the latest STRATUM release and do not currently provision the services. That means these settings do nothing today.

SettingIntended purpose
controllerServices.chronyController-provided time synchronization.
controllerServices.coreDNSController-managed DNS service.
controllerServices.dhcpProvisioning or PXE DHCP. This is separate from the working nat0 guest DHCP service.
controllerServices.pxePXE, TFTP, and HTTP boot provisioning.
controllerServices.nfsShared image or content exports.
controllerServices.imageToolsController-side disk and image conversion tooling.

Demo Data

demo.enabled adds synthetic capabilities, including simulated GPU Fabric, for demonstrations, development, and scale testing. It still exercises the real guest, STRATUM's virtualized hardware GPU device, lease, and STRATUM GPU Fabric paths, but it is not a performance substitute for physical GPU hardware.

Advanced

SettingMeaning
security.adminDevelopmentBypassPasses STRATUM_ADMIN_DEV_BYPASS=1 into the STRATUM container for development-only admin startup or authentication testing. It may weaken the intended admin security boundary. Never enable it in production.
compatibility.hostIDEight-hex-character Cisco IOL compatibility host ID. It is unrelated to the worker FQDN or Continuum identity. STRATUM neither endorses or ships with anything from Cisco.
compatibility.hostIDEndianUses native or reversed byte order for the legacy IOL host ID.
runtime.debugStarts an interactive debug shell instead of the normal STRATUM startup path.
Advanced means exceptional. These settings exist for qualified compatibility, diagnostics, or development, not normal installation.
stratum-host Demo Data menu
Runtime configuration in stratum-host.
14 TUI PagesOperationsStatus Host actions and deployment roles Inspect one host action at a time and keep implemented roles separate from reserved ones.

Host Action Status

This page shows whether each enabled host feature is already correct, needs change, is disabled, or requires a reboot.

KeyAction
↑ ↓ or j kSelect an action.
EnterInspect details.
SpaceEnable or disable the selected action in configuration.
PPlan only the selected action.
AApply only the selected action.

Deployment Roles

RoleCurrent behavior
controllerRuns the control plane and can also participate in workload execution when configured.
workerProvides compute, memory, network, KVM, and GPU capabilities to managed scheduling.
storageRuns a dedicated STRATUM Storage replica container with its own identity, container name, persistent path, and host UDP port. Storage communication is over STRATUM Continuum.
fabric-gatewayReserved. Configuration is visible, but startup is intentionally blocked.
failover-controllerReserved. Configuration is visible, but startup is intentionally blocked.

A worker and storage container may share one physical host when they use separate configuration files, container names, storage roots, and non-conflicting storage.connectHostPort values.

stratum-host Host Action menu
Host Action Statusconfiguration in stratum-host.
15 Architecture NoteOperationsPersistent data Storage for STRATUM The platform automates VM storage access; you provide the durable, performant application storage beneath it.

In the current production architecture, STRATUM serves VM disk I/O through controller-managed storage services. The operator does not manually connect each VM disk—the platform APIs handle that—but the underlying persistent storage is still your responsibility.

STRATUM is Datacenter as an Application. Applications need storage. Give storage.root durable, adequately sized, protected, and fast storage; STRATUM handles the rest of the virtual-datacenter relationships.

Standalone Linux

The default persistent root is /persistent/stratum. Back it with the design appropriate to the deployment: fast local SSD/NVMe, RAID, DAS, a NAS mount, SAN-backed filesystem, or another qualified backend. Capacity, latency, redundancy, snapshots, backups, and disaster recovery remain infrastructure decisions.

Kubernetes

Use the appropriate PersistentVolume and StorageClass design. Kubernetes can mount the storage into the STRATUM workload, but the cluster operator must still select the performance, topology, access mode, backup, and retention behavior.

Inspect and repair

stratum-host storage status
sudo stratum-host storage fix-permissions

The permission repair resolves the numeric stratum UID/GID from the installed OCI image and fixes mutable data ownership. It does not rewrite certificate files.

Do not size only for the container image. Persistent state can include databases, virtual datacenters, templates, VM disks, snapshots, Arsenal content, logs, certificates, and recovery-critical configuration.
16 OperationsOperationsLifecycle Start, stop, reconcile, verify, and view logs Daily commands keep the container predictable without hiding persistent state.

Normal lifecycle

sudo stratum-host start
sudo stratum-host stop
sudo stratum-host restart
sudo stratum-host start --reconcile
sudo stratum-host restart --reconcile

--reconcile recreates a container when the selected image or configuration has drifted, while retaining persistent data. Image drift is the local container image being updated, or a container registry with a new updated STRATUM image. Just as regular containerized applications get new updates or pushing updates to a Kubernetes cluster.

Status and logs

stratum-host status
stratum-host status --deep
stratum-host status --json --strict
stratum-host logs --tail 200
stratum-host logs --follow

status --deep includes relevant listening sockets and the tail of the in-container supervisor log. Container output is available through stratum-host logs. Persistent logs normally appear below <storage.root>/var/log, which is /persistent/stratum/var/log with defaults.

Validation and drift

stratum-host config validate
stratum-host config diff
stratum-host verify
stratum-host verify --json

Remove the container, not the data

sudo stratum-host remove --yes

Removal deletes the configured container. Persistent host data and certificates are retained.

17 ReferenceCli ReferenceComplete CLI arguments and usage Every supported command in the reviewed stratum-host source.

Global options must appear before the command:

stratum-host [--config PATH] [--set KEY=VALUE] COMMAND
CommandPurpose
stratum-host / tuiOpen the interactive host manager.
help [COMMAND]Show general or command-specific help.
start [--reconcile]Create/start the configured container; optionally recreate drifted configuration.
stopStop without deleting.
restart [--reconcile]Stop and start; optionally recreate drifted configuration.
remove --yesRemove the container while retaining persistent data and certificates.
status [--json] [--strict] [--deep]Show resolved role, runtime, certificate, hardware, network, host actions, sockets, and reboot state.
logs [--follow] [--tail N]Show Docker/Podman container output.
planRun the same host-action engine as apply, but with dry-run operations.
apply --yesApply enabled host actions. Root is required.
verify [--json]Check enabled host features without changing the host.
config init [--force]Create a default YAML configuration.
config showPrint the resolved YAML.
config validateValidate settings and resolved deployment configuration.
config diffShow differences from defaults.
config settingsList every registry key, type, default, help text, and environment mapping.
config set KEY VALUEPersist one setting atomically to the selected host.yaml.
cert status|trust|ensure [--json]Inspect, verify, or create/validate certificate material.
network status|confirm|rollback [--json]Inspect and control a pending physical-bridge rollback transaction.
storage status|fix-permissions [--json]Inspect or repair persistent storage ownership and modes.
pci list [--json]List BDF, IDs, class, driver, IOMMU group, and boot-VGA state.
gpu list [--json]List GPUs matching the configured vendor.
rolesList implemented and reserved roles.
versionPrint application version and build revision.

Global options

OptionBehavior
--config PATHUse an alternate configuration file, such as a dedicated storage-role YAML.
--set KEY=VALUERepeatable one-run override. It does not modify host.yaml.

Configuration precedence is: built-in defaults, saved YAML, supported environment variables, then one-run --set overrides. Use config set for a persistent change.

18 ExamplesAutomationCLI Automation patterns Use the same plan, apply, verify, and reconcile model from one host to an enterprise fleet.

Automation-safe host preparation

sudo ./stratum-host --config ./host.yaml plan
sudo ./stratum-host --config ./host.yaml apply --yes
sudo install -m 0755 ./stratum-host /usr/local/sbin/stratum-host
sudo install -m 0600 ./host.yaml /etc/stratum/host.yaml
sudo stratum-host config validate
sudo stratum-host verify --json
sudo stratum-host start --reconcile
sudo stratum-host status --json --strict

For unattended runs, treat a non-zero exit as failure and retain the plan output with the change record.

One-run overrides

sudo stratum-host \
  --set role=worker \
  --set identity.site=us-east \
  --set network.pnetMode=routed \
  start --reconcile

These values affect only that invocation. Persist them with config set or write the desired YAML through your preferred configuration-management system.

Separate worker and storage containers on one host

sudo stratum-host --config /etc/stratum/host.yaml start --reconcile
sudo stratum-host --config /etc/stratum/storage.yaml start --reconcile

Give the storage configuration its own container name, persistent role directory, node identity, and unique storage.connectHostPort.

Recommended pattern: configuration management copies the binary, OCI tarball, YAML, controller CA, and enrollment material; runs plan; applies approved host changes; starts with --reconcile; then records JSON verification and status.
No matching sections.Try a broader term or clear the category filter.
Operational note
Host networking, VFIO binding, firewall changes, storage permissions, and container reconciliation can affect running workloads. Plan, validate, and maintain an out-of-band recovery path for remote hosts.
STRATUM HOST

Helm chart details

STRATUM / KUBERNETES

Prepare the cluster. Deploy the datacenter.

The STRATUM Helm chart deploys a persistent controller and an optional worker fabric onto Kubernetes nodes that have already been prepared for virtualization, Continuum networking, provider networking, and host-level device access.

15 sections shown
01 GuideKubernetesStart here What this Helm chart deploys Kubernetes runs STRATUM; node preparation remains outside Kubernetes.

The chart deploys one STRATUM controller StatefulSet and, optionally, one worker DaemonSet. It generates each role’s /etc/stratum/host.yaml, mounts certificates and enrollment material from Secrets, initializes controller storage, and schedules STRATUM on prepared nodes.

Host preparation
stratum-host
Golden image
Ansible / Puppet / Chef
Node provisioning

Prepare KVM, TUN, pnet0, IOMMU/VFIO, GPU drivers, firewall policy, kernel settings, and the host itself.

Kubernetes runtime
Helm
StatefulSet controller
DaemonSet workers
PVC + Secrets + Services

Schedule and operate the already-prepared STRATUM infrastructure workload.

This is not a normal restricted application pod. STRATUM needs privileged host access, host networking, host PID visibility, KVM, TUN, and selected host filesystems. Use Kubernetes you control. Locked-down managed services commonly block one or more required capabilities.

STRATUM helm chart requires Kubernetes 1.25+.

02 ChecklistBefore InstallRequired Kubernetes prerequisite checklist Complete these items before running Helm.
RequirementWhat to verify
Controlled clusterThe cluster permits privileged infrastructure pods, hostNetwork, hostPID, hostPath devices, and the required namespace security policy.
VirtualizationSelected nodes expose /dev/kvm. Hardware virtualization must be enabled, and nested virtualization must be available when Kubernetes itself runs in VMs.
Continuum networkingSelected nodes expose /dev/net/tun, can resolve the controller hostname, and permit the STRATUM ports between nodes.
Provider networkpnet0 is prepared on every selected node when VMs need provider/LAN access. Helm only checks it; Helm never creates it.
Persistent storageA default StorageClass exists, a chart-specific StorageClass is selected, or an existing PVC is supplied.
Image availabilityThe STRATUM OCI image is in a reachable registry, authenticated with an imagePullSecret, or preloaded on every selected node.
Placement labelsExactly one controller node has stratum.io/role=controller; worker nodes have stratum.io/role=worker.
DNS and TLSstratum.publicHost resolves where users and workers need it, and the controller certificate matches that hostname.

Useful node checks

test -c /dev/kvm && echo KVM-ready
test -c /dev/net/tun && echo TUN-ready
ls -ld /sys/class/net/pnet0/bridge
kubectl get storageclass
kubectl get nodes --show-labels
Do not make Helm discover these failures for you. Validate the cluster and nodes first, especially storage binding, KVM, TUN, pnet0, and namespace policy.
03 GuideStorageCritical Give STRATUM persistent storage The controller PVC is the Kubernetes equivalent of the host’s persistent STRATUM directory.

STRATUM is a datacenter as an application. Like any stateful application, it needs durable storage. The chart uses one controller PVC and mounts subdirectories into the runtime for:

/stratum/var
/stratum/labs
/stratum/data
/stratum/arsenal
/stratum/opt/mariadb/data
/stratum/auth/config/users.yml
/stratum/proxy/certs        # self-signed mode
Size and performance matter. The current STRATUM controller serves VM storage for the fabric. The PVC backend must be appropriate for VM templates, ISO images, virtual disks, databases, and normal controller state. A small or slow general-purpose volume may deploy successfully but perform poorly in real use.
ValueMeaning
controller.persistence.enabledKeep controller state outside the replaceable OCI image. Leave enabled for normal deployments.
initializeSeed image defaults into empty PVC paths. Existing content is never overwritten during upgrades.
storageClassSelect the Kubernetes storage backend. Empty uses the cluster default.
existingClaimMount a PVC you created and manage separately.
accessModesDefaults to ReadWriteOnce, which matches a single controller pod.
sizeDefaults to 100Gi. Treat that as a demo starting point, not a production sizing recommendation.

Production storage questions

  • Is the backend fast SSD/NVMe, RAID, SAN, NAS, or another persistent platform suitable for VM I/O?
  • Can the volume reattach if the controller is rescheduled, or must the controller remain pinned to one node?
  • Are snapshots, backups, replication, capacity monitoring, expansion, and reclaim policy configured?
  • Does the backend have enough capacity for future templates, ISO images, VM disks, and database growth?
controller:
  persistence:
    enabled: true
    initialize: true
    storageClass: fast-rwo
    accessModes: [ReadWriteOnce]
    size: 2Ti

STRATUM handles the virtual datacenter above this layer. Kubernetes and your storage platform remain responsible for the durability and performance of the PVC beneath it.

04 RunbookCluster SetupPrivileged Prepare the namespace Pod Security Admission must permit STRATUM’s host-level runtime.

Create and label the namespace before installation. Using only helm --create-namespace does not add the required Pod Security labels.

kubectl create namespace stratum
kubectl label namespace stratum \
  pod-security.kubernetes.io/enforce=privileged \
  pod-security.kubernetes.io/audit=privileged \
  pod-security.kubernetes.io/warn=privileged

On OpenShift, grant the chart ServiceAccount an approved privileged SCC according to your organization’s policy.

ServiceAccount tokens are off by default. serviceAccount.automountServiceAccountToken=false because STRATUM does not need to call the Kubernetes API for normal operation.
Security boundary: the chart defaults to privileged execution, hostNetwork: true, hostPID: true, host /sys, host cgroups, /dev/kvm, and /dev/net/tun. Review this as an infrastructure workload, not as a tenant application.
05 RunbookNode SetupBefore Helm Prepare and label controller and worker nodes Use stratum-host or your own node automation, then control placement with labels.

The chart intentionally disables all hostSetup actions inside the pod. Prepare each selected node before it can run STRATUM.

# On each selected node
sudo stratum-host plan
sudo stratum-host apply --yes

# Kubernetes placement
kubectl label node ctrl1 stratum.io/role=controller
kubectl label node worker001 stratum.io/role=worker
Do not label the controller node as a worker. Both roles use host networking and may contend for fixed STRATUM ports. The chart deliberately pins the controller away from worker-labelled nodes.

What node preparation must handle

LayerHandled outside Helm
Firmware and kernelBIOS virtualization, KVM modules, IOMMU, VFIO readiness, kernel tuning, KSM, and device permissions.
Networkingpnet0 bridge or routed provider network, firewall rules, MTU planning, and any physical uplink changes.
GPUGPU drivers, IOMMU grouping, VFIO binding, and selection of host-display versus passthrough devices.
Image deliveryContainer image registry access, registry credentials, or image preload on disconnected nodes.

Golden images, Cluster API bootstrap, MachineConfig, Ansible, Puppet, Chef, Terraform-driven provisioning, or another fleet system may replace direct use of stratum-host as long as the resulting node contract is the same.

06 GuideNetworkingHost prepared pnet0, host networking, DNS, and ports Helm describes the expected network; the node supplies it.

hostAccess.hostNetwork=true is required and enforced by the chart. STRATUM binds directly to the Kubernetes node network rather than relying only on pod-network address translation.

Bridge mode
VM → pnet0 → physical LAN

Default VMware-style behavior. Prepare the physical bridge on each node. Initial host setup may interrupt SSH while the uplink moves under the bridge.

Routed mode
VM → pnet0 → host routing/NAT → uplink

Cloud- and SSH-safe mode. The management NIC remains intact, prioritizing node reachability during setup.

# Cloud-safe node preparation
sudo stratum-host config set network.pnetMode routed
sudo stratum-host plan
sudo stratum-host apply --yes

The Helm values must match the already-prepared host:

stratum:
  network:
    pnetMode: bridge       # bridge or routed
    pnetBridgeName: pnet0

hostAccess:
  pnet:
    requireExisting: true

requireExisting adds an init-container check for /sys/class/net/pnet0/bridge. It fails fast when the bridge is absent or not a Linux bridge. It does not create, convert, repair, or reconfigure anything.

Default network endpoints

PortUse
8443/TCPSTRATUM web interface, APIs, and controller enrollment. This port is currently fixed by chart validation.
51820/UDPDefault STRATUM Continuum transport. The chart value is configurable.
6817/TCPWorker scheduling/control service.
6818/TCPWorker runtime service.
6819/TCPAccounting service.

Allow the applicable traffic between controller and workers, and ensure every worker can resolve the controller hostname. Because pods use host networking, port conflicts occur on the node itself.

07 GuideImagesOCI Make the STRATUM image available Use a versioned registry image, a digest, or preload it for disconnected operation.
ValueUse
image.repositoryRegistry path or local image name.
image.tagRelease tag. Pin a tested version in production.
image.digestOptional immutable digest; when set, it takes precedence over the tag.
image.pullPolicyIfNotPresent for normal use; Never for preloaded disconnected nodes.
image.pullSecretsSecrets used by Kubernetes to authenticate to a private registry.

Private registry

kubectl -n stratum create secret docker-registry stratum-registry \
  --docker-server=registry.example.com \
  --docker-username='USERNAME' \
  --docker-password='PASSWORD'
image:
  repository: registry.example.com/stratum/stratum
  tag: r018
  pullPolicy: IfNotPresent
  pullSecrets:
    - name: stratum-registry

Disconnected cluster

image:
  repository: stratum
  tag: r018
  pullPolicy: Never

Preload the exact image on every controller and worker node. The controller initialization container uses the same image and does not need internet access.

08 RunbookInstallController first Deploy the controller A production-friendly installation sequence begins with the controller and persistent storage.

1. Create production values

image:
  repository: registry.example.com/stratum/stratum
  tag: r018
  pullPolicy: IfNotPresent

stratum:
  publicHost: ctrl1.stratum.lab
  site: datacenter-a

controller:
  enabled: true
  certificates:
    mode: self-signed
  persistence:
    enabled: true
    initialize: true
    storageClass: fast-rwo
    size: 2Ti

workers:
  enabled: false

hostAccess:
  pnet:
    requireExisting: true

2. Validate rendering

helm lint ./stratum -f values-production.yaml
helm template stratum ./stratum \
  --namespace stratum \
  -f values-production.yaml > rendered.yaml

3. Install

helm upgrade --install stratum ./stratum \
  --namespace stratum \
  -f values-production.yaml

4. Confirm readiness

kubectl -n stratum get pods,pvc,svc
kubectl -n stratum get pods \
  -l app.kubernetes.io/component=controller
kubectl -n stratum logs \
  -l app.kubernetes.io/component=controller \
  -c stratum --tail=200

Open https://<stratum.publicHost>:8443/stratum/ after DNS, firewall policy, service exposure, and certificates are correct.

Initialization is one-time and non-destructive. The init container seeds empty PVC directories from the immutable STRATUM image, fixes ownership, and leaves existing data untouched on later upgrades.
09 GuideSecurityTLS Certificates, public DNS, and controller access The public hostname is part of both connectivity and certificate validation.

stratum.publicHost is used for the public URL, controller endpoint advertisement, STRATUM server name, and certificate expectations. Choose a stable hostname that users and workers can resolve.

Self-signed
controller:
  certificates:
    mode: self-signed

The controller creates demo CA, server, client, and PKCS#12 material in the persistent certificate directory.

External PKI
controller:
  certificates:
    mode: external
    existingSecret: stratum-controller-certs

Kubernetes mounts your supplied certificate Secret read-only. STRATUM does not overwrite it.

External certificate Secret

kubectl -n stratum create secret generic stratum-controller-certs \
  --from-file=ca.crt \
  --from-file=server-cert.key \
  --from-file=server-key.key \
  --from-file=client-cert.key \
  --from-file=client-key.key \
  --from-file=demo.p12
The certificate must match stratum.publicHost. A worker may reach the controller IP and still fail enrollment when the certificate name, public hostname, or CA is wrong.

The chart creates a public controller Service for HTTPS and Continuum, a headless Service for StatefulSet identity, and an internal same-name Service for worker control and accounting. Configure controller.service.type and any platform-specific annotations when external service exposure is required.

10 RunbookWorkersTwo stage Add worker nodes Create enrollment material after the controller is available, then enable the DaemonSet.

Workers require a Kubernetes Secret containing the controller CA and an enrollment token. A self-signed controller plus workers is normally a two-stage deployment:

  1. Deploy the controller.
  2. Generate or download enrollment material from the STRATUM administrator interface.
  3. Create the worker Secret.
  4. Enable workers and upgrade the release.
kubectl -n stratum create secret generic stratum-worker-enrollment \
  --from-file=stratum-enroll \
  --from-file=controller-ca.crt=ca.crt
workers:
  enabled: true
  existingSecret: stratum-worker-enrollment
  updateStrategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 10%
helm upgrade stratum ./stratum \
  --namespace stratum \
  -f values-production.yaml

Each worker’s STRATUM identity comes from Kubernetes spec.nodeName. Use stable, unique Kubernetes node names and working controller DNS.

Dynamic workers default to enabled. Workers enroll and appear automatically. When stratum.slurm.dynamicNodes=false, worker definitions must be managed manually in the STRATUM administrator GUI workflow; this is intended for more controlled or advanced deployments.

Workers-only release

A separate workers-only Helm release is supported, but it must set controller.enabled=false, stratum.slurm.controllerHost, the controller endpoint under stratum.connect.controllerHost, and the worker enrollment Secret.

11 GuidePlacementRegions Use sites to divide the enterprise fabric A site label groups workers into a region or facility for default VM placement.

stratum.site may be any useful region or facility name, such as:

default
us-east
boston
tokyo
taiwan
edge-plant-07

Workers report the site during enrollment. When a virtual datacenter or VM is deployed to that site, STRATUM can prefer the matching worker pool by default. This lets one enterprise fabric contain multiple geographic regions or operational zones while keeping placement intentional.

One Helm value applies to the release. All workers in the same release receive the same stratum.site. Use separate worker releases or separate clusters/values when worker groups must report different sites.
stratum:
  site: us-east
12 ReferenceConfigurationKey values Important values and why they matter The settings most operators should review before production installation.
ValuePlain-English meaning
stratum.publicHostBrowser- and worker-facing controller hostname. It must resolve and match the certificate.
stratum.siteRegion/facility label reported by controller and workers.
stratum.connect.*Continuum interface, UDP port, mode, payload MTU, TLS behavior, and optional controller hostname.
stratum.network.pnetModeWhether the already-prepared provider network is direct bridge mode or cloud-safe routed mode.
stratum.network.nat*nat0 bridge, subnet, gateway, DHCP range, and DNS presented to NAT-connected guests.
stratum.slurm.dynamicNodesAutomatically admit discovered workers. False requires manual worker definitions.
stratum.slurm.maxNodeCountMaximum dynamically managed worker capacity. Default is 65536.
stratum.gpu.discoveryauto inventories GPUs exposed to the pod; off disables discovery.
controller.nodeSelectorDefaults to stratum.io/role: controller.
controller.resourcesKubernetes CPU and memory requests/limits. Define these for production scheduling and capacity control.
controller.serviceControls ClusterIP, LoadBalancer, annotations, traffic policy, and optional load-balancer IP.
workers.nodeSelectorDefaults to stratum.io/role: worker.
workers.resourcesRequests/limits for each worker pod. Avoid limits that starve host-level virtualization workloads.
hostAccess.pnet.requireExistingFail fast if the configured pnet bridge is absent.
serviceAccount.automountServiceAccountTokenDefaults false because normal STRATUM runtime does not need Kubernetes API credentials.
Generated host configuration is read-only. Helm renders /etc/stratum/host.yaml from values. Make persistent changes in values and run helm upgrade; do not edit the file inside the pod.
13 GuideHardwareOptional GPU, VFIO, host devices, and extra mounts The chart exposes prepared hardware; it does not configure the hardware for you.
ValueMeaning
hostAccess.devices.kvmMount /dev/kvm for hardware-accelerated virtual machines. Enabled by default.
hostAccess.devices.tunMount /dev/net/tun for the STRATUM Continuum. Enabled by default.
hostAccess.devices.vfioMount the host /dev/vfio tree for devices already bound to VFIO. Disabled by default.
stratum.gpu.discoveryDiscover GPUs visible from the host-mounted environment.
stratum.slurm.gresOptional explicit Generic RESources declaration, such as gpu:a100:8: eight A100-class GPUs on that worker definition.
stratum.slurm.gpuAutoDetectOptional lower-level GPU discovery override used by the current worker startup contract.
hostAccess.extraVolumesAdd Kubernetes volumes not modeled by the chart.
hostAccess.extraVolumeMountsMount those additional volumes into the STRATUM container.
Enabling VFIO is not VFIO setup. The node must already have IOMMU enabled, devices isolated into acceptable IOMMU groups, the intended PCI functions bound to vfio-pci, and any host-display GPU decisions resolved before the pod starts.

The optional runtime socket mount is disabled by default. Enable it only for a specific supported integration that truly needs a local Docker-compatible socket; exposing the host container socket gives the pod broad host control.

14 GuideSecretsOptional Licensing and sensitive material Use Kubernetes Secrets for certificates, enrollment, registry credentials, and optional license files.
SecretRequired keys
Controller certificatesca.crt, server-cert.key, server-key.key, client-cert.key, client-key.key, and demo.p12.
Worker enrollmentstratum-enroll and controller-ca.crt.
Controller licensestratum.license, referenced by controller.licenseSecret.
Worker licensestratum.license, referenced by workers.licenseSecret.
RegistryStandard Kubernetes Docker registry authentication Secret referenced by image.pullSecrets.
kubectl -n stratum create secret generic stratum-controller-license \
  --from-file=stratum.license

# values.yaml
controller:
  licenseSecret: stratum-controller-license

When a license Secret is not supplied, STRATUM uses its persistent license path under /stratum/var/lib/stratum-license/. Keep Secrets out of values files and source control.

15 ReferenceOperationsDay two Upgrade, inspect, roll back, and troubleshoot Use Helm for desired state and Kubernetes for runtime inspection.

Common operations

# Current values and release history
helm -n stratum get values stratum
helm -n stratum history stratum

# Upgrade
helm upgrade stratum ./stratum \
  --namespace stratum \
  -f values-production.yaml

# Inspect
kubectl -n stratum get pods,pvc,svc -o wide
kubectl -n stratum describe pod <pod-name>
kubectl -n stratum logs <pod-name> -c stratum --tail=300

# Roll back chart resources
helm -n stratum rollback stratum <revision>
SymptomCheck first
Controller pod stays PendingController node label, PVC binding, StorageClass, node affinity, taints/tolerations, and privileged admission.
pnet preflight failsThe configured bridge does not exist on the scheduled node or is not a Linux bridge. Prepare the node; do not try to repair it from Helm.
Pod fails to create/dev/kvm, /dev/net/tun, /sys, cgroups, and optional /dev/vfio hostPath availability.
Worker will not enrollWorker Secret keys, controller CA, enrollment token, controller DNS, certificate hostname, 8443/TCP, and Continuum UDP connectivity.
UI is unreachablestratum.publicHost, DNS, firewall, Service type/annotations, controller node reachability, and certificate validity.
Upgrade did not refresh seeded filesExpected behavior: initialization copies defaults only into empty destinations. Existing persistent data is preserved.
Managed Kubernetes rejects the chartThe platform may prohibit privileged pods, host networking, host PID, KVM/TUN, nested virtualization, or prepared pnet bridges.

Uninstall carefully

helm uninstall stratum --namespace stratum
Uninstall is not a backup policy. Verify the controller PVC and the StorageClass reclaim policy before deleting anything. Preserve or snapshot the volume when the virtual datacenter must survive chart removal.

Current chart boundaries

fabricGateway.enabled and failoverController.enabled are reserved placeholders and deliberately fail validation when enabled. The current chart deploys one controller replica; controller failover is not implemented by these values.

No matching sectionTry a broader search or clear the active filter.
Deployment boundary
STRATUM handles the virtual datacenter. Kubernetes schedules it. Your node, network, and storage platforms provide the foundation beneath it.
STRATUM / HELM

PATENT PENDING