Skip to main content

Product · Scale AI & bare metal

Deploy a private AI cloud like it's one product.

Run your own AI cloud on GPUs you already own, instead of renting capacity from a hyperscaler — deployed, managed, and scaled across data center, edge, and hybrid. Zynera is the Kubernetes-native platform underneath, built on open standards, with no lock-in.

See Zyra decide · The 8 components

GPU pools & hard quotaseBPF network intelligencePrivate AI cloud · zero egressApproval-gated AI ops (Zyra)Open standards, no lock-in
Why it exists

Zynera

Your AI infrastructure control plane: a private model gateway, governed releases, RAG, GPU scheduling and cost control. Idle GPU capacity is money burning on a rack — and without tenant isolation, quota enforcement, and topology-aware scheduling, that's exactly what happens. Zynera's scheduler policies push utilization up across every tenant sharing the pool, while Zyra reasons over placement, cost, and security as one picture instead of a dozen disconnected tools.

Feature pillars

What you get

Kubernetes-native GPU control

GPU pools, tenants, quotas, VMs with GPU live migration, and extended workload types as Kubernetes resources.

Private AI gateway

An OpenAI-compatible /v1 endpoint on your own GPUs: guardrails, routing, caching, weighted model aliases with canary and rollback, batch, speech-to-text, and image generation.

Evaluation-gated promotion & RAG

Promotion to production is blocked unless the vertical’s golden sets pass (admin override needs a written rationale). RAG Studio adds connectors, citations, and federated retrieval.

Zyra — embedded AI ops engineer

Approval-gated autonomous workflows, multi-agent specialists, and hybrid RAG grounded in your own cluster state. Actions stay advisory until a human approves.

Network intelligence

eBPF-powered flow visibility and security governance across the fleet.

FinOps & federation

Chargebacks, budget guardrails, and cost-preferring placement across clusters.

Differentiators

Built for production

Compute — GPU fabric after cluster bootstrap with HyperCluster.

Why teams choose us

  • Production-ready out of the box — operators and resources to start from, not a blank slate
  • Local AI platform — chat, a license-tagged open-weight model hub, per-key usage metering, and an OpenAI-compatible /v1 gateway on your own GPUs, zero external LLM key required
  • Governed model release — promotion to production is gated on golden-set evaluation, with every decision recorded
  • Pairs with IronWolf bare metal and Zeus OS VM GPU passthrough
  • Self-hosted economics vs always-on public GPU clouds

When to choose

  • You run AI/ML on owned Kubernetes with NVIDIA GPUs
  • You want a private OpenAI-compatible endpoint with guardrails and evaluation-gated model releases, on hardware you own
  • You need multi-tenant GPU quotas rather than best-effort sharing
FAQ

Zynera questions

How it flows

How it works

CLIzynera gpu inspect NODE --rdma

Migrate from the platforms you already run

  • Enterprise Hypervisors
  • HCI Platforms
  • AWS
  • Azure
  • Google Cloud
  • Windows Hypervisor
  • Oracle Cloud
  • Machina fleet cloud
  • KubeVirt
  • KVM
  • KVM-based Platforms
SOC2-ready postureRBAC + audit logging99.9% SLA targetAir-gap programs

Architecture

One operator per layer. One control plane.

Every layer of your GPU infrastructure — scheduling, storage, networking, multi-tenant cost control — gets its own purpose-built Kubernetes operator, instead of being stitched together by hand.

GPU Operator

NVIDIA-native GPU scheduling with MIG partitions, time-slicing, topology awareness, and automatic driver management.

AI Operator

Elastic training and workload orchestration for distributed AI and ML pipelines.

Storage Operator

Lustre, GPFS, and BeeGFS integration to keep your GPUs fed with data.

Network Operator

InfiniBand and RoCE v2 support with automatic NVLink tuning for fast GPU-to-GPU traffic.

Quota Operator

Multi-tenant GPU quota management with cost tracking and chargeback across teams and projects.

Network Intelligence

See GPU traffic patterns and bandwidth use, and enforce network policy, with eBPF-powered observability.

Virtualization Operator

Run virtual machines alongside GPU containers on the same control plane.

Private AI cloud

Host your own AI cloud — LLMs on your GPUs, nothing phones home.

Zynera is a self-hosted inference platform. Serve open models on your own hardware, run them locally, and keep every token, prompt, and embedding inside your cluster.

Many serving backends, one spec

Pick the serving engine per model — Ollama, vLLM, TensorRT-LLM, Triton, SGLang, or TorchServe — from one spec. Scale to zero when idle and canary new versions safely.

Ollama, built in

Pull and serve Llama, Mistral, Qwen, and more, managed for you — no external API, no per-token bill.

Private routing by policy

Sensitive prompts are routed to a local model automatically, so they never leave the building even when cloud models are also configured.

Embeddings on your hardware

Chunk, embed, and index your documents for RAG entirely on your own hardware — no hosted embedding API.

Air-gapped Zyra

Zyra, the AI ops engineer, can run on a local model — it keeps working with zero internet access.

Model serving in the docs

No per-token billingNo prompt egressOllama · vLLM · SGLangOpenAI-compatible endpointLocal-only inference option

In the console

See it in the console.

Screens from a Zynera install: model catalog, Mission Control, and the live pod console.

Zynera Model Fabric Studio listing catalog model templates with Deploy and Train actions
Model Fabric Studio — catalog templates with Deploy and Train
Zynera AI Mission Control showing model catalog, deployments, training jobs, and spend tiles
AI Mission Control — model catalog, deployments, training jobs, and spend
Zynera live pod console with running, pending, and failed pod counts
Live pod console

AI platform

A full AI platform, not just a scheduler.

One dashboard for the whole AI lifecycle — build, ship, and govern models on the same control plane that schedules your GPUs.

Build

Author, fine-tune, and test models and agents without leaving the platform.

Operate

Promote a model to a served endpoint and watch it under load — no CLI required.

Govern

Traces, policy, chargeback, and multi-tenant workspaces in one governed surface.

The differentiator

Zyra — an AI ops engineer that works under approval.

Most infrastructure platforms automate infrastructure. None ship an embedded, approval-gated AI SRE that reasons about your fleet, proposes an action, shows its blast radius, and waits for a human to sign off.

Watch Zyra work through a GPU imbalance.

Zyra scans fabric telemetry, spots an imbalance, and proposes a topology-aware rebalance with its blast radius shown up front. Nothing executes until a human clicks Approve. This is a scripted illustration with invented numbers, not a live cluster.

Part of Zyra — see the same AI persona across the whole suite

zyra · session #4192 · zynera-system

Open source companion

Test Zynera scheduling without physical GPUs

Zyvor Janus is the Apache 2.0 discrete-event simulator for Zynera — MIG, topology, quotas, and LLM serving metrics on your laptop.

Zynera + Zyvor Janus, side by side (~3 min)

Intelligence plane

Observe, learn, and act on GPU fleets

FinOps waste detection, placement bandit, training/inference optimizers, and NL cluster copilot with policy-gated remediation.

Cluster copilot

Natural-language queries and remediation suggestions with policy guardrails — closed-loop auto-tuning without blind automation.

FinOps & placement

Waste detection, cost rollups, and bandit-based placement recommendations tied to actual GPU utilization.

FabricAutoPolicy

Orchestrator tick reconciles extended CRDs into auto-remediation — federation APIs for multi-cluster AI factory operations.

Platform surface

CLI, API gateway, and dashboard

zynera CLI

Rust CLI for jobs, quotas, queues, GPU inspection, network, and security operations.

API gateway & SDKs

FastAPI gateway for UI and automation; Python and Go SDKs for custom controllers.

101-page web UI

Jobs, quotas, nodes, costs, network flows, policies, workspaces, model registry, inference, and auto-tuning.

Extended CRDs & eBPF

Kubernetes-native GPU fabric control, kernel-level eBPF visibility, and Grafana views for observability.

Resources

Signature deck

Download the h2kvm-format PDF or open the HTML preview for stakeholder reviews.

Overview

A4 portrait · purple cover · orange accent (h2kvm style)

Browse all 9 Zynera decks

30-day free trial

Install in 30 seconds — no sign-up required.

Install Zynera with a single Helm chart. No account, no email, no sales call.

Kubernetes1.28 and later
GPU nodesNVIDIA GPU nodes
Architecturex86-64 or ARM64
Network fabricInfiniBand / RoCE optional
InstallSingle Helm chart
Sign-upNone — the 30-day trial starts automatically
Step 1 — Install
helm install zynera ./charts/zynera-core \
  --namespace zynera-system \
  --create-namespace \
  --set namespace.create=false

Full docs, including the eight components of the AI Infrastructure OS: Read the Zynera docs

Free trial

30-day free trial, no sign-up — Zynera

Run your own AI cloud on GPUs you already own — Kubernetes-native, no lock-in.

  • One Helm install, GPU nodes detected automatically
  • Trial clock starts from the build date — no account, no sales call
  • Licence key from [email protected] after the trial

After the trial window, it is completely up to you whether to continue. There is absolutely no pressure or obligation from our side.

Ready to own your GPU compute?

Kubernetes-native orchestration for the GPUs you already own — fully self-hosted.