Zynera
Your AI infrastructure control plane: a private model gateway, governed releases, RAG, GPU scheduling and cost control. Idle GPU capacity is money burning on a rack — and without tenant isolation, quota enforcement, and topology-aware scheduling, that's exactly what happens. Zynera's scheduler policies push utilization up across every tenant sharing the pool, while Zyra reasons over placement, cost, and security as one picture instead of a dozen disconnected tools.
What you get
Kubernetes-native GPU control
GPU pools, tenants, quotas, VMs with GPU live migration, and extended workload types as Kubernetes resources.
Private AI gateway
An OpenAI-compatible /v1 endpoint on your own GPUs: guardrails, routing, caching, weighted model aliases with canary and rollback, batch, speech-to-text, and image generation.
Evaluation-gated promotion & RAG
Promotion to production is blocked unless the vertical’s golden sets pass (admin override needs a written rationale). RAG Studio adds connectors, citations, and federated retrieval.
Zyra — embedded AI ops engineer
Approval-gated autonomous workflows, multi-agent specialists, and hybrid RAG grounded in your own cluster state. Actions stay advisory until a human approves.
Network intelligence
eBPF-powered flow visibility and security governance across the fleet.
FinOps & federation
Chargebacks, budget guardrails, and cost-preferring placement across clusters.
Built for production
Compute — GPU fabric after cluster bootstrap with HyperCluster.
Why teams choose us
- Production-ready out of the box — operators and resources to start from, not a blank slate
- Local AI platform — chat, a license-tagged open-weight model hub, per-key usage metering, and an OpenAI-compatible /v1 gateway on your own GPUs, zero external LLM key required
- Governed model release — promotion to production is gated on golden-set evaluation, with every decision recorded
- Pairs with IronWolf bare metal and Zeus OS VM GPU passthrough
- Self-hosted economics vs always-on public GPU clouds
When to choose
- You run AI/ML on owned Kubernetes with NVIDIA GPUs
- You want a private OpenAI-compatible endpoint with guardrails and evaluation-gated model releases, on hardware you own
- You need multi-tenant GPU quotas rather than best-effort sharing
Zynera questions
How it works
zynera gpu inspect NODE --rdmaMigrate from the platforms you already run
- Enterprise Hypervisors
- HCI Platforms
- AWS
- Azure
- Google Cloud
- Windows Hypervisor
- Oracle Cloud
- Machina fleet cloud
- KubeVirt
- KVM
- KVM-based Platforms
Architecture
One operator per layer. One control plane.
Every layer of your GPU infrastructure — scheduling, storage, networking, multi-tenant cost control — gets its own purpose-built Kubernetes operator, instead of being stitched together by hand.
GPU Operator
NVIDIA-native GPU scheduling with MIG partitions, time-slicing, topology awareness, and automatic driver management.
AI Operator
Elastic training and workload orchestration for distributed AI and ML pipelines.
Storage Operator
Lustre, GPFS, and BeeGFS integration to keep your GPUs fed with data.
Network Operator
InfiniBand and RoCE v2 support with automatic NVLink tuning for fast GPU-to-GPU traffic.
Quota Operator
Multi-tenant GPU quota management with cost tracking and chargeback across teams and projects.
Network Intelligence
See GPU traffic patterns and bandwidth use, and enforce network policy, with eBPF-powered observability.
Virtualization Operator
Run virtual machines alongside GPU containers on the same control plane.
Private AI cloud
Host your own AI cloud — LLMs on your GPUs, nothing phones home.
Zynera is a self-hosted inference platform. Serve open models on your own hardware, run them locally, and keep every token, prompt, and embedding inside your cluster.
Many serving backends, one spec
Pick the serving engine per model — Ollama, vLLM, TensorRT-LLM, Triton, SGLang, or TorchServe — from one spec. Scale to zero when idle and canary new versions safely.
Ollama, built in
Pull and serve Llama, Mistral, Qwen, and more, managed for you — no external API, no per-token bill.
Private routing by policy
Sensitive prompts are routed to a local model automatically, so they never leave the building even when cloud models are also configured.
Embeddings on your hardware
Chunk, embed, and index your documents for RAG entirely on your own hardware — no hosted embedding API.
Air-gapped Zyra
Zyra, the AI ops engineer, can run on a local model — it keeps working with zero internet access.
In the console
See it in the console.
Screens from a Zynera install: model catalog, Mission Control, and the live pod console.



AI platform
A full AI platform, not just a scheduler.
One dashboard for the whole AI lifecycle — build, ship, and govern models on the same control plane that schedules your GPUs.
Build
Author, fine-tune, and test models and agents without leaving the platform.
Operate
Promote a model to a served endpoint and watch it under load — no CLI required.
Govern
Traces, policy, chargeback, and multi-tenant workspaces in one governed surface.
The differentiator
Zyra — an AI ops engineer that works under approval.
Most infrastructure platforms automate infrastructure. None ship an embedded, approval-gated AI SRE that reasons about your fleet, proposes an action, shows its blast radius, and waits for a human to sign off.
Watch Zyra work through a GPU imbalance.
Zyra scans fabric telemetry, spots an imbalance, and proposes a topology-aware rebalance with its blast radius shown up front. Nothing executes until a human clicks Approve. This is a scripted illustration with invented numbers, not a live cluster.
Part of Zyra — see the same AI persona across the whole suite →
Open source companion
Test Zynera scheduling without physical GPUs
Zyvor Janus is the Apache 2.0 discrete-event simulator for Zynera — MIG, topology, quotas, and LLM serving metrics on your laptop.
Zynera + Zyvor Janus, side by side (~3 min)
Observe, learn, and act on GPU fleets
FinOps waste detection, placement bandit, training/inference optimizers, and NL cluster copilot with policy-gated remediation.
Cluster copilot
Natural-language queries and remediation suggestions with policy guardrails — closed-loop auto-tuning without blind automation.
FinOps & placement
Waste detection, cost rollups, and bandit-based placement recommendations tied to actual GPU utilization.
FabricAutoPolicy
Orchestrator tick reconciles extended CRDs into auto-remediation — federation APIs for multi-cluster AI factory operations.
CLI, API gateway, and dashboard
zynera CLI
Rust CLI for jobs, quotas, queues, GPU inspection, network, and security operations.
API gateway & SDKs
FastAPI gateway for UI and automation; Python and Go SDKs for custom controllers.
101-page web UI
Jobs, quotas, nodes, costs, network flows, policies, workspaces, model registry, inference, and auto-tuning.
Extended CRDs & eBPF
Kubernetes-native GPU fabric control, kernel-level eBPF visibility, and Grafana views for observability.
Signature deck
Download the h2kvm-format PDF or open the HTML preview for stakeholder reviews.
30-day free trial
Install in 30 seconds — no sign-up required.
Install Zynera with a single Helm chart. No account, no email, no sales call.
| Kubernetes | 1.28 and later |
| GPU nodes | NVIDIA GPU nodes |
| Architecture | x86-64 or ARM64 |
| Network fabric | InfiniBand / RoCE optional |
| Install | Single Helm chart |
| Sign-up | None — the 30-day trial starts automatically |
helm install zynera ./charts/zynera-core \ --namespace zynera-system \ --create-namespace \ --set namespace.create=false
Full docs, including the eight components of the AI Infrastructure OS: Read the Zynera docs →
Free trial
30-day free trial, no sign-up — Zynera
Run your own AI cloud on GPUs you already own — Kubernetes-native, no lock-in.
- One Helm install, GPU nodes detected automatically
- Trial clock starts from the build date — no account, no sales call
- Licence key from [email protected] after the trial
After the trial window, it is completely up to you whether to continue. There is absolutely no pressure or obligation from our side.
Zynera · GPU / AI infra
Ready to own your GPU compute?
Kubernetes-native orchestration for the GPUs you already own — fully self-hosted.