01 · Cloud Computing
Clusters that survive the failure of a node, a disk or a link — on Proxmox, Ceph, ZFS, VMware and Hyper-V, and Kubernetes with autoscaling for containerized applications. Designed, built, tested and maintained 24/7/365.
Service details
High Availability is not one product — it is an architecture in which no single failure stops the service. A node dies and the virtual machines start elsewhere. A disk fails and the data is already on other drives. A link goes down and traffic takes the second path. Users notice a few seconds, not a few hours.
We design, build and maintain such environments end to end: from choosing the platform and hardware, through clustering compute, storage and network, to testing the failover and running it 24/7/365 under an SLA.
Availability levels
Every additional nine costs real money. We start by agreeing what downtime actually costs your business, and then design the cheapest architecture that meets that target — not the most impressive one.
| SLA | Downtime per year | Typical architecture |
|---|---|---|
| 99.9% | ≈ 8 h 45 min | Single server with redundant components, backup with a tested restore |
| 99.95% | ≈ 4 h 23 min | Two-node cluster with shared or replicated storage, redundant links |
| 99.99% | ≈ 52 min | Three-node cluster with Ceph, redundant network and power paths, monitoring 24/7 |
| 99.99% + DR | ≈ 52 min, site loss covered | Cluster plus a second location with replication and a tested recovery plan |
RTO and RPO matter as much as the percentage: how fast the service must be back and how much data you can afford to lose. Both are agreed before the first diagram is drawn.
Architecture
Cluster of at least three nodes with live migration and automatic restart of workloads on a healthy node.
Ceph with replication or erasure coding, or a redundant array with multiple controllers — no single copy of your data.
Redundant switches and uplinks, bonded interfaces, separate storage and management networks — see network engineering.
Replication with automatic or controlled failover, tested under load so the promotion of a replica actually works.
Several instances behind a load balancer, health checks and rolling updates instead of maintenance windows.
HA is not backup: we pair the cluster with backup and DR, including immutable copies in a second location.
Kubernetes
For applications built from containers, Kubernetes gives you two things at once: availability that survives the loss of a node, and elasticity that adds capacity exactly when traffic demands it. We build and operate clusters on your hardware, in our data center or in the public cloud.
Three control-plane nodes with a quorum of etcd, so losing one does not stop deployments or the cluster API.
Failed pods are restarted, nodes that die have their workloads rescheduled elsewhere, health probes keep traffic away from unhealthy instances.
More pods when CPU, memory or your own metrics (queue length, requests per second) grow — and fewer when the peak passes.
New worker nodes added automatically when pods cannot be scheduled, removed when they are no longer needed — capacity follows the load.
Rolling updates, readiness gates and pod disruption budgets, so releases and maintenance happen during the day without an outage.
Anti-affinity rules and topology constraints keep replicas on different nodes, racks or locations instead of all in one place.
Ceph through a CSI driver for volumes that survive a pod moving to another node, with snapshots and replication.
Redundant ingress controllers, TLS certificates managed automatically, traffic distributed across healthy instances.
Cluster upgrades, monitoring, backups of etcd and workloads, capacity planning and on-call — see DevOps services.
| Kubernetes | VM cluster (Proxmox / VMware / Hyper-V) | |
|---|---|---|
| Best for | Containerized applications, microservices, APIs, CI/CD-driven teams | Business systems, databases, legacy applications, mixed environments |
| Failover | Seconds — a new pod starts on another node | From seconds to minutes — the VM restarts on another host |
| Scaling | Automatic, in both directions, down to a single pod | Manual or scheduled, by adding resources or machines |
| Deployments | Rolling updates with no downtime | Usually a maintenance window |
| Requirements | Application must be container-ready and stateless where possible | Works with what you already run, no application changes |
| Operations | More moving parts — worth outsourcing if you have no platform team | Simpler day-to-day administration |
In practice the two live together: Kubernetes for the applications, a VM cluster for the databases and systems that are not going to be containerized — on the same storage and network foundation.
Why it matters
Today's businesses cannot afford downtime. Every minute of system unavailability means real financial, reputational and operational losses. That is why we deploy High Availability (HA) solutions that guarantee uninterrupted operation of key IT services — even in the event of hardware or software failure. Why High Availability from Remote Admin? We specialize in designing and maintaining highly available IT environments. For years we have delivered end-to-end deployments for enterprises with elevated availability requirements.
Technology tailored to your needs We build HA architectures based on both licensed products and open-source solutions — optimizing the budget without sacrificing quality or reliability. We work with proven technologies and respected platforms: • Fortinet • Cisco • Dell • VMware • Proxmox • ZFS • Ceph
Operations
A 30-minute call with an engineer — we'll outline the scope and ballpark budget, with no sales pitch.
Storage
Ceph is an open-source, highly scalable, distributed enterprise-class storage system designed for high availability (HA), fault tolerance and automatic data rebalancing. It can deliver block, file and object storage simultaneously within a single platform. The system is built on RADOS technology, which provides automatic data replication, self-healing and load distribution across cluster nodes.
Ceph was designed for virtually unlimited growth:
Despite its enormous capabilities, Ceph is not always optimal. It is not recommended for:
Ceph requires suitable network infrastructure (ideally 10/25/40GbE), and its efficiency increases with scale.
Storage
ZFS (Zettabyte File System) is a modern, extremely fault-tolerant file system combined with a storage management layer. Originally created by Sun Microsystems and now developed as an open-source project, it is one of the most respected technologies in environments where data integrity and high performance are key.
As the name suggests, the ZFS architecture is designed to handle zettabyte-scale data volumes.
Scaling is done by adding disks or additional vdevs to the pool.
It excels wherever the top priorities are:
Most commonly found in:
It is not always a one-size-fits-all solution:
Platforms
Proxmox Virtual Environment (Proxmox VE) is an advanced open-source platform for server virtualization, containerization and management of highly available IT clusters. It combines features known from enterprise-class solutions (VMware, Hyper-V) without licensing costs and with full flexibility to adapt.
Proxmox integrates in a single environment:
Administration is done through an intuitive web panel and API.
Proxmox performs well in both small deployments and large clusters:
It can work with both local storage (ZFS, RAID) and distributed Ceph for larger deployments.
We see the greatest benefits in:
Platforms
Hyper-V is a type 1 hypervisor developed by Microsoft, designed for server virtualization and building highly available IT environments based on Windows Server and Microsoft Azure. It enables centralized management of compute, storage and network resources — both on-premises and in a hybrid model.
Hyper-V is an integral part of the Microsoft ecosystem, which ensures compatibility and full support for Windows environments and many business applications.
Hyper-V handles environments of all sizes:
Performance and capacity grow as the cluster expands — with no downtime for critical services.
Leveraging Windows Server Standard/Datacenter licenses helps optimize costs compared to third-party hypervisors.
Especially recommended for:
Many organizations choose Hyper-V to ensure business continuity for mission-critical systems.
24/7/365
Book a call with our advisor. We'll select services tailored strictly to your needs, with no artificial costs.
Questions and answers
An architecture in which the failure of a single element — a server, a disk, a switch or a power supply — does not stop the service. Workloads move automatically to healthy components, and users experience seconds of disruption instead of hours.
No, and one does not replace the other. HA protects against hardware and software failures in the present; backup protects against deletion, corruption and ransomware, and lets you go back in time. A serious environment needs both.
Two nodes give you basic redundancy, but three are the practical minimum for a cluster that can decide which side is healthy (quorum) — which is why most of our HA deployments start at three nodes plus redundant network and power.
Yes, with the right architecture: a three-node cluster, distributed storage such as Ceph, redundant network paths and 24/7 monitoring. Whether it is worth it depends on what an hour of downtime costs you — we do that calculation together before designing anything.
On two levels. The Horizontal Pod Autoscaler adds and removes application instances based on CPU, memory or your own metrics such as queue length. The Cluster Autoscaler adds and removes worker nodes when there is nowhere to schedule new pods. Together they keep performance stable during peaks and the bill low outside them.
No. Kubernetes is excellent for containerized applications, but a Proxmox, VMware or Hyper-V cluster delivers high availability for systems that are not containerized — including most ERP systems and databases. We often run both side by side.
Yes — with replication and controlled or automatic failover, tested under load. For transactional systems we usually keep databases on fast local NVMe with replication rather than on distributed storage, because latency matters more than elasticity there.
Always, before acceptance and periodically afterwards: we pull the plug on a node, break a link and simulate a disk failure, then document how the environment behaved and how long recovery took. An untested cluster is a promise, not a guarantee.
That needs a second site: replication of data, a plan for DNS and traffic, and regular disaster recovery tests. We design it together with backup and DR, because the fastest cluster in the world does not help if the building is gone.
Both. We build on Proxmox, Ceph and ZFS as well as VMware and Hyper-V, and we choose based on your applications, your team's skills and the licence costs — not on what we prefer to install.
You, with the documentation and training we hand over, or our team under a 24/7/365 SLA — monitoring, updates, capacity planning and incident response. See server administration.
Related services
First step
30 minutes, no slide deck. We'll tell you straight whether this service solves your problem, what scope makes sense and how much it costs.
A proposal with scope and pricing within 48 hours of the call.
Go to contact