It is possible to install two UPS systems and still have one point whose failure removes both. It is also possible to have N+1 power modules but no safe way to service the frame without exposing the load. Redundancy therefore needs to be analysed as a path through sources, switchgear, UPS modules, batteries, bypass, distribution and the connected equipment. Static bypass and maintenance bypass deserve particular attention because they change the source and topology during overload, fault or service. Terms such as N+1 and 2N are valuable shorthand, but only after the design team agrees what “N” includes and which common components sit outside the label. The most reliable way to expose weakness is to draw every credible operating state and ask what remains if each important component is unavailable.

Define N before claiming redundancy

N is the capacity or number of components required to perform the intended function. In a modular UPS, four modules may be needed to carry the design load; a fifth creates N+1 at the module level. But the frame, control system, static bypass or battery may still be common. A complete resilience claim should therefore state the boundary: “N+1 power modules within one UPS frame” is more precise than simply “N+1 UPS”.

For 2N, two independent systems each have full required capacity. Independence should be tested across upstream and downstream infrastructure. If two UPS units share one input switchboard, one bypass transformer or one downstream distribution point, the design may not deliver the expected fault isolation. Draw the boundary around the service being protected and identify every common component.

What the static bypass does

In a double-conversion UPS, the inverter normally carries the load. The static bypass is an electronic transfer path that can connect the load to an alternative AC source when the inverter cannot support it, for example during certain overloads or faults. This can preserve continuity, but it means the bypass source, synchronisation window and protection behaviour are part of the critical design.

If the bypass source is unavailable or outside acceptable limits, the UPS may be unable to transfer. If the downstream fault current required to operate a protective device exceeds what the inverter can supply, the system may use bypass to help clear the fault, depending on design. These details influence selectivity and resilience. Request manufacturer data and coordinate protection for both inverter and bypass operating states.

What maintenance bypass adds

A maintenance bypass provides a manually controlled path that allows the UPS electronics to be isolated for service while the load remains energised. It may be integrated or external. An external wrap-around bypass can allow a complete UPS to be removed electrically, which can be valuable for replacement and major service.

The bypass route must be correctly rated, interlocked and labelled. Switching should follow a controlled procedure because an incorrect sequence can interrupt the load or create an unsafe parallel condition. From a resilience perspective, ask what is lost while on maintenance bypass. The load may remain energised but no longer have battery protection or power conditioning. In a critical site, maintenance windows should be treated as temporary changes in risk rather than ordinary operating conditions.

Common-mode failures

Redundant systems can fail together when they share an environment, control dependency or human process. Examples include both UPS paths in one room losing cooling, both battery strings exposed to the same high temperature, both systems updated with faulty firmware, or both feeds switched incorrectly during a procedure. Common-mode risk is broader than the one-line diagram.

Physical separation, independent controls and diverse sources can reduce some risks, but every additional layer adds complexity. The goal is not infinite duplication. It is to identify plausible common causes with consequences that justify mitigation. Include fire compartments, flooding, HVAC, monitoring networks, earthing, maintenance access and operator actions in the review. A simple architecture that staff understand can outperform a complex redundant design operated inconsistently.

Maintenance states reveal weak designs

A system may meet its resilience objective in normal operation but not while a component is isolated. If maintenance is frequent, the facility may spend significant time in a reduced-resilience state. Model planned service conditions during design. Can one UPS be maintained while the other carries the full load? Can a battery string be tested without losing required autonomy? Can bypass switchgear itself be inspected?

Where concurrent maintainability is required, every component that needs routine service should have a strategy for isolation. That includes upstream and downstream breakers, not just the UPS power modules. If a component cannot be maintained without interrupting the load, document that limitation and schedule an appropriate shutdown rather than pretending the system is continuously maintainable.

A/B distribution and single-cord loads

Dual UPS paths are most effective when the connected equipment can accept two independent supplies. Data-centre servers often have dual power supplies, allowing A and B feeds. Single-cord devices create a problem because a transfer switch or other point of convergence is needed, and that device can become a single point of failure.

Inventory single-cord loads and decide how they are protected. Static transfer switches can provide rapid transfer between sources, but their failure modes, maintenance and fault-current behaviour must be understood. For non-IT environments, duplicated power supplies may exist at the control-system level rather than the appliance. The resilience model should follow the load all the way to the function being protected.

Operational procedures are part of redundancy

Redundant hardware does not prevent an operator from opening the wrong breaker. Clear labels, mimic diagrams, keyed interlocks, switching schedules and peer checks can reduce human-error exposure. Procedures should state preconditions: source availability, synchronisation, load level, alarm state and the expected result of each step.

Train staff on abnormal states as well as normal switching. During a real failure, the system may already be partly degraded when a manual decision is required. Event logs and monitoring should make the current topology visible. After any significant switching event, verify that the system has actually returned to the intended normal configuration; temporary bypass states can otherwise persist unnoticed.

How to perform a single-point-of-failure review

Take the one-line diagram and choose one component at a time: transformer, breaker, UPS module, frame, battery string, bypass source, PDU, transfer switch, cooling unit or control network. Assume it is unavailable. Does the protected function continue? At what load? For how long? Can the failed item be isolated safely? Repeat the exercise for planned maintenance states and credible common causes.

Record findings in a simple table with component, failure effect, remaining capacity, detection method and mitigation. Prioritise weaknesses that can interrupt a high-consequence service without warning. The review should be repeated after major expansion because new distribution and temporary connections often introduce convergence points that were not present in the original architecture.

Primary references and further reading

Standards and official guidance may be amended. Confirm the edition and project-specific requirements with a competent professional before design, procurement or maintenance work.