MaximoWorld: Where Maximo users unlock more of their Maximo investment.

Join the leaders shaping the future of reliability at IMC

Sign Up

Please use your business email address if applicable

Why a Single Safety Controller Failure Triggered a Gas Turbine Trip in a 2oo3 TMR Architecture

Why a Single Safety Controller Failure Triggered a Gas Turbine Trip in a 2oo3 TMR Architecture

Triple Modular Redundancy (TMR) architectures with 2-out-of-3 (2oo3) voting are designed to tolerate a single hardware fault without interrupting processes. However, an unexpected Gas Turbine Generator (GTG) trip occurred when only one of three safety controllers failed. This case study analyzes the event, demonstrating the critical distinction between fault tolerance and fail-safe behavior when safety integrity is compromised.

Introduction

Engineers often assume a 2oo3 voting architecture guarantees continued operation during a single controller failure. This holds true only if voting integrity, controller synchronization, and deterministic communications remain healthy. When these conditions are violated, the Safety Instrumented System (SIS) intentionally prioritizes safety over process availability.

System Overview

The GTG was protected by an IEC 61508-compliant SIS utilizing a TMR architecture (Controllers R, S, and T) with 2oo3 voting logic and Hardware Fault Tolerance (HFT) = 1. Under normal conditions, a single failure should not cause a trip. The conceptual architecture and the diagnostic boundaries governing this behavior are illustrated in Figure 1.

reliabilityweb.com



Figure 1. Conceptual TMR Architecture and Design Intent

Event Description

At 21:56:55, the GTG unexpectedly tripped. Sequence-of-Events (SOE) logs showed simultaneous activation of multiple shutdown signals including Fire Protection, Gas Detection, and Emergency Push Buttons. Process data revealed no actual emergencies, shifting the investigation focus entirely to the safety system's automated response.

Fault Propagation and Investigation Findings

The investigation traced a progressive degradation of Controller R over a 10-minute window before the trip. The exact progression timeline is mapped out in Figure 2.



Figure 2. Event Timeline from Degradation to Trip.

The systematic engineering breakdown of how these sequential alarms escalated into a protective shutdown is detailed across seven distinct phases in Figure 3.

Figure 3. Fault Propagation Logic and Phase Breakdown.

Phases 1 & 2 (Early Degradation & Inconsistency): Repeated diagnostic alarms—Application Overrunning Frame and Frame Sync Monitor—indicated that Controller R lost deterministic execution. It quickly lost synchronization with Controllers S and T, causing Logic Signal Voting Mismatches.

Phase 3 & 4 (I/O Impact & Fault Classification): Safety I/O pack timeouts occurred. Under IEC 61508-2:2010, Clause 7.4.8.1, the system classified this total loss of communication and deterministic execution as an unrecoverable dangerous detected fault.

Phases 5, 6 & 7 (Fail-Safe Action & Trip): Following the OEM design philosophy, Controller R executed its predefined fail-safe response. Its outputs to the I/O packs were disabled, forcing all safety channels to safe values. The simultaneous fire, gas, and ESD signals were consequences of this fail-safe execution, not independent process causes.

Root Cause Analysis

The investigation eliminated external field triggers and hardware fault tolerance failures. As show in the cause map (Figure 4), the definitive cause chain traces back to one specific source

Figure 4. Root Cause Cause Map / Fault Tree

Root Cause Statement: Internal failure of Safety Controller R progressively degraded deterministic execution and controller synchronization, preventing reliable participation in the TMR architecture and causing the Safety Instrumented System (SIS) to execute its designed fail-safe response, resulting in an automatic protective shutdown of the gas turbine.

Functional Safety & Design Verification

The system's behavior perfectly aligns with GE Vernova’s Mark VIeS Control Systems System Guide (GEH-6721, Section 3.2.7.2), which dictates that upon detecting critical internal faults, the controller must disable outputs to the I/O packs to force a safe state. Furthermore, this response satisfies IEC 61508-2:2010, Clause 7.4.8.1. The trip was an intended protective action, not a spurious malfunction.

Engineering Lessons

  • Fault Tolerance Has Boundaries: A 2oo3 TMR architecture delivers fault tolerance only while synchronization and deterministic execution remain healthy. Once integrity cannot be verified, safety overrides availability.
  • Diagnostics Are Leading Indicators: Framework synchronization and overrunning frame alarms must be treated as critical leading indicators, not nuisance alarms. Prompt intervention during these early stages may prevent a full plant shutdown.
  • Availability Is Secondary to Safety: In functional safety, protecting people, assets, and the environment always takes precedence over maximizing production uptime.
  • Reliability and Functional Safety Must Cooperate: Reliability engineers must understand functional safety philosophies to avoid misclassifying designed protective actions as system failures.
  • IEC 61508-2:2010, Functional Safety of Electrical/Electronic/Programmable Electronic Safety-Related Systems — Part 2, Clause 7.4.8.1.
  • GE Vernova, Mark VIe and Mark VIeS Control Systems, Volume I: System Guide (GEH-6721), Section 3.2.7.2.

Conclusion

This event demonstrates that a single controller failure can trigger a process trip even within a 2oo3 TMR architecture. The internal failure of Controller R compromised safety integrity, triggering a validated fail-safe response. The shutdown represents the successful execution of a system designed to prioritize protection over production.

References
1.IEC 61508-2:2010, Functional Safety of Electrical/Electronic/Programmable Electronic Safety-Related Systems — Part 2, Clause 7.4.8.1.
2.GE Vernova, Mark VIe and Mark VIeS Control Systems, Volume I: System Guide (GEH-6721), Section 3.2.7.2.

Nguyen Khac Sau

Nguyen Khac Sau, CMRP, is a Reliability Engineer with over 14 years of experience in refinery, petrochemical, and power generation industries. His expertise includes instrumentation reliability, functional safety, root cause analysis (RCA), reliability-centered maintenance (RCM), and maintenance strategy optimization. He is passionate about sharing practical engineering case studies that help improve asset reliability, process safety, and operational excellence.

You can ask anything about maintenance, reliability, and asset management.