Architecture of Survival

Technological Sovereignty and ERP Continuity in Times of Full-Scale Crisis

October 7, 2023, marked a definitive point of no return for Israel a moment when the concept of Disaster Recovery (DR) shifted from theoretical modeling to the realm of national economic survival. For a systems engineer, this was a moment of truth. As the horizon filled with smoke and sirens wailed for hours, the operational objective remained immutable: over 1,500 enterprise users within the Priority ERP ecosystem could not be allowed to feel that the world around them was fracturing. At MedaTech System Ltd, we understood the gravity of the stakes. A failure of the ERP backbone at that hour would have resulted in the immediate paralysis of medical supply chains, food distribution networks, and defense logistics. We were managing a digital fortress where every byte of data had to reach its destination, notwithstanding direct physical threats to the underlying infrastructure.

The primary challenge of that “Black October” was absolute unpredictability. Direct kinetic strikes on communication hubs and rolling blackouts in frontline sectors threatened to instantaneously destabilize the nation’s logistics. My engineering mission was to transform systemic fragility into architectural invulnerability, designing an IT landscape capable of “breathing” and reconfiguring in real-time as primary backbone links ceased to exist.

The technical response to this chaos was total decentralization and the final abandonment of the single-command-center paradigm. We finalized the elimination of reliance on local on-premises nodes in favor of a Multi-Region Cluster a distributed network of data centers geographically dispersed to preclude data loss from any single catastrophic event. However, building a cloud environment is merely the foundation. The true art of engineering in a conflict zone is ensuring persistent connectivity when provider infrastructures are under constant kinetic pressure. We implemented a hybrid data transmission topology where traditional fiber backbones were reinforced by autonomous microwave radio-relay bridges and satellite terminals. Leveraging dynamic routing protocols (BGP) with extremely aggressive convergence timers (Keepalive/Hold-time) allowed for automated traffic rerouting within fractions of a second. During those critical days, there were instances where backbone nodes failed, yet the system instantly pivoted data flows to redundant paths, preserving active user sessions and transactional integrity. This represents the pinnacle of High Availability where the mathematical logic of protocols triumphs over physical destruction.

Special emphasis was placed on a data storage strategy rooted in absolute resilience. In an active combat environment, a standard daily backup is an act of professional negligence. We implemented a synchronous replication scheme for mission-critical SQL databases across geographically distant sites. Key technical parameters were pushed to the edge, our Recovery Point Objective (RPO) was locked at zero. Through synchronous transaction commits across multiple sites, data loss became physically impossible, even in the event of the instantaneous destruction of a primary data center. The Recovery Time Objective (RTO) was reduced to under 30 seconds, with automated failover ensuring full system restoration without manual engineering intervention. Every Priority ERP operation whether it was the dispatch of tactical gear or a life-saving medical order was committed simultaneously to two remote storage facilities. This created a “digital mirror” effect, where data possessed higher survivability than the physical media upon which it was written.

This experience taught me the most vital lesso, technology remains soulless until it is infused with an engineer’s will and foresight. We deployed monitoring systems that analyzed not only CPU load and disk array health but also the stability of external backbones, predicting signal degradation through indirect indicators (Jitter, Packet Loss) before they reached a fatal threshold. We identified failures at provider nodes before the providers themselves could issue notifications. Every time the system continued to operate normally following an incident, I knew we were doing more than maintaining servers we were ensuring the viability of the nation’s economy under existential threat. Today, looking at service availability reports at 99.9%, I see more than just statistics. Behind every decimal point lies the stability of supply chains and the confidence of hundreds of companies protected by technological armor. My experience is the knowledge of how to build systems that do not break under extreme pressure and how to remain an engineer when reality demands you be an architect of survival.

January 2024
Dmitry Bogoliubov