In the SaaS world, maintenance windows are a relic of the past. Customers expect your service to be always on, always available, and seamlessly updated. For years, Automic Automation has delivered on this promise through our zero downtime upgrade (ZDU) capability.
ZDU has long been a robust and trusted feature, allowing us to seamlessly upgrade the underlying engine and components. This enabled us to bridge the gap between Automic Automation versions, without interrupting active workflows and tasks.
However, updating systems in wildly different environments with diverse workloads changes the laws of probability. When a system processes millions of executions, demanding conditions and highly varying environmental influences are no longer theoretical edge cases; they are statistical inevitabilities.
To ensure absolute reliability at SaaS scale, we recently evolved our ZDU testing strategy. To explain how we did it, let's use a metaphor: Building enterprise automation software is a lot like building a Formula 1 race car.
Before we look at how we test the upgrade, it helps to understand the technical process happening behind the scenes. An Automic ZDU is a highly orchestrated process in which two versions of the Automation Engine (AE) run simultaneously against the same database.
In our V24 baseline, the technical sequence looked like this:
The database shift: The database schema is upgraded in a strictly backward-compatible way, allowing the current version to keep functioning normally.
Parallel processes: The new version's AE work processes (xWPs) and communication processes (xCPs) are spun up alongside the old version.
Seamless handoff: Agents and UI connections are instructed to disconnect from the old CPs and reconnect to the new ones.
The drain and finalize: The system waits for all message queues (MQs) and the MQMEM table to empty completely, which means that all workload has changed to the new version. Once the old processes are drained of active tasks and queues, they are safely shut down, and the upgrade is finalized.
It is an elegant architectural flow, but ensuring this works flawlessly across thousands of concurrent executions requires an uncompromising testing strategy.
From the beginning, our ZDU testing was highly automated, focusing on validating the orchestration of the upgrade itself.
The technical reality: We would automatically spin up a system, initiate the ZDU, and verify the transition from one version to the other. We ran these on default test systems with different combinations of databases and operating systems. This proved that the upgrade worked, database schemas were updated, and component routing shifted exactly as designed.
The metaphor: This is our wind tunnel. Inside a sterile, high-tech aerodynamic laboratory, our unbranded F1 car is locked onto a testing platform. With smooth smoke trails flowing over its curves, the engineers are validating that the underlying mechanics work perfectly in theory. But a wind tunnel doesn't win races; it just proves the car won't fall apart when the engine starts.
Knowing the core process was solid, our natural next step was to introduce active workload into the equation.
The technical reality: We extended our automation by adding dedicated test cases directly into the ZDU sequence. Now, while the upgrade was orchestrating, tests would actively run in parallel. This proved that standard workloads experienced zero disruption during the transition. However, because these tests were explicitly written for the ZDU pipeline, they were polite and predictable.
The metaphor: We took our F1 car out of the lab and put it on the perfect track. The sun is shining, the sky is blue, and the tarmac is pristine. The sleek F1 car speeds flawlessly from the v24 starting gantry to the v26 finish line. Everything is operating under optimal, stress-free conditions.
It was at this point we realized a fundamental paradox in enterprise software: What is best practice for executing an upgrade is the exact opposite of best practice for testing it.
If you are planning an upgrade in a production environment, you carefully schedule a window and try to minimize system load to create the safest possible conditions. But if you are an engineering team trying to make that upgrade truly bulletproof, relying on those perfect conditions during testing is a trap.
The reality check: A Formula 1 car is a masterpiece of engineering, but it comes with a major catch: It only performs on a flawlessly maintained, predictable circuit. What happens if you take that perfectly tuned, low-clearance F1 car and drop it onto a heavy-duty off-road trail? It gets stuck.
Perfect, predictable race tracks don't reflect how enterprise software is actually used. While our underlying SaaS infrastructure is rock-solid, the sheer complexity of our customers' operations creates an intensely demanding environment. Every tenant brings a wildly different mix to the track: massive daily workloads, thousands of diverse agents, intense API traffic, and complex data transports.
In our world, the complexity isn't coming from the platform—it is the extreme variables created by massive scale, high concurrency, and unique customer usage. If we only test our ZDU against polite, predictable workloads, we are building a car for a scenario that doesn't actually exist.
To catch sporadic, micro-timing issues like deep race conditions, we couldn't just orchestrate polite tests around an upgrade. We needed to embrace the mud.
This led to our real breakthrough: Instead of adding tests to our carefully controlled ZDU pipeline, we added a ZDU to our heaviest, most chaotic test pipeline.
The technical reality: Every night, our automated CI/CD pipelines execute thousands of complex tests, putting the system under massive, sustained load. Rather than running a separate ZDU test, we now randomly trigger a ZDU right in the middle of these heavy nightly test runs. The architecture shifts randomly right when the database is redlining and queues are packed to the brim.
This phase 3 approach completely changed the game. By triggering the upgrade entirely unannounced during our highest-load scenarios, we forced dormant race conditions out of hiding.
In fact, this "off-road" chaos testing was exactly how we fortified Automic V26.1. We forced a variety of active agents with different operating systems and versions to disconnect and reconnect to the new CPs simultaneously, while the system was redlining with nightly test data. In this way, our pipelines exposed microscopic race conditions in the handoff. Because we caught these transient errors in the mud, our engineers were able to completely rewrite and fortify the agent reconnect logic, making it virtually indestructible.
The V26.1 finalization (no more waiting on queues): While revisiting the ZDU process, we also realized that the technical checks required before finalization (like waiting for empty MQs and other tables like MQMEM) were overly cumbersome. In V26.1, we eliminated this by replacing those confusing, deep table checks with a simple, intuitive activity check.
Now, the system only verifies that old version activities are finished or switched over. You can easily track the status of this right in the AWI process monitoring view. Once clear, you can finalize the upgrade immediately, while the AE automatically handles the technical cleanup quietly in the background.
To ensure all of this works perfectly, we instituted an uncompromising quality gate: Not a single test is allowed to fail.
If even one test out of thousands drops or stalls due to the ZDU, the build is flagged. When a build survives this nightly gauntlet without dropping a single active job, we know with absolute certainty that our codebase is rock solid.
Evolving our ZDU testing strategy was about taking an already great feature and hardening it for the uncompromising demands of SaaS.
We threw out the playbook, moving out of the test lab and into real-world conditions with high-volume, unpredictable nightly chaos. By doing so, we have eliminated the risk of sporadic edge cases. Because we simplified the finalizing checks and bulletproofed the agent reconnects, rolling out Automic V26.1 in our SaaS environment is faster and more resilient than ever.
Importantly, this updated ZDU is also available for our on-premises environments. We invite our on-premises customers to take advantage of all the rigorous testing and optimization we have done. It is now simple to put these optimizations into practice in your own deployments.
Today, when we roll out an upgrade, it is a single code base across all deployment modes, and our customers don't notice a thing. Plus, our engineers sleep soundly knowing the release has already survived far worse.