<img height="1" width="1" style="display:none;" alt="" src="https://px.ads.linkedin.com/collect/?pid=1110556&amp;fmt=gif">
Skip to content
    October 2, 2026

    Going Off-Road: How We Bulletproofed Automic's Zero Downtime Upgrades

    9 min read

    Summary
    To ensure scalability and uninterrupted availability, Broadcom evolved the testing of Automic’s zero downtime upgrade capability. Teams perform unannounced, random testing under high-volume workloads rather than testing in controlled lab environments. This approach boosts resilience and speeds upgrade finalization, allowing customers to execute transparent updates across both SaaS and on-premises environments.
    Key Takeaways
    • Continuous workflow execution: Maintain uninterrupted operations during software upgrades through continuous workload testing.
    • Accelerated upgrade finalization: Leverage intuitive activity checks rather than table queues to finalize system upgrades faster.
    • Cross-environment confidence: Deploy resilient automation releases across both SaaS and on-premises environments.

    In the SaaS world, maintenance windows are a relic of the past. Customers expect your service to be always on, always available, and seamlessly updated. For years, Automic Automation has delivered on this promise through our zero downtime upgrade (ZDU) capability.

    ZDU has long been a robust and trusted feature, allowing us to seamlessly upgrade the underlying engine and components. This enabled us to bridge the gap between Automic Automation versions, without interrupting active workflows and tasks.

    However, updating systems in wildly different environments with diverse workloads changes the laws of probability. When a system processes millions of executions, demanding conditions and highly varying environmental influences are no longer theoretical edge cases; they are statistical inevitabilities.

    To ensure absolute reliability at SaaS scale, we recently evolved our ZDU testing strategy. To explain how we did it, let's use a metaphor: Building enterprise automation software is a lot like building a Formula 1 race car.

    Under the hood: The mechanics of an Automic V24 ZDU

    Before we look at how we test the upgrade, it helps to understand the technical process happening behind the scenes. An Automic ZDU is a highly orchestrated process in which two versions of the Automation Engine (AE) run simultaneously against the same database.

    In our V24 baseline, the technical sequence looked like this:

    • The database shift: The database schema is upgraded in a strictly backward-compatible way, allowing the current version to keep functioning normally.

    • Parallel processes: The new version's AE work processes (xWPs) and communication processes (xCPs) are spun up alongside the old version.

    • Seamless handoff: Agents and UI connections are instructed to disconnect from the old CPs and reconnect to the new ones.

    • The drain and finalize: The system waits for all message queues (MQs) and the MQMEM table to empty completely, which means that all workload has changed to the new version. Once the old processes are drained of active tasks and queues, they are safely shut down, and the upgrade is finalized.

    It is an elegant architectural flow, but ensuring this works flawlessly across thousands of concurrent executions requires an uncompromising testing strategy.

    Phase 1: The "pure ZDU" baseline (the wind tunnel)

    From the beginning, our ZDU testing was highly automated, focusing on validating the orchestration of the upgrade itself.

    • The technical reality: We would automatically spin up a system, initiate the ZDU, and verify the transition from one version to the other. We ran these on default test systems with different combinations of databases and operating systems. This proved that the upgrade worked, database schemas were updated, and component routing shifted exactly as designed.

    • The metaphor: This is our wind tunnel. Inside a sterile, high-tech aerodynamic laboratory, our unbranded F1 car is locked onto a testing platform. With smooth smoke trails flowing over its curves, the engineers are validating that the underlying mechanics work perfectly in theory. But a wind tunnel doesn't win races; it just proves the car won't fall apart when the engine starts.

    Phase 2: ZDU with dedicated tests (the perfect track)

    Knowing the core process was solid, our natural next step was to introduce active workload into the equation.

    • The technical reality: We extended our automation by adding dedicated test cases directly into the ZDU sequence. Now, while the upgrade was orchestrating, tests would actively run in parallel. This proved that standard workloads experienced zero disruption during the transition. However, because these tests were explicitly written for the ZDU pipeline, they were polite and predictable.

    • The metaphor: We took our F1 car out of the lab and put it on the perfect track. The sun is shining, the sky is blue, and the tarmac is pristine. The sleek F1 car speeds flawlessly from the v24 starting gantry to the v26 finish line. Everything is operating under optimal, stress-free conditions.

    The reality check: Perfect tracks don't exist in the wild

    It was at this point we realized a fundamental paradox in enterprise software: What is best practice for executing an upgrade is the exact opposite of best practice for testing it.

    If you are planning an upgrade in a production environment, you carefully schedule a window and try to minimize system load to create the safest possible conditions. But if you are an engineering team trying to make that upgrade truly bulletproof, relying on those perfect conditions during testing is a trap.

    The reality check: A Formula 1 car is a masterpiece of engineering, but it comes with a major catch: It only performs on a flawlessly maintained, predictable circuit. What happens if you take that perfectly tuned, low-clearance F1 car and drop it onto a heavy-duty off-road trail? It gets stuck.

    Perfect, predictable race tracks don't reflect how enterprise software is actually used. While our underlying SaaS infrastructure is rock-solid, the sheer complexity of our customers' operations creates an intensely demanding environment. Every tenant brings a wildly different mix to the track: massive daily workloads, thousands of diverse agents, intense API traffic, and complex data transports.

    In our world, the complexity isn't coming from the platform—it is the extreme variables created by massive scale, high concurrency, and unique customer usage. If we only test our ZDU against polite, predictable workloads, we are building a car for a scenario that doesn't actually exist.

    Phase 3: The breakthrough (the off-road evolution)

    To catch sporadic, micro-timing issues like deep race conditions, we couldn't just orchestrate polite tests around an upgrade. We needed to embrace the mud.

    This led to our real breakthrough: Instead of adding tests to our carefully controlled ZDU pipeline, we added a ZDU to our heaviest, most chaotic test pipeline.

    • The technical reality: Every night, our automated CI/CD pipelines execute thousands of complex tests, putting the system under massive, sustained load. Rather than running a separate ZDU test, we now randomly trigger a ZDU right in the middle of these heavy nightly test runs. The architecture shifts randomly right when the database is redlining and queues are packed to the brim.

    • The metaphor: We didn't just change the tires; we made the vehicle thoroughly rugged. Our F1 car evolved into an off-road beast. The vehicle is equipped with lifted suspension, deep-tread tires, and rally lights. Now, when faced with the exact same flooded, muddy forest track, the armored car powerfully drifts through the mud, roaring victoriously toward the v26 finish line, despite unannounced obstacles and extreme environmental stress.

    The ultimate quality gate: Squashing bugs and evolving to V26.1

    This phase 3 approach completely changed the game. By triggering the upgrade entirely unannounced during our highest-load scenarios, we forced dormant race conditions out of hiding.

    In fact, this "off-road" chaos testing was exactly how we fortified Automic V26.1. We forced a variety of active agents with different operating systems and versions to disconnect and reconnect to the new CPs simultaneously, while the system was redlining with nightly test data. In this way, our pipelines exposed microscopic race conditions in the handoff. Because we caught these transient errors in the mud, our engineers were able to completely rewrite and fortify the agent reconnect logic, making it virtually indestructible.

    The V26.1 finalization (no more waiting on queues): While revisiting the ZDU process, we also realized that the technical checks required before finalization (like waiting for empty MQs and other tables like MQMEM) were overly cumbersome. In V26.1, we eliminated this by replacing those confusing, deep table checks with a simple, intuitive activity check.

    Now, the system only verifies that old version activities are finished or switched over. You can easily track the status of this right in the AWI process monitoring view. Once clear, you can finalize the upgrade immediately, while the AE automatically handles the technical cleanup quietly in the background.

    To ensure all of this works perfectly, we instituted an uncompromising quality gate: Not a single test is allowed to fail.

    If even one test out of thousands drops or stalls due to the ZDU, the build is flagged. When a build survives this nightly gauntlet without dropping a single active job, we know with absolute certainty that our codebase is rock solid.

    The result: Invisible upgrades, absolute reliability

    Evolving our ZDU testing strategy was about taking an already great feature and hardening it for the uncompromising demands of SaaS.

    We threw out the playbook, moving out of the test lab and into real-world conditions with high-volume, unpredictable nightly chaos. By doing so, we have eliminated the risk of sporadic edge cases. Because we simplified the finalizing checks and bulletproofed the agent reconnects, rolling out Automic V26.1 in our SaaS environment is faster and more resilient than ever.

    Importantly, this updated ZDU is also available for our on-premises environments. We invite our on-premises customers to take advantage of all the rigorous testing and optimization we have done. It is now simple to put these optimizations into practice in your own deployments.

    Today, when we roll out an upgrade, it is a single code base across all deployment modes, and our customers don't notice a thing. Plus, our engineers sleep soundly knowing the release has already survived far worse.

    Andreas Ronge

    Andreas Ronge is the Chief Architect for Automic Automation. He has been with Broadcom for 20+ years in different roles, starting as a developer and is now the technical lead for the product responsible for all new feature designs as well as the technical roadmap.

    Other resources you might be interested in

    icon
    Video August 28, 2026

    Automic Automation Cloud Integration: Workday

    This video explains the Automic Automation Workday agent integration and its benefits. Find out how to install, configure, and use the agent.

    icon
    Blog August 26, 2026

    Your Orientation Guide to the automic.com Migration

    All automic.com sites have been migrated to Broadcom platforms. Review this post to find out how to access downloads, Docker images, and documentation.

    icon
    Product Education July 10, 2026

    Automic Integration Brochure

    This brochure serves as your guide to the diverse tools, platforms, and systems that can connect seamlessly with Automic.

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v26 in Azure

    Master the deployment of Automic Automation v26 on Azure AKS. Cover database setup, TLS certificates, and the new Kubernetes Gateway API.

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v24 in Azure

    Deploy Automic Automation Kubernetes Edition v24 on Azure AKS with this step-by-step installation and configuration guide.

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v26 in AWS

    Learn how to deploy Automic Automation Kubernetes Edition v26 on AWS EKS with this step-by-step guide for configuring databases, secrets, and agents.

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v24 in AWS

    See how to deploy Automic Automation v24 on AWS EKS. Learn about using Fargate, Helm charts, PostgreSQL, and AWS Load Balancer Controller.

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v24 in GCP

    This guide walks you through the steps to deploy Automic Automation Kubernetes Edition v24 into Google Kubernetes Engine (GKE) on the Google Cloud Platform (GCP).

    icon
    White Paper June 5, 2026

    How to Install Automic Automation Kubernetes Edition v26 in GCP

    Discover the steps needed to deploy Automic Automation Kubernetes Edition v26 into Google Kubernetes Engine (GKE) on the Google Cloud Platform (GCP).