LIVE Is Not Your Test Environment

Michael Jones · Oct 8, 2026 · 9 min read

LIVE Is Not Your Test Environment

All posts

Your customers are not test scripts. Functional success does not prove that a live service is operationally ready.

Customers do not experience software as a collection of passed tests. They experience a service.

They do not care that User Acceptance Testing was completed, that every functional requirement passed, or that the deployment pipeline reported success. They care whether the service responds when they need it, whether their data remains safe, whether transactions complete, and whether somebody can restore the service when something goes wrong.

The uncomfortable reality is that a system can be functionally correct while remaining operationally unsafe.

It can pass its tests and still fail under load. It can return the right answer and still be impossible to recover. It can satisfy a requirement and still expose the business to unacceptable risk.

The software can work fine, while the service does not.

That distinction explains why Operational Acceptance Testing matters and why so many organisations overlook it until they experience the consequences firsthand.

We test the software. Customers use the service.

Most organisations naturally focus on functional behaviour. Functional testing answers visible questions:

  • Can the customer log in?
  • Can they complete a transaction?
  • Does the new feature produce the correct result?
  • Does the user journey meet the requirement?

These tests are relatively easy to connect to a roadmap, customer request or contractual commitment.

Operational Acceptance Testing asks different questions:

  • What happens when transaction volume triples?
  • Can the service recover after a database failure?
  • Will alerts identify the real customer impact?
  • Can support teams diagnose the problem?
  • Can a failed deployment be rolled back safely?
  • Will data remain consistent during recovery?
  • What happens when a supplier, network or shared service becomes unavailable?

These questions do not demonstrate a new feature. They demonstrate whether the organisation can safely operate that feature after release.

That distinction matters because a production service includes far more than application code. It includes infrastructure, databases, networks, cloud configuration, monitoring, alerting, deployment mechanisms, suppliers, operational procedures, support teams and recovery capabilities.

A component can pass its functional tests while the overall service remains unsafe to operate.

Effective OAT therefore tests the complete service, including software, platforms, suppliers, people, processes, monitoring and support.

Why functional testing gets prioritised

Functional testing produces visible progress. Features can be demonstrated. Delivery milestones can be measured. Requirements can be signed off.

Operational readiness is less visible.

Product teams are usually measured against functionality, milestones and release dates. OAT can expose information that challenges those commitments.

A performance test may show that projected demand cannot be supported. A recovery exercise may prove that documented recovery objectives are unrealistic. A deployment rehearsal may reveal that rollback is unsafe.

Good OAT can produce answers that organisations may not want to hear.

Unless operational evidence is allowed to influence release decisions, investment priorities and delivery plans, OAT becomes an inconvenience rather than a meaningful control.

Many organisations say they want greater reliability while continuing to prioritise functionality and dates. Risks are recorded, but do not consistently change investment, sequencing, testing or release decisions.

Functional tests passing while the live service shows degradation, queue growth and dependency timeouts
Passed functional tests do not tell you how the service behaves under real operational conditions.

OAT is often introduced too late

OAT is sometimes treated as a final checklist immediately before release. By that point:

  • The architecture has already been selected.
  • The environment has already been built.
  • Supplier arrangements have already been committed.
  • Non-functional requirements may be missing or poorly defined.
  • There is insufficient time to address significant findings.
  • The release date has acquired commercial or management importance.

This turns OAT into a late approval gate rather than an operational readiness activity.

A more effective approach begins earlier by identifying operational risks, critical customer journeys, demand assumptions, failure scenarios, required evidence and accountable owners.

Operational readiness should be treated as a service-level, evidence-led activity, not a final-stage performance test or a mechanism for transferring engineering risk to a test team.

Representative environments are difficult and expensive

A meaningful OAT environment does not need to be an identical copy of production, but it must reproduce the behaviour and constraints relevant to the risks being tested.

Companies frequently discover that:

  • Production configuration is not fully documented.
  • Lower environments behave differently.
  • Some changes have been made manually.
  • Real customer traffic patterns are not understood.
  • Production data volumes cannot be reproduced.
  • External dependencies cannot be exercised safely.
  • The service cannot be rebuilt consistently.

When an organisation tries to build another environment, it may discover that nobody can explain which configuration is authoritative or why production behaves differently. At that point, what looked like an infrastructure exercise becomes an investigation into how the service actually works.

Why start-ups are particularly vulnerable

Start-ups often have entirely understandable reasons for not introducing comprehensive OAT immediately. In the early stages:

  • The customer population is small.
  • Traffic levels are limited.
  • The engineering team understands the whole platform.
  • Production issues can often be managed manually.
  • Recovery depends on a small number of knowledgeable people.
  • The immediate business risk is often failing to reach the market.

Heavy governance at this stage can be disproportionate.

The problem is not that start-ups take risks in those early days. The problem is when their risk controls do not evolve as the business grows.

Early success can hide operational weaknesses

Capable engineers often compensate for missing monitoring, incomplete automation, weak documentation and immature operational processes.

Because incidents are resolved, management may conclude that the operating model is working. In reality, people may have become undocumented parts of the service.

The organisation may have a service that is repeatedly recovered by resilient individuals rather than an inherently resilient service.

Growth changes the risk faster than the organisation changes

Growth brings:

  • More customers and transactions
  • More integrations and data
  • Larger and less predictable peaks
  • More engineers and team boundaries
  • More markets and regulatory obligations
  • Greater dependency on suppliers
  • Greater financial and reputational consequences

The approaches that worked for a small customer base may fail at much larger scale. Meanwhile, engineers can become increasingly occupied managing incidents and production issues, reducing the time available to remove underlying weaknesses.

Failures may expose a weakness. Success applies more customers, transactions, data and expectations to every weakness at once.

This is where OAT belongs

Operational Acceptance Testing is the risk-based process used to determine whether a product, service or significant change can be operated safely, reliably, securely and supportably under normal, peak, degraded and failure conditions.

It is not simply a performance test. It is not a QA activity performed at the end of delivery. It is not a late-stage approval gate.

Its purpose is to provide objective evidence that important operational risks are understood, tested and either controlled or consciously accepted by an accountable decision-maker.

OAT does not guarantee that incidents will never happen. Its purpose is to avoid discovering critical operational weaknesses for the first time during a live incident.

What proportionate OAT should look like

Companies do not need to begin with a large, bureaucratic testing programme. They need risk-based evidence. For each significant release, OAT should answer:

  • What could materially harm customers or the business?
  • What demand and customer behaviour are expected?
  • Which failures and degraded conditions are credible?
  • What evidence will demonstrate readiness?
  • Can the organisation detect, diagnose, contain and recover?
  • Can the change be rolled back safely?
  • Who can accept the remaining risk?

Proportionate does not mean applying every test type to every change. It means selecting evidence according to service criticality, the nature of the change, credible failure scenarios and the level of uncertainty being introduced.

Robust Operational Acceptance Testing infographic: test today, safer tomorrow
Robust OAT converts operational uncertainty into evidence for a readiness decision.

A practical scope may include:

  • Realistic workloads and customer journeys
  • Peak, stress and sustained-load testing
  • Dependency and partial-failure scenarios
  • Backup, restore and recovery exercises
  • Deployment, migration and rollback rehearsals
  • Monitoring, alerting and diagnostic validation
  • Data integrity and reconciliation
  • Support, escalation and incident communications
  • Controlled release and post-release verification

The resulting decision should be clear: GO | CONDITIONAL GO | NO-GO | RE-TEST

Residual risk should be accepted by the accountable service owner rather than transferred to the test team.

Conclusion

Functional correctness is necessary. A service that does the wrong thing is a problem.

But many of the incidents that create significant customer, financial, regulatory and operational consequences do not begin with a functional defect. They begin when a service encounters conditions it was never adequately tested to handle.

The service cannot scale. The dependency fails unexpectedly. The alerting does not expose customer impact. The rollback process has never been rehearsed. The recovery takes longer than expected.

The organisation discovers these weaknesses at exactly the wrong moment: when customers are already affected.

Your live service should not be your operational testing environment, and customers should not be your principal error-alerting mechanism.

The question is whether you want to discover operational limits in a controlled test or during a live incident.

Watchmen helps organisations establish proportionate, evidence-led operational acceptance testing before customer growth, service complexity and production incidents make the gaps expensive. If your releases pass their tests but operational readiness remains uncertain, we should talk.
#OperationalReadiness#Testing#Resilience