TrendSane

Why Digital Public Services Need Failure Budgets

Why Digital Public Services Need Failure Budgets

Published on Sep 29, 2026 · 13 min read

Digital public service reliability should be judged by whether people can still complete essential tasks when technology fails—not merely by whether a website remains online. A shopping site that crashes may frustrate a customer, who can postpone a purchase or take their money elsewhere. A government portal can be different. It may be the only route to file a tax return, renew a licence, receive a benefit, prove identity, register a birth, book a healthcare appointment or appeal an official decision.

When that system fails, the cost is often transferred to the person trying to use it. They may miss a deadline, lose wages while repeatedly calling support, be unable to obtain documentation, or face uncertainty about an application that was submitted but never acknowledged. The service may report excellent uptime while the citizen is stuck inside a broken form, an identity check that cannot be completed, or a queue with no human route out.

This is why government digital services need something more demanding than a performance dashboard. They need a failure budget: an explicit, public agreement about the kinds of failure a service can tolerate, the harm those failures must not cause, and the recovery routes available when they occur. It is not a licence for poor service. It is a way to recognise that complex systems fail and to make agencies responsible for what happens next.

For public institutions, reliability is not only a technical property. It is a condition of access to rights, obligations and shared services. A reliable digital service is one that helps people make progress even when the software stops cooperating.

What a failure budget means in public services

In reliability engineering, teams often use service-level objectives: measurable targets for how a system should perform. An error budget is commonly used in that context to describe the amount of unreliability a service can sustain before it has exceeded its target. The concept helps teams balance the pressure to release new features against the need to protect stability.

A public-service failure budget should adapt that logic rather than copy it mechanically. Government services have different duties from commercial platforms. Their users may have no alternative provider, no ability to defer the task and no realistic way to absorb an error. The question is therefore not simply, “How many minutes of downtime are acceptable?” It is, “What failures can occur without causing someone to lose access, money, time or legal standing?”

A useful failure budget covers more than outages. It establishes tolerances and response commitments for:

  • service unavailability, including planned maintenance during critical periods;
  • slow or timed-out pages that prevent transactions from finishing;
  • failed payments, duplicate charges or missing receipts;
  • identity-verification failures and account lockouts;
  • incorrect records, missing documents and data-transfer errors;
  • applications left unresolved beyond a defined period;
  • accessibility barriers that prevent people from using the primary route; and
  • support failures, including a lack of reachable human assistance.

The point is not to normalise these outcomes. It is to anticipate them, measure them and constrain their consequences. If an agency knows that an online filing service may occasionally fail, it should decide before launch what evidence a person needs to preserve their deadline, who can investigate the case, how quickly that review will happen and how the system will prevent a repeat failure.

That is a more meaningful form of accountability than a generic assurance that a platform is “available.”

Reliability is more than uptime

Uptime remains useful. If a public website is unreachable, that is plainly a reliability failure. But uptime is a narrow infrastructure measure, not a full account of whether a service works for the public.

A system can be technically online and practically unusable. A form may reject a valid address. A document uploader may fail on a mobile device. An authentication process may assume a person has a particular identity document, smartphone or stable connection. A page may meet basic availability targets yet be impossible to navigate with assistive technology. A call centre may be so overloaded that the supposed fallback route exists only on paper.

For citizens, the meaningful outcome is usually completion: the application was received, the payment was correctly recorded, the appointment was booked, the case was reviewed, the record was corrected, or the appeal was preserved. Public sector technology should therefore measure successful outcomes alongside technical availability.

Questions that expose the gap

  • What proportion of people complete the task, rather than merely start it?
  • How often does a completed submission receive a durable confirmation?
  • How many people contact support after attempting the digital route?
  • How long does it take to resolve a disputed transaction or account error?
  • Can people using screen readers, keyboards, translation support, older devices or intermittent connectivity complete the same journey?
  • What happens to an application when a third-party identity, payment or messaging service fails?

These measures move the discussion from server health to public outcomes. They also reveal why averages can mislead. A service may work quickly for people with new devices, strong broadband and familiar documentation while repeatedly failing those who are already navigating language, disability, poverty, unstable housing or complex case histories.

Digital public service reliability is distributional. It matters not only how often a system fails, but who bears the failure and whether they have a way to recover.

The four safeguards every service should publish

A failure budget becomes real only when users can see what protection exists. Every important online government service should plainly publish four safeguards.

1. Fallback paths

People need to know what to do when the main route does not work. Depending on the service and the jurisdiction, that may mean a phone line, a staffed service desk, a postal route, an appointment, an authorised representative, a community-based assisted-digital option or another secure channel. Not every task requires every alternative. But a service that is essential, deadline-driven or difficult to reverse should not leave users with a single fragile path.

A fallback route must be operational, not ceremonial. A phone number that reaches an automated menu instructing callers to return to the broken website is not a contingency plan. Nor is an in-person option that requires a digital booking code the person cannot obtain.

2. Response times

When a transaction fails, uncertainty is itself a burden. Agencies should state when they will acknowledge a report, when they expect to investigate it and when a person can expect a decision or recovery action. The necessary speed will vary. A delayed update to a non-urgent record is not the same as a failed benefit payment, an expiring document or a time-limited appeal.

Response commitments should apply to the citizen’s case, not only to internal incident tickets. A technical team may restore a system quickly while thousands of people remain unsure whether their submissions survived the disruption.

3. Human escalation

There must be a route to a trained person who can examine exceptions, interpret evidence and correct an error. Human escalation is especially important where automated identity checks, eligibility rules, fraud controls or document-processing systems produce a result that does not fit the person’s circumstances.

The human reviewer needs authority as well as access. If staff can only repeat what the interface says, the agency has created an exception queue without a remedy. Effective escalation includes case ownership, the ability to request or assess alternative evidence, a clear record of decisions and a path for further review when the first response does not resolve the problem.

4. Deadline protection

People should not lose eligibility, submission time or appeal rights because a government system failed. Where a documented technical problem prevents completion, agencies should have a defined method for preserving the person’s place in the process. That can include recording an attempted submission, accepting supporting evidence later, extending a deadline, or opening a manual case.

The precise legal mechanisms will differ across services and jurisdictions. The durable design principle is simpler: an agency should not treat its own technical failure as the citizen’s procedural failure.

Forced digitization raises the reliability standard

Digital-first service delivery can make routine administration faster and easier for many people. But making a service digital-only, or making the digital route effectively unavoidable, changes where risk sits. The agency may gain a streamlined workflow while citizens become responsible for devices, connectivity, passwords, document formats, identity tools and error recovery.

Those resources are unevenly distributed. Some people have limited or costly internet access. Some share devices or cannot safely use a shared device for sensitive matters. Some cannot receive verification codes reliably. Others need language assistance, accessible formats, help understanding official terminology or time with a person who can explain what a request means.

Digital inclusion is therefore not an optional outreach programme sitting beside the service. It is part of service continuity and administrative fairness. Offline and assisted options are not evidence that digital transformation has failed. They are evidence that designers understand the public they serve.

This does not mean every government process must retain a paper replica forever. It means agencies should assess the consequences of failure before narrowing channels. The more essential the service, the more severe the consequence of delay, and the less able users are to switch providers, the stronger the fallback obligations should be.

Design failure modes before launch

Resilient public service design begins by assuming that the happy path will not be the only path. Before launch, teams should map likely failure scenarios from the user’s perspective. This exercise should include ordinary technical problems and the awkward cases that are often pushed into manual back offices.

Common scenarios include a payment that leaves a bank account but does not appear in a government record; an identity match that rejects a legitimate applicant; duplicate submissions caused by a browser timeout; a document upload that silently fails; an account locked after a change of phone number; missing data from a legacy system; or a third-party provider outage that interrupts an otherwise functional government service.

A practical tool is a service-level failure matrix. For each scenario, it should identify:

  1. the likely user impact and its severity;
  2. how the person will recognise that something has gone wrong;
  3. the fallback route available immediately;
  4. the accountable service owner;
  5. the target time for acknowledgement, investigation and recovery;
  6. how eligibility, payments or deadlines will be protected; and
  7. what data, audit trail or evidence is needed to restore the case.

This matrix should be tested with real users, including people with disabilities, limited connectivity, low digital confidence, different language needs and urgent deadlines. Technical testing can establish whether a system behaves as intended. Recovery testing establishes whether people can get unstuck when it does not.

Error messages deserve special attention. A useful message tells people what happened in plain language, whether the agency received any part of the submission, what they should do next and where they can obtain help. “Something went wrong” is not merely unhelpful copy. In a public-service context, it can leave someone unable to prove they acted on time.

The human should not become the exception queue

Automation can help public agencies handle high volumes and reduce repetitive work. Yet an automated process is never the whole service. It has boundaries: facts it cannot see, evidence it cannot interpret, life events it does not represent well and mistakes it cannot recognise in its own outputs.

Human escalation should be designed as a core capability, not as a hidden concession for the most persistent users. People need clear contact routes, realistic waiting times and staff who can see the relevant case history. They should not have to retell a failed journey to multiple teams, or repeatedly submit the same evidence through the same malfunctioning interface.

Case ownership is particularly important when services rely on several systems or suppliers. A citizen should not be asked to determine whether the problem belongs to an identity provider, payment processor, software vendor, local office or central agency. The public body remains responsible for the service outcome.

That responsibility requires audit trails. Agencies should be able to establish what the user attempted, what the system did, which automated rules were applied, who reviewed the exception and how a correction was made. Such records support both fairness for the individual and learning for the organisation.

How to measure a public service failure budget

A failure budget should turn vague aspirations into visible operating measures. No single metric can capture reliability, but agencies can combine technical, operational and user-outcome indicators.

  • failed and incomplete transactions, by journey stage;
  • abandonment after error messages or verification checks;
  • repeat contact rates and repeat submissions;
  • cases unresolved beyond published response times;
  • time from failure report to human review;
  • the share of escalated cases successfully recovered;
  • deadline protections or manual interventions triggered by incidents;
  • accessibility-related support requests and completion barriers; and
  • the recurrence of known defects after they have been identified.

These results should be broken down where lawful, appropriate and privacy-preserving. Aggregate success can conceal sharply different experiences among users of assistive technology, people using mobile devices, people in remote areas or those with complex circumstances. The goal is not intrusive profiling. It is to avoid declaring a service reliable because it works for the easiest cases.

Public reporting also matters. Agencies need not expose sensitive security details, but they can publish clear incident histories, service status information, the effect on users, available protections and remediation plans. Honest communication is often more valuable than a polished claim of seamless transformation. It tells people whether they should keep trying, use an alternative route or preserve evidence of their attempt.

Most importantly, thresholds must trigger action. If unresolved cases, failed identity checks or support queues exceed the agreed budget, the response should include investment, redesign, staffing changes or a pause on additional rollout—not merely a postmortem after harm has accumulated.

Procurement and governance must preserve accountability

Government digital services are frequently built from a mix of internal systems, cloud infrastructure, specialist software, identity services and contracted support. Outsourcing technical components does not outsource public responsibility. Reliability commitments must survive the contract boundary.

Procurement should require suppliers to support incident reporting, recovery procedures, accessible service delivery, data portability, documented interfaces and meaningful cooperation during failures. Agencies should understand how to retrieve records, continue critical work and transition away from a supplier if necessary. A contract that leaves a public body unable to inspect, correct or recover a failed workflow creates governance risk as well as technical risk.

Each service also needs a named public owner with authority across organisational silos. That owner should be accountable for the end-to-end journey, including the handoffs users do not see. If a resident cannot complete a task because two suppliers each claim the fault lies elsewhere, the service has failed regardless of its contractual arrangements.

Failure budgets can build trust without promising perfection

No complex public system will be failure-free. Software has defects, networks fail, records conflict, demand surges and people’s lives refuse to fit neat categories. Pretending otherwise encourages brittle design and defensive communication.

A failure budget offers a more credible promise: failures will be detected, their impact will be limited, people will have another route, and the institution will remain answerable until the case is resolved. That is not a lower standard. It is a mature standard for systems that people depend on.

Trust in government is shaped not only by whether rules are fair in theory, but by what happens when those rules meet real life. A person who encounters an error may accept that technology sometimes fails. They are less likely to accept being trapped, ignored or penalised for it.

Build an exit ramp into every digital service

The essential question for online government services is not whether a digital journey looks seamless on a successful day. It is whether a person can still move forward on a bad one.

Failure budgets make that question operational. They require agencies to define acceptable levels of failure, monitor the outcomes that matter, publish fallback paths, protect deadlines and provide meaningful human accountability. They also make digital inclusion part of reliability rather than an afterthought reserved for users who cannot conform to an idealised online journey.

Citizens should never be trapped inside a failed interface. Digital government becomes genuinely reliable when a technical error does not end the person’s ability to act, be heard or receive the service they need.

Image by ArtisticOperations on Pixabay.