Reliability & Resilience · Keep the light on, whatever the weather

When your app goes dark, who notices first?

Reliability and resilience engineering means making sure what your customers depend on keeps working and stays quick to respond, hearing about trouble before they do, and coming back calmly when something breaks, because you practised for it.

Send a storm

Who’s who in the story

In our story, Ingrid keeps the lighthouse and Bosun is her seal. The lighthouse is your app or website, and the ships are your customers. A storm is a rush of customers, a bad update or plain bad luck. The gauges, the one bell that matters and the buoy far out (a pretend customer trying your app all night) are your early warnings; the bell stuffed with a sock is an alarm nobody hears. The spare lamp and the oil in another shed are your backups. Tested on a calm day, the spare lamp spluttered: Bosun had hidden a fish in its oil can, a weak spot nobody knew about.

What does the keeper have ready?
Pick a storm

For you: a part simply fails

Plain bad luck: a long dark night.

How long the light is outa long dark night
  1. Who notices first?A ship notices first: “Hello? Anyone there?”
  2. How long is it dark?The spare splutters, and out flops a fish. A long dark night.
  3. Who does what?Everyone runs to the lamp. The radio goes quiet. Bosun hides in the foghorn.
  4. Next morningThe logbook asks: who hid the fish?

Plain bad luck: a long dark night. A ship notices first: “Hello? Anyone there?” The spare splutters, and out flops a fish. A long dark night. Everyone runs to the lamp. The radio goes quiet. Bosun hides in the foghorn. The logbook asks: who hid the fish?

Nothing ready · a long dark night

  • Who notices first? A ship notices first: “Hello? Anyone there?”
  • How long is it dark? The spare splutters, and out flops a fish. A long dark night.
  • Who does what? Everyone runs to the lamp. The radio goes quiet. Bosun hides in the foghorn.
  • Next morning The logbook asks: who hid the fish?

Everything ready · a short blink

  • Who notices first? The buoy blinks red and one bell rings. Ingrid hears first.
  • How long is it dark? A short blink, and the tested spare takes over.
  • Who does what? Ingrid: lamp. Bosun: foghorn. The radio: someone named on the plan, talking to the ships. The ships get home.
  • Next morning The logbook asks: what do we fix?

Worth knowing: Practice doesn’t stop storms. It aims to make the dark shorter and the recovery calmer, and some problems still surprise everyone.

An illustration, not a measurement. In every storm, something still breaks.

Watch · 2 min 14 sec

Reliability & Resilience: the lighthouse story

A lighthouse goes dark, and the first to notice is a passing ship. A short story about noticing early, practising on calm days and keeping a spare you have tested, so the next storm is a bit boring.

Read the story instead
  1. The shipHello? Anyone there?
  2. On screenWhen a lighthouse goes dark, who notices first?
  3. NarratorWhen a lighthouse goes dark, who notices?
  4. NarratorNot Ingrid the keeper, or Bosun her seal. The ship did.
  5. NarratorThe oil ran out. Its warning bell? Stuffed with a sock.
  6. NarratorYour app is a lighthouse. Customers are the ships, and they often notice first.
  7. On screenapp = lighthouse, ships = your customers
  8. NarratorStorms will come, like a rush of customers or a bad update. So expect breaks, notice early, bounce back.
  9. On screenstorm = a rush of customers · a bad update · bad luck
  10. NarratorThe grown-up name: reliability and resilience.
  11. NarratorFirst, notice early. Gauges for oil and wind.
  12. NarratorA buoy far out checks the light like a ship would. For you: a pretend customer, trying your app all night. So Ingrid hears first.
  13. NarratorOne bell, before the oil runs out. Too many, and nobody listens.
  14. NarratorSo on a calm day, Ingrid switches the main lamp off. On purpose.
  15. BosunOn PURPOSE?!
  16. NarratorThe spare lamp should take over. It splutters, and out flops a fish.
  17. BosunIt was hiding from storms!
  18. NarratorA weak spot, found on a calm day, not in a storm.
  19. NarratorEnough radios and hands for the busiest night. For you: your busiest sales day.
  20. NarratorToo slow can be as bad as broken.
  21. NarratorA tested spare lamp, and spare oil in another shed. And a plan for who does what.
  22. NarratorThis time, Ingrid knows first. A short blink, and the spare takes over.
  23. On screennot perfect, just ready
  24. BosunFoghorn! HONK!
  25. The shipHello? Anyone there?
  26. IngridRight here. Come on home.
  27. NarratorShips home. A bit boring.
  28. NarratorNext morning: what to fix, not who hid the fish.
  29. NarratorAlgoshred looks for weak spots with you, and tests spares.
  30. NarratorWe add early warnings and agree with you what’s reliable enough. Then we practise switching to spares on calm days.
  31. NarratorWhere’s the fish hiding in your business? Visit algoshred.com, or write to contact@algoshred.com.

The idea in plain words

Four ideas from one calm keeper

Each one in everyday words first, then the grown-up name experts use for it.

A papercut island lighthouse at dusk, its amber beam sweeping over the sea towards a small buoy and a trawler

Practise the storm before it comes.

  1. 01

    Aim to know before your customers do

    A few gauges tied to what customers feel, a pretend customer trying your app all night, and one bell that rings only when someone needs to act. Too many bells, and nobody listens.

    Experts call this: monitoring, alerting, observability, synthetic checks

    Worth knowing: Gauges show trouble early, but someone still has to look, and some storms come without warning.

    See which bells should ring
  2. 02

    Practise failure on a calm day

    Switch one part off on purpose, small and planned, with a hand on the switch, and find the hidden weak spot before a storm does.

    Experts call this: chaos engineering, game days

    Worth knowing: Practice switch-offs start small, on a calm day, with a way to stop. We agree what is off-limits before anything is switched off.

    See a calm-day switch-off
  3. 03

    Plan for your busiest day

    Enough room and hands for the busiest night, tried before it arrives. Too slow can be as bad as broken.

    Experts call this: performance and capacity, load testing

    Worth knowing: Trying your busiest day ahead shows where things strain. Real crowds can still surprise you, so someone should be watching on the day.

    Try a busy night on a quiet week
  4. 04

    A tested spare and a plan

    A spare you have switched to on purpose, copies kept somewhere else, and a plan on the wall for who does what, so you can keep going another way while the main thing is fixed.

    Experts call this: backups, disaster recovery, business continuity, incident response

    Worth knowing: A spare that has not been tested is a hope, not a plan. Keep copies somewhere else, and if a mistake gets copied into your backup too, you need an older, dated copy to go back to.

    See why a tested spare matters

Bells that mean something

Too many bells, and nobody listens

Every bell on Ingrid’s wall rings for something. Sort them: which should wake a person, and which belong in the logbook? Then run a night and hear the difference.

For you: The bells are the alerts your team gets. The buoy is a pretend customer trying your app all night.

A wall of clanging brass bells, one with a woolly sock stuffed into its mouth and one ringing in amber light, while the keeper watches and the seal guiltily hugs another sock

Which bells should wake someone?

Pressed bells wake someone. The rest go in the logbook.

Every bell rings tonight. Sort the wall, then press Run a night.

Ding, ding, ding. Ingrid stops listening, and Bosun reaches for a sock. The bell that mattered is lost in the noise.

Wakes someone

  • Gulls: a harmless blip
  • Wet paint: a routine change was made
  • Kettle: a routine job finished
  • Post: the daily report arrived
  • Tide: the usual busy spell of the day
  • Fog: pages are getting slow for customers
  • Oil low: something is about to run out
  • Lamp out: customers can’t get in (Bosun muffled it with a sock)

Goes in the logbook

  • Nothing here yet. Press a bell to send it here.

The grown-up name:monitoring and alertingalert fatigueobservabilitysynthetic checks (the buoy)

Worth knowing: Gauges show trouble early, but someone still has to look, and some storms come without warning.

An illustration, not a measurement.

Calm-day practice

Switch one thing off, on purpose, on a calm day

Pick one thing, keep the test small, and pull the lever. What everyone guessed and what really happens are often different, and a calm day is the time to find out.

On a calm grey day the keeper holds a brass lever with the main lamp switched off and a small spare lamp glowing, while the seal hides behind the lens
What will you try on a calm day?
How much do we switch off?
only this lamp, only a moment
  1. What should happen

    Ships can still see the light.

  2. What everyone guessed

    The spare takes over.

  3. What the practice showed

    The spare splutters, and out flops a fish. Bosun faints.

  4. Fixed on the calm day

    Can cleaned, spare tried again: it glows.

The grown-up name:chaos engineering, failover test

Worth knowing: Practice switch-offs start small, on a calm day, with a way to stop. We agree what is off-limits before anything is switched off.

More worth knowing

Worth knowing: A spare that has not been tested is a hope, not a plan. Keep copies somewhere else, and if a mistake gets copied into your backup too, you need an older, dated copy to go back to.

Worth knowing: Trying your busiest day ahead shows where things strain. Real crowds can still surprise you, so someone should be watching on the day.

An illustration, not a measurement. A hand stays on the switch the whole time.

Sounds familiar?

The signs your customers notice before you do

If any of these ring true, your customers are probably finding the problems before you do.

  • “It went down, and we found out from an angry customer.”
  • “Our biggest day was the day it fell over.”
  • “We had backups. Then we tried to bring one back.”
  • “It isn’t down, it’s just so slow that people give up.”
  • “The alarms ring all night, and most of them are nothing.”
  • “When it broke, everyone panicked and nobody knew who does what.”

A quick self-check

Seven plain questions. Answer what you can, and we’ll point to where we’d look first.

01If your app went down tonight, would you hear before a customer told you?
02When an alarm rings, does it usually need someone to act?
03Have you agreed what reliable enough means for your most important service?
04Have you brought back a backup, or switched to a spare, on a calm day recently?
05Have you tried your busiest day before it arrived?
06Does everyone know who does what when something breaks, even if one person is away?
07After an outage, does the review end in a fix rather than a name?

Where we’d look first

Answer any question and the places we’d look first appear here.

Talk it through

Answer “No” or “Not sure” to any question to email the list.

Nothing is stored or sent unless you choose to email it.

What we help with

From one bad night, looked at calmly, to a plan your team practises

Eight pieces of work. Start with the one that worries you most, or bring them together service by service.

Notice

Hear about trouble before customers do.

A reliability check-up, and “reliable enough” agreed

A short list of the risks that matter most, and targets everyone has agreed.

We look at what your customers rely on most, your past outages, your alarms and backups, and who answers at night. Perfect isn’t possible, and chasing it costs more than it’s worth. We help you agree what reliable enough means for each part, with the people who own and use it.

Experts call this: reliability targets (SLOs), error budgets, site reliability engineering (SRE)

Gauges and bells that mean something

Trouble can show up before customers notice, and nights on alarm duty can get quieter.

Readings tied to what customers feel, a pretend customer that tries your key pages around the clock, and alarms reworked so fewer ring and each one asks for action.

Experts call this: observability, monitoring and alerting, synthetic checks

Also in this area

  • Reliability targets agreed with the business (SLOs, and an error budget: how much trouble is acceptable before new work pauses for fixes)
  • Dashboards that follow a customer’s path through your app (customer-journey dashboards), built on the four golden signals: how long people wait, how many are asking, how many are turned away, and how full you are
  • Alarm clean-up, so each bell means act now
  • Pretend-customer checks, and readings of how your key pages feel to real visitors (real-user monitoring)
  • Being able to look inside and find out why something happened (observability)

Practise

Find the weak spots on a calm day.

Switching things off on purpose, on a calm day

So weak spots can be found and fixed on a calm day rather than in the storm.

Small, planned experiments that switch one part off with agreed limits and a way to stop, plus team practice days that rehearse the people as well as the systems. Practice switch-offs start small, on a calm day, with a way to stop. We agree what is off-limits before anything is switched off.

Experts call this: chaos engineering, game days

Ready for your busiest day

Your busiest day becomes something you have tried before, and slow pages get the attention outages do.

Trying the crowd before it arrives, finding the slow parts, and planning the room you need for sales, launches, results days and festival rushes. Trying your busiest day ahead shows where things strain. Real crowds can still surprise you, so someone should be watching on the day. If you’d like, we can watch alongside your team on the day.

Experts call this: performance engineering, load testing, capacity planning

Also in this area

  • Small, limited switch-off experiments (chaos engineering)
  • Team practice days that rehearse people as well as systems (game days)
  • Load, stress and long-running tests before your big days
  • Storms walked through on paper with leaders (tabletop exercises)

Recover

Keep one break from becoming many, and come back calmly.

Built to bounce back

Aims to keep one failure from taking everything else down with it.

Removing “only one of” weak spots, rolling changes out a little at a time with an off switch, polite limits on repeated requests, and a simpler working mode when a part struggles.

Experts call this: high availability, safe rollouts, graceful degradation, rate limits

Spares and backups you have tested

“How long to come back” becomes something you have practised, not guessed.

Choosing how ready each spare should be, protected copies kept apart from everyday systems, and supervised restore and switch-over practice. A spare that has not been tested is a hope, not a plan. Keep copies somewhere else, and if a mistake gets copied into your backup too, you need an older, dated copy to go back to.

Experts call this: disaster recovery, backup and restore testing, recovery time and data-loss limits (RTO/RPO)

Also in this area

  • Removing single points of failure (the “only one of” weak spots)
  • Gradual rollouts with off switches
  • Limits on repeated requests, and simpler working modes when a part struggles
  • Protected, separate backups with regular restore tests
  • Choosing how ready each spare should be, service by service: a spare in a box (backup and restore), the essentials kept ready but switched off (pilot light), a small spare already running (warm standby), or two lamps lit together (active-active)

Plan and people

A plan everyone has practised.

Calm incident handling and no-blame reviews

So people know their part when the bell rings, and the same thing is less likely to catch you out twice, because the fix is written down.

Clear roles, written steps for known problems, message templates for customers, and a review habit that asks what to fix, not who to blame.

Experts call this: incident response, on-call, runbooks, blameless post-incident reviews

The storm plan, people and suppliers included

A plan that gives everyone their part on the bad day, and records of what you practised.

Deciding what comes back first, a plan that covers people, suppliers and customer messages, storms walked through on paper, and the test records your compliance team asks for.

Experts call this: business continuity, business impact analysis, tabletop exercises, operational resilience

Also in this area

  • Incident roles, written steps (runbooks) and customer messages
  • No-blame reviews that end in a fix
  • Continuity plans that cover people and suppliers
  • Records for financial regulators’ reviews (for example the EU’s Digital Operational Resilience Act, DORA; the UK’s Financial Conduct Authority and Prudential Regulation Authority, FCA and PRA; or India’s Reserve Bank and securities regulator, RBI and SEBI), prepared with your compliance and legal teams (whether they meet a rule is their call, not ours)
  • Training and handover, including what AI helpers may do on alarm duty (sort alarms and suggest next steps; a person approves any change to the live system)

For your technical team: we work with what you already use, for example OpenTelemetry, Prometheus, Grafana, Datadog, PagerDuty, k6, JMeter, Chaos Mesh, LitmusChaos, AWS Fault Injection Service, Azure Chaos Studio, Velero and your cloud’s own backup tools.

A path across the rock to a second shed holding a spare oil can and a spare lamp, with a blank plan board by the lighthouse door

Ways to start

Walk through one bad night

A no-blame look at one real outage, and what would have helped.

One storm on paper

A walk-through on paper (a tabletop exercise) for one service you worry about.

Test your spare

A supervised restore or switch-over practice on one service, in a safe window.

Tell us about one bad night

How we work

Look first, agree what’s reliable enough, then practise on calm days

Five steps in plain words, from the first look to your team keeping the habit.

  1. Look

    Look for the weak spots, with you

    We map what your customers rely on most and read the last few outages without blame. Together we find the “only one of” weak spots: one server, one supplier, one person who knows.

  2. Reliable enough

    Agree what’s reliable enough

    We agree targets for each part with the people who own it: how long it may be down, and how much recent work you could afford to lose. Perfect is not the goal: the payments desk and the staff newsletter deserve different answers.

  3. Early warnings

    Add gauges and one bell that matters

    We watch what customers feel, add a pretend customer that tries your key pages, and rework the alarms so each one means act now. The rest goes in the logbook.

  4. Calm days

    Practise on calm days

    On a calm day we switch one part off on purpose, try the busiest day before it arrives, restore a backup and walk a storm through on paper. Each surprise gets written down and fixed.

  5. Keep practising

    Your team keeps the plan and the habit

    Your team owns the plan, the written steps and the practice calendar, and we train them to run it. We can share the watch for a while, or come back for regular practice days, if you prefer. It’s built with your team, one service at a time, not handed over in a box.

What you get along the way

  • A map of what customers rely on most, and its “only one of” weak spots
  • Reliability targets agreed for each part
  • Reworked alarms and customer-journey dashboards
  • One calm-day switch-off, with what was found and fixed
  • A tested restore of one backup
  • A storm plan with roles and customer messages
  • Hand-over sessions and a practice calendar for your team
The morning after a storm: an open logbook and a mug on a table, the seal peeking out, and boats safe in the harbour beyond the window

Where it fits

Where it becomes real

Illustrative examples, not customer stories: the kind of work this approach suits.

Retail and e-commerce
For example, before a festival sale the checkout could be tried with a crowd that hasn’t arrived yet, the slow step found, and the team agreed on what gets switched off first if the night gets rough.
Banking and payments
For example, a payments service could have a spare that is practised on a schedule, with records of each practice ready when a partner bank asks how recovery would work.
Healthcare
For example, a clinic’s booking system could have one bell for “patients can’t book”, and a written plan for taking appointments by phone while it is being fixed.
Logistics
For example, when the tracking page struggles, it could show a simpler view instead of an error, while polite limits stop every app asking again and again at once.
Education and ticketing
For example, results day or a ticket release could be tried on a quiet week, so the busiest hour is not the first time the system meets that crowd.
Teams adding AI helpers to night-time alarm duty
For example, an AI assistant could sort alarms and suggest next steps from clean readings and written steps, while a person still approves any change to the live system.

Good to know

Questions we are often asked

Plain answers to what owners and tech leads ask first. Something else on your mind?

Ask us directly

What is reliability and resilience engineering, in plain words?

Making sure what your customers rely on keeps working, hearing about trouble before they do, and coming back calmly when something breaks, because you have practised for it.

Is this the same as DevOps, or as having an operations team?

It overlaps with both. DevOps is about builders and operators sharing the work of shipping and running software; reliability engineering zooms in on keeping what’s shipped working. Site reliability engineering (SRE) treats that as an engineering job: agreed targets, useful alarms, practised recovery and fixing causes, so it depends on habits rather than heroes.

Can you stop it from ever going down?

No one honestly can. Perfect isn’t possible, and chasing it costs more than it’s worth. We help you agree what reliable enough means for each part, with the people who own and use it. Then we watch it, and practise recovery so bad days are shorter and calmer. Experts call these targets SLOs, and the room they leave for trouble an error budget.

Isn’t switching things off on purpose risky?

Practice switch-offs start small, on a calm day, with a way to stop. We agree what is off-limits before anything is switched off. Often the first practice happens in a test copy or on paper. Finding a weak spot that way is gentler than finding it in the storm.

Do we need new monitoring tools, or a second data centre?

Not necessarily. We start with what matters most and the tools you already have. How ready a spare needs to be is a business choice, made service by service: the payments desk and the staff newsletter deserve different answers.

Can you help with what our regulator expects?

We can help you map your important services, plan and run the practice, and keep records of what was tested, working with your compliance and legal teams. Whether that meets a particular rule (for example the EU’s DORA; the UK’s financial regulators, the FCA and PRA; or India’s central bank and market regulator, the RBI and SEBI) is their call, not ours.

How will we know it’s working?

We agree a few plain signs with you at the start: who notices problems first (you or your customers), how long until someone is working on it, when each spare was last tried, and whether each review led to a fix. The numbers stay with you, and we make no promises about them.

What happens after the project ends? Do you answer the alarms at night for us (on-call)?

Your team keeps the targets, alarms, written steps, plan and practice calendar, and we train them to run it. If you prefer, we can share the watch for a while or come back for regular practice days; we agree the scope in writing first.

Heard it before?

Myths, answered

Myth: “Good systems don’t break.”

Answer: Everything breaks sometimes. Good teams expect it, notice early and bounce back.

Myth: “We have backups, so we’re fine.”

Answer: A backup you have not tried bringing back is a hope, not a plan. In our story, the spare lamp spluttered: Bosun had hidden a fish in its oil can.

Myth: “Switching things off on purpose is reckless.”

Answer: Practice switch-offs start small, on a calm day, with a way to stop. We agree what is off-limits before anything is switched off. Finding the weak spot on a calm day beats finding it in the storm.

Myth: “We should aim for perfect.”

Answer: Perfect isn’t possible, and chasing it costs more than it’s worth. We help you agree what reliable enough means for each part, with the people who own and use it.

Myth: “More alarms means safer.”

Answer: Too many bells, and people stop listening. A bell should ring when someone needs to act.

Myth: “Outages are someone’s fault, so find them.”

Answer: Outages usually come from gaps in the steps, not bad people. The morning logbook asks what to fix, not who hid the fish.