15:51 · 18 września 2026

This Is Fine? Yes, Indeed. 6 Patterns for Building More Resilient Systems

Najważniejsze wnioski
What happens when a system that looks perfect in testing meets the reality of production? Failures, retries, lost messages and inconsistent data can expose weaknesses that aren’t visible on a dashboard full of green indicators. In this article, Wojciech Kołodziej, our Principal Software Architect explores six architectural patterns that can help make distributed systems more resilient - from Circuit Breaker and idempotency to optimistic locking, Outbox/Inbox and Saga. You’ll also learn why resilience isn’t just a technical decision, but a business one, and how thoughtful architecture can help systems recover safely when things inevitably go wrong.
Spis treści
Najważniejsze wnioski

What happens when a system that looks perfect in testing meets the reality of production? Failures, retries, lost messages and inconsistent data can expose weaknesses that aren’t visible on a dashboard full of green indicators. In this article, Wojciech Kołodziej, our Principal Software Architect explores six architectural patterns that can help make distributed systems more resilient - from Circuit Breaker and idempotency to optimistic locking, Outbox/Inbox and Saga. You’ll also learn why resilience isn’t just a technical decision, but a business one, and how thoughtful architecture can help systems recover safely when things inevitably go wrong.

You build a system that looks flawless on paper and in testing: a well-designed domain architecture, clearly defined bounded contexts, and carefully designed tests. Everything runs like clockwork, and your monitoring dashboards glow with a reassuring green.

And yet, every now and then, something goes wrong in production. A few messages disappear, or the system starts behaving in a way that nobody expected.

I still remember a sleepless night caused by an incident that resulted in accounts being charged multiple times. A small bug, missing idempotency, and an endless retry loop led to a situation where customer accounts were wiped out. In another case, incorrectly implemented queue communication caused some messages to be lost, resulting in incomplete accounting entries.

All it took was one seemingly harmless deployment to break data consistency and block business processes for several hours. Our code was doing exactly what we had designed it to do. The problem was that we had designed it for an ideal world - a world without failures.

We all know the iconic This Is Fine meme, where a dog calmly drinks coffee while sitting in a room on fire. Looking back, I’ve learned that we should design systems in a way that lets us eventually say, calmly and with genuine satisfaction “This is fine.”

System resilience is primarily a business decision, not a technical one

When we think about resilience, it’s easy to fall into a technical mindset. We start discussing library configurations, timeout values, or thread pool parameters.

But true resilience starts much earlier - at the level of business process logic.

In my work as a Principal Software Architect at XTB, I’ve learned that the key is to change the perspective. Instead of asking, “Which library should I use to protect my system from failures?”, we should ask, “What business outcome are we trying to achieve, and what risks does that outcome need to withstand?”

Resilience also comes at a cost and it’s a cost we need to be willing to accept. It doesn’t always make sense to implement every possible resilience pattern. Sometimes, accepting a certain level of risk is more reasonable when the cost of eliminating it outweighs the potential benefits.

With this approach, business and engineering can start speaking the same language. We talk about states, ownership, and consequences rather than technology for technology’s sake.

So, what does this mean in practice?

Let’s look at six patterns that can help make distributed systems more resilient.

Circuit Breaker: A safety switch for your resources

What is it?
Circuit Breaker is a classic protection mechanism for service-to-service communication. When an external service starts returning large numbers of errors or stops responding, the Circuit Breaker “opens the circuit”. Instead of sending more requests and waiting for timeouts, the system immediately returns a fast, controlled fallback.

When should you use it?
Whenever you integrate with external systems or services whose failure could tie up your application’s resources, such as database connection pools or HTTP threads. It’s particularly useful when a process involves multiple stages of communication between dependent services.

What does it give you?
An immediate signal that a system is unavailable. Instead of waiting for a response that may never arrive, we move straight to a fallback plan. From a business perspective, this can be the difference between “one payment method is temporarily unavailable” and “the entire application is down”.

One of the biggest advantages of a Circuit Breaker is that it can also test whether a service has recovered. After a certain amount of time, the breaker can enter a half-open state, allowing a small number of test requests through. If they succeed, the circuit closes and normal operation resumes. If they fail, it stays open and continues protecting the system from further failures.

It’s also worth asking whether you actually need a Circuit Breaker when working with reactive systems that don’t synchronously block a thread pool. In that case, there may be no risk of exhausting the thread pool, but requests can still accumulate and queue up in memory and somewhere at the end of that chain, someone is still waiting for a response.

Of course, you don’t need to implement this pattern from scratch. There are ready-made libraries for virtually every language, such as Resilience4j for Java or Polly for .NET. The library itself, however, is only part of the solution. The important thing is configuring the thresholds for opening and closing the breaker correctly and defining an appropriate fallback.

And the technical definition isn’t enough. It’s worth asking one more question: “What happens when the Circuit Breaker opens?” Will the customer see an error message, or will the system automatically switch to an alternative provider?

That’s where the real design decision begins.

Idempotency: Safe retries

What is it?
Idempotency ensures that sending the same request again, identified by a unique idempotency key, produces the same result without triggering duplicate side effects in the system.

When should you use it?
Whenever a client or external system may retry an operation after losing a connection or experiencing a timeout. Personally, I recommend using idempotency wherever we’re dealing with state-changing operations, such as charging a card or posting a financial transaction.

What does it give you?
It removes a fundamental dilemma for clients and client systems: “I got a timeout, so did the transaction go through? Should I retry, or am I risking a duplicate charge?” In payments, idempotency is one of the foundations of trust. Nobody wants to be charged twice for the same purchase.

Many people starting out in software development misunderstand the relationship between transactions and HTTP communication. It’s important to remember that HTTP communication is not transactional in itself and is inherently unreliable.

Think of it like sending a letter. You send it, but if you don’t receive confirmation, you don’t know whether the package arrived, whether the recipient read the contents, or whether the return receipt was lost on the way back.

In that situation, the safest option is to send the letter again but in a way that lets the recipient recognize it as the same letter and avoid treating it as a new one.

That’s exactly the role idempotency plays in HTTP communication and distributed systems.

From the recipient’s perspective, every request without such an identifier can be interpreted as a new, independent operation. That’s why a lack of idempotency in systems that support retries can lead to serious consequences.

So, how do you implement idempotency in practice?

A common approach is to use a unique identifier for each operation, generated on the client or server and sent in the request headers. The server stores these identifiers and checks whether a given operation has already been processed. If it has, the server returns the result of the previous operation instead of executing it again.

Sometimes there’s a temptation to return a dedicated status code for an idempotency conflict. In practice, it can be simpler to return the original successful result so that the client doesn’t have to handle an additional case. An important exception is when the same idempotency key is used with a different payload. In that situation, returning a 409 Conflict is a common approach.

There’s one more important scenario: handling a major outage.

Imagine maintaining a highly complex system that processes millions of operations every day. Its components have been developed over years, and nobody remembers every process and dependency anymore.

When such a system goes down and is later restored, some operations may have been processed while others may not. Figuring out exactly what happened and how to safely resume every process can be extremely difficult.

Idempotency allows us to safely retry operations even when we’re not sure which ones have already been processed. We can resume processing without risking duplicate execution and maintain data consistency.

At 3 a.m., a “Retry and forget” button can feel like a blessing.

Versioning and optimistic locking

What is it?
Optimistic locking adds a sequential version number to a record. Every update checks whether the version in the database is still the same as the version we originally read. If someone has modified the data in the meantime, the transaction is rejected.

Every change increments the version number, allowing us to detect conflicts caused by concurrent updates.

When should you use it?
In concurrent environments where a process can be split into smaller parallel steps and responses can arrive in a different order. It also works well in systems where multiple people work with the same data, such as document editing or reservations.

What does it give you?
It prevents silent data corruption caused by lost updates. A write conflict becomes explicit, allowing us to handle it safely and retry the update in a controlled way.

For large relational data structures, it’s worth remembering that the version should be maintained at the level of the main record rather than individual child tables. This is particularly important when working with DDD aggregates.

If we change data in a child table, the aggregate as a whole has changed and its version should be updated as well. Otherwise, changes to child records may go undetected by the versioning mechanism, potentially leading to conflicts and data loss.

When there are many concurrent updates, optimistic locking can result in frequent conflicts and rejected transactions, making it inefficient. In such cases, it may be worth considering other approaches, such as pessimistic locking or event sourcing with immutable messages — but those are topics for another article.

One relatively simple and effective approach is to split the process into smaller processes, each with its own state records that can be updated independently. A separate scheduled process can then monitor the overall state of the task and trigger the next steps when needed.

Acknowledging a task: Save first, respond second

 

  • What is it?
    This architectural principle says that a system should confirm that an operation has been accepted for processing only after the request has been durably stored in a database together with a unique correlation ID.

 

  • When should you use it?
    In asynchronous processes and wherever processing is complex and involves multiple stages.

 

  • What does it give you?
    It eliminates “black holes”. Even if the server crashes a fraction of a second after sending the confirmation, the request is safely stored and can be automatically resumed by recovery processes.

A natural extension of this principle is the use of Outbox and Inbox patterns, which we’ll cover in the next section. But the general rule is simple: secure the information first, then send the confirmation.

And this doesn’t always mean writing to a database. If we’re implementing a proxy adapter, for example, the confirmation should only be sent after we’ve transformed and sent the data to the external system and received confirmation that it has accepted it.

One less obvious example is consuming messages from a queue.

A common mistake is to rely on the queue’s default configuration, which automatically acknowledges a message as soon as it is read. In practice, if the message-processing service crashes after reading the message but before saving the result, the message can be lost.

That’s why it’s worth using manual acknowledgements and acknowledging the message only after it has been fully processed and the result has been persisted.

It’s also important to remember that default transactions in application frameworks such as Spring apply only to a single local database. Distributed transactions such as two-phase commit (2PC) can be configured, but in practice they can be expensive and inefficient.

Outbox and Inbox: A reliable bridge between the database and the broker

 

  • What are they?
    The Transactional Outbox pattern ensures that changes to the database and the creation of a message to be published to a broker, such as Kafka, are coordinated through a local database transaction. The message is first written to a dedicated Outbox table, and a separate process then publishes it to the broker. The Inbox pattern works on the receiving side. It records incoming messages, helps detect duplicates, and ensures that messages can be safely processed at least once.

 

  • When should you use them?
    When communicating with external systems while relying on transactions provided by local databases. These patterns are particularly useful when you need to avoid losing messages while keeping database changes consistent with message publication.
  • What do they give you? They help provide at-least-once delivery semantics and reduce the risk of ending up in a situation where the database has been updated but the broker knows nothing about the change or the other way around.

There’s an important catch, though: these patterns need to be combined with idempotency because the same message may be delivered more than once.

With Outbox, we first store the message in a dedicated table in the same local database transaction that updates the main business table.

A separate process, such as a worker, then reads messages from the Outbox table and publishes them to the broker. Once the broker confirms receipt, the message can be marked as sent or removed from the Outbox table.

If the publishing process fails, the messages remain in the Outbox and can be processed again.

Inbox works in a similar way on the receiving side. When a message arrives from the broker, we record it in the Inbox table and check whether it is a duplicate, for example by using an idempotency key. We can then process the message and update the system state in a controlled way.

This approach makes message processing resilient to failures and restarts, but it doesn’t magically make a distributed system “exactly once”. Duplicate delivery still needs to be handled explicitly.

There’s another important piece here: retry handling.

If the system crashes, restarts, or temporarily loses power, Outbox and Inbox processing should be able to resume from where it stopped. Database transactions help us maintain a consistent local state, while messages that haven’t been processed remain available in the Outbox or Inbox tables.

Processing can resume automatically when the system starts again. For large message volumes, however, it’s worth considering queuing and batching mechanisms so that recovery doesn’t overload the system.

These patterns are closely connected to local database transactions. However, when designing high-performance distributed systems, it can sometimes be more effective to move away from transactional guarantees between services altogether and use a well-designed Saga instead.

In such an approach, completing a process step means changing its state through orchestration or publishing an event through choreography. Any incomplete step can be retried from a known state, while idempotency protects the system from inconsistent results.

And that brings us to the final pattern.

Saga: Controlling long-running processes in distributed systems

  • What is it? A Saga breaks a distributed, multi-step business process into a series of smaller local transactions. After each step is completed, its state is persisted and an event or command triggers the next stage.

 

  • When should you use it?
    When a business process involves multiple independent microservices, such as reserving a product, processing a payment, and arranging delivery, and you don’t want or cannot use distributed database transactions such as 2PC.

 

  • What does it give you?
    Instead of relying on the illusion of atomicity across an unreliable network, you gain control over the entire process. You know exactly which stage an operation has reached and whether it needs to be compensated or moved forward.

With orchestration, you get a clear central control point. With choreography, you can achieve a more decentralized and flexible model.

Before listing the benefits of this pattern, it’s worth pointing out that in the event of a failure, we’re generally talking about compensating actions rather than a classic rollback.

In distributed systems, we can’t assume that every microservice will be able to undo its changes. Instead, we design processes that can handle the situation correctly.

For example, if a customer’s account is blocked during a payment process, the Saga can initiate a refund and cancel the product reservation.

Technical failures, such as network outages or temporary unavailability, are a different matter and should generally be handled through mechanisms such as retries.

The Saga pattern can be implemented in two main ways: orchestration and choreography.

Orchestration

In this approach, a central component - the orchestrator - manages the entire process.

The orchestrator sends commands to individual microservices, waits for their responses, and decides when to move to the next step or trigger a compensating transaction in case of an error. Orchestration is relatively straightforward to implement. We can create a state machine diagram describing all possible states and transitions in the process. The state is maintained in a database, and the process is divided so that each step is an independent local transaction that ends with the state being persisted. If something fails, we can resume the process from the last saved state. Orchestration is highly readable, easy to diagnose, and works particularly well for sequential processes with a relatively small number of branches.

Choreography

In this approach, there is no central component managing the process. Instead, the process state is propagated through events published by individual microservices. Each service reacts to events and decides whether to move to the next step or trigger a compensating transaction. Each service is responsible for its own part of the process and communicates with other services through events. There is no single central point of control, and new states and transitions can be added more dynamically. This approach can be more scalable and flexible, but it is also more difficult to implement, design, and diagnose. It requires greater discipline when designing events and handling them to avoid unwanted side effects.

Before choosing this approach, it’s worth spending time on a solid design, consistent testing practices, and proper event monitoring.

Closing the loop

In the This Is Fine meme, the dog says these words while calmly looking at the flames around him. He’s fully aware of the situation and its consequences. I’m convinced the situation would look very different if he knew there was a fire but couldn’t see its scale or what exactly was burning. I assume none of us would say “This is fine” unless we were confident that things really were fine. Monitoring and alerting around these patterns is a great topic for another article.

For now, there are a few things worth keeping an eye on:

  • The number of messages rejected because of idempotency - if this number remains high or keeps increasing, it may indicate that some part of the system has entered a retry loop where operations repeatedly fail.
  • Circuit Breaker and health-check status - sometimes a provider doesn’t know that its service is unavailable, or there may be network issues between systems.
  • The number of messages in queues or Inbox and Outbox tables - if the number isn’t decreasing, it may indicate a failure or insufficient processing capacity.
  • Memory usage in reactive services or thread-pool utilization in synchronous communication.
  • CPU utilization - unusually low CPU usage under heavy load can indicate that the system is spending too much time waiting for synchronous responses, such as database or external-service responses. Very high CPU usage can point to excessive thread-pool size relative to the available CPU cores or inefficient memory-management configuration.

Maturity means being aware of imperfection

Some of you may be thinking: “I don’t use any of these patterns, and my system works just fine.”

And you may be right.

Most of the situations described here happen relatively infrequently. If you have 100 users a day and your system mainly provides information, the probability of a serious failure may be relatively low. In transactional, payment, or trading systems, however, where the volume and business impact of operations can be much higher, these failure scenarios become a much more important part of everyday risk.

That doesn’t change the main point: resilience should be considered when designing a system, not only after the first major incident.

Building bulletproof payment systems is about making conscious design decisions: breaking processes into atomic steps, explicitly modelling states, introducing safe retries, automating recovery processes, and building good observability. Only when we accept that failure is simply one of the normal scenarios our systems need to handle can we start talking about real architectural maturity.

Are you dealing with similar challenges in your projects? How do you maintain data consistency without distributed transactions? I’d be happy to discuss it. You can find me on LinkedIn.

Good luck building systems that can take a punch 🙂

18 września 2026, 09:45

My XTB Journey: Learning My Way Around Fintech

11 września 2026, 12:41

Gen Z at XTB: Beyond the Job Description

4 września 2026, 16:37

Digital Accessibility in Investing: What UX Research Can Teach Us About Building More Inclusive Products

27 sierpnia 2026, 21:36

My XTB Journey: Turning Ideas into Impact