Architecture11 min

Online or asynchronous? How to choose how to integrate your systems

By Dorian Chávez · founder of Hábil and integration architect ·

A decision tree for leadership: online or asynchronous first; then REST or gRPC, or queue, callback or event, so the core does not go down.

Imagine a Thursday at month-end close. Your operations team uploads a file with 50,000 policy endorsements to adjust the portfolio's rates. The portal receives them and sends them, as fast as it can, to the core system. Within minutes the core starts to slow down. Then it stops responding. The agents who were issuing new policies at that very moment see frozen screens, and nobody knows which of the 50,000 endorsements were applied and which were not.

Nobody made a mistake in the code: a single way of integrating (the online call) was chosen, without saying so, for a job that called for another. This article is for whoever is responsible for operations or technology at a regulated company or a retailer: it comes down to three questions, in order.

Integrating is not "connecting systems." It is deciding how the business behaves when one of them slows down, goes down or receives ten times more work than usual.

First question: do you need the answer at that instant, or can it wait?

Ask it per process, not per company.

  • Online (synchronous): whoever asks stays waiting for the answer in order to continue. A quote on screen: the person cannot move forward without the price. It is a phone call.
  • Asynchronous: whoever asks leaves the work, receives an acknowledgment and carries on; the other system finishes at its own pace and notifies. A policy issuance or an invoice stamping: what matters is that it is done correctly and that you know what state it is in, not that it happens in the same second. It is a transaction at a service counter.

Four criteria:

  1. Is there a person looking at the screen, waiting? If so, and the answer is short, lean toward online.
  2. Does the work take longer than that person's patience? Long seconds, minutes or hours: asynchronous.
  3. Can the other system withstand the volume at its worst moment? If it saturates under a peak, an online call drags down with it the channel that depends on it. Asynchronous.
  4. Does it depend on a third party that may be slow or fail? (a bank, the tax authority, a provider). Asynchronous, and with a tracking number.

Rule of thumb: what a screen decides goes online; what moves money, documents or inventory tends to go asynchronous.

If it is online: REST or gRPC?

The decision is technical; the criterion is for leadership:

  • REST (the style of web APIs over HTTP, almost always with JSON) when compatibility with the callers (a partner, an app, a provider) matters more, or when anyone must be able to test and debug easily. It is the common language of the internet: almost every tool understands it. It can also have versioned contracts.
  • gRPC (a remote-call framework from Google that works from a strict contract and sends compact binary messages over HTTP/2) when a strict contract, clients generated from it, continuous streaming or measured performance justify having the callers support it; that is why it tends to appear between your own systems, where both ends speak it.

A frequent guideline, not a rule: REST outward, gRPC between your own systems. The detail is in the article "REST or gRPC: when to use each, and why your core speaks neither" (/blog/rest-or-grpc-when-to-use-each).

If it is asynchronous: which flavor?

For this decision it helps to distinguish six frequent options, which can be combined; the enterprise messaging literature (Hohpe and Woolf) and the documentation of the large cloud providers [1][2] describe many more. None is "the good one": each solves a different problem.

Callback: "we'll notify you with your reference number"

Use it when the work takes seconds or minutes and whoever asked can receive notifications. The receiving system answers immediately "received, your reference is 4821" and gets to work. When it finishes, it notifies back with the reference. Technically, the result comes back through a callback; if it is delivered over HTTP to a registered address, it is usually called a webhook.

The reference number is the tracking number that links each response to its request (correlation identifier [3]). Without it, a notification is a piece of paper with no owner.

Status polling: "come back with your reference number and ask"

Use it when whoever asked cannot receive notifications (firewall, mobile app, third parties with restrictions). You come back with your reference number and ask "is it ready?" every so often.

Microsoft's documentation describes it this way: the initial response is an "accepted" (HTTP code 202) that indicates where to check, and the query returns states such as pending, in progress, succeeded or failed [4]. It works where the notification does not arrive, and it proposes requiring a key per request so that it is not processed twice [4].

Queue: the line of jobs

Use it when the receiving system has a pace limit and whoever sends can produce much more. The sender leaves its work and withdraws; the processor picks it up at its own pace. It is, above all, a buffer: Microsoft's documentation calls it Queue-Based Load Leveling and describes it as a buffer between whoever asks and the service, which smooths intermittent loads that could make the service fail [5]. For more speed, more workers are added on the same line (Competing Consumers), if the jobs are independent [6].

Scenario (illustrative figures, no client). A portal receives 50,000 endorsements; the core handles 20 per minute. Without a buffer, the portal pushes everything at once and the core saturates. At 20 per minute, the theoretical time is about 42 hours (50,000 ÷ 20 = 2,500 minutes); the real time depends on validations, retries and operating windows. The queue regulates the intake; the reference per batch, the progress dashboard and the alerts are designed separately, and in return there is no core down at peak hour.

The queue does not make the core faster: it makes its limit stop being a surprise. It is not suitable if an immediate response is needed or if the volume is low and steady [5].

Event or topic: the loudspeaker announcement

Use it when a single fact must move several things at once: documents, customer notification, CRM, fraud. The previous ones go from one to one; the event, from one to many. Instead of telling one system "do this," whoever lives the fact announces it: "policy issued," "payment applied." Everyone subscribed receives a copy, and whoever emits it does not know how many there are. In the messaging literature, it is the difference between a point-to-point channel (one receiver) and a publish-subscribe one (all the interested parties) [7].

Its cost: everything is eventually consistent; for a while, some systems know the news and others do not. It is not suitable if a single atomic transaction is needed between sender and receivers [8].

High-volume event streaming: the tape that can be rewound

Use it when events are very many, continuous, and it matters to keep them so they can be re-read. Apache Kafka defines it as capturing events in real time, storing them durably to query them later, and processing them both as they happen and retrospectively [9]. The image: a tape of facts that several teams read at their own pace, for as long as the configured retention period lasts [9].

It makes sense for telemetry, inventory or fraud; not for a one-off order, since it is machinery that someone must operate. According to each provider's documentation, the guarantees on delivery, ordering, duplicates and replay vary by product and configuration [10][11].

Exchange through tables or files: when the other system can accept nothing more

Use it when the destination system is an old core that speaks none of the previous languages. A system like that usually has four limitations:

  • It exposes no interfaces for others to ask it for things, and it also does not notify you when something changes inside it.
  • It only accepts data in an exchange zone: agreed tables or files, where the information is left.
  • It processes them at its own pace, often in batches, triggering a process of its own that picks them up, runs them and writes the result.
  • It can take little load, and the response arrives when that process finishes, not when it is requested.

The exchange zone is a fixed-format service counter: each instruction is left there with its reference number, and the core leaves the result in the same place. Its merit: nothing is written into the core's internal tables. Its cost: someone must watch the zone and the response takes as long as the batch takes.

On top of that zone goes an anti-corruption layer: a translator from modern contracts to the fields and paces the core understands [12]. That is how you modernize in stages (Strangler Fig) while the core keeps operating [13].

Received is not processed

If you take away just one idea from this article, let it be this one: an acknowledgment of receipt is not a confirmation of outcome. It is often believed that, because the acknowledgment arrived, what was sent has been resolved.

There are three levels of response, and it helps to know which one each of your processes is at:

Received is not processed
LevelWhat it saysSPEI payment (Mexico's interbank transfer system)Invoice (CFDI, Mexico's digital tax invoice)Order
Technical acknowledgment"It arrived"The app confirms it sent the instructionThe stamping provider acknowledges receipt of the XML"Order received"
Functional acknowledgment"It is well formed"Account and amount with a valid formatThe XML meets the structureComplete data and the product exists
Business confirmation"It is done," or "rejected, and this is the reason"Settled status and, with the credit confirmed, Banxico's CEP (the proof-of-payment receipt, Comprobante Electrónico de Pago) [16]Stamp with UUID and SAT sealDelivered, or rejected with a reason

Only the third level says whether the invoice exists, the payment arrived or the order will be fulfilled. Banxico illustrates it: settled means SPEI settled and notified the receiving institution, and the CEP records the credit in the beneficiary account [16].

The industry already separates receiving from processing. HTTP defines code 202 as "accepted for processing, but the processing has not been completed"; the request might never be carried out, and the protocol has no way to resend the status later [17]. In AS2, the standard for exchanging documents between companies, the official example of a signed receipt carries this comment: "This is not a guarantee that the message has been completely processed or understood by the receiving translator" [18].

What we saw. In an order integration between two corporate systems that we reviewed, out of about eight calls, two returned the identifier of the document created; the other six, only an acknowledgment. The errors arrived by email from the receiving area, about three days later. In the same batch there were more than ten identical cases that nobody had seen: only the ones someone happens to check are discovered, and the rest are indistinguishable from the ones that went well. It is a measurement from a single case, not an industry figure. The cheapest thing that worked: checking that the reference that comes back corresponds to the right record. Across 234 records, it did not produce a single false alarm.

What to ask for, by flavor:

  • In all of them: a reference per operation that travels both ways, to link each result with its request [3].
  • Callback: that it notifies the result; a "finished" that does not say whether it was applied is just another acknowledgment.
  • Status polling: a clear status per reference (pending, in progress, succeeded or failed) and, if it failed, the reason [4].
  • Queue: that messages that fail are set aside for review, with their origin noted [15].
  • Tables or files: one result line for each record sent.

And two more defenses. An explicit state for what has not come back ("unconfirmed since 10:40"): the worst thing is not the error, it is not knowing. And a periodic reconciliation that compares what was sent against what was confirmed, with an alarm for what did not come back in time; that way silence becomes a task with an owner.

Sector problem, first question, flavor

A conversation guide, not a recipe: combining flavors is the norm.

Sector problem, first question, flavor
Sector problemFirst question: online or asynchronous?Flavor that tends to fit
Policy issuanceThe quote, online; the issuance can waitQuote online; issuance with a reference and callback; "policy issued" as an event for documents, collections and CRM
Mass endorsements with a limited coreAsynchronous: the core sets the paceQueue with a controlled pace, or exchange through tables or files toward the core; reference per batch and progress dashboard
CFDI stamping (when integrating with a certification provider; with the SAT's free tool, confirm its capabilities and limits first [19])Asynchronous: it depends on a third party that may be slow or failQueue with one command per invoice; result by callback or polling; exceptions set aside; a retry must not stamp twice
PaymentsMinimum validations online; the rest asynchronousInstruction with a reference; explicit states (received, sent, accepted, rejected); reconciliation separately
Omnichannel inventoryThe availability query online; movements, asynchronousEvents for reservation, release and movement; continuous stream for visibility; reservation rule with a single owner, so as not to sell twice
ClaimsIntake with a reference is immediate; the rest, asynchronousReference instantly; documents, estimate, fraud and assignment react to events

The risks, in business language

An asynchronous flavor does not eliminate problems: it moves them. Five for the table.

  • "It was charged twice." Most queues deliver each message at least once: it can arrive repeated. Microsoft asks that processing it several times produce the same result, to avoid "duplicate records or repeated charges" [5]: each instruction carries a unique key and the system records which ones it has already processed.
  • Order. When several workers take from the same line, there is no guarantee that jobs finish in the order they came in. If "cancel the policy" arrives before "issue the policy," you have a problem [5][6]. Define the ordering key for your business (policy, payment, claim) and respect it where it matters.
  • Retries without brakes. A transient error deserves another attempt; a functional rejection ("account does not exist") is not fixed by repeating it. The recommendation is to set it aside in an exception queue and monitor it [5][6].
  • Not knowing what state something is in. It is the risk from the previous section; with events, it also helps to have an identifier that follows each operation from end to end [8].
  • Each party sees a different version, for a while [8]. Decide in advance who owns the final data and how a half-completed operation is reconciled.

The format of notifications is an agreement that gets versioned; CloudEvents requires that source plus id be unique per event, which helps detect resends [14].

Five questions for your next committee

  1. Which decisions demand an immediate answer and which can be resolved with a reference number?
  2. What is the maximum pace our core sustains, and what happens if ten times that volume arrives?
  3. Who owns the final state of each operation, and how is one left incomplete reconciled?
  4. How do we verify that a retry does not duplicate a payment, a policy, a CFDI or a movement?
  5. Of what we sent yesterday, which operations have a confirmation of outcome (applied or rejected) and which only an acknowledgment of receipt?

If in two or more the answer is "I don't know," there may be your risk.

Where to start this week

Write down your most important processes (issuance, collection, stamping, inventory), which system answers each one and how much work per minute it can take; and whether whoever starts it needs the answer now or a reference number would be enough. You will see where an online call is being used for a job that called for a line.

Then take one interface and count how many of the operations sent yesterday have a confirmation of outcome, not just an acknowledgment.

What Hábil does

Hábil is a Mexican engineering consultancy, founded in 2006, that works between the systems a company already has and the channels it needs to open. See the page Unify your operations and your channels.

The first conversation is at no cost: it starts from a map of your current integrations and of where your operation could lose money, time or evidence. Write to us on WhatsApp.

Descriptions of vendor services reflect their documentation as of October 6, 2026; they are the vendor's attributions, not our own measurements.

References

  1. Gregor Hohpe and Bobby Woolf, Enterprise Integration Patterns (Addison-Wesley, 2003) and online catalog. https://www.enterpriseintegrationpatterns.com/
  2. Microsoft, Azure Architecture Center, cloud design patterns catalog. https://learn.microsoft.com/en-us/azure/architecture/patterns/
  3. Enterprise Integration Patterns, Correlation Identifier. https://www.enterpriseintegrationpatterns.com/patterns/messaging/CorrelationIdentifier.html
  4. Microsoft, Asynchronous Request-Reply pattern (202 response with status location, states, idempotency key). https://learn.microsoft.com/en-us/azure/architecture/patterns/asynchronous-request-reply
  5. Microsoft, Queue-Based Load Leveling pattern (buffer between task and service; at-least-once delivery; ordering; exception queue; when not to use it). https://learn.microsoft.com/en-us/azure/architecture/patterns/queue-based-load-leveling
  6. Microsoft, Competing Consumers pattern. https://learn.microsoft.com/en-us/azure/architecture/patterns/competing-consumers
  7. Enterprise Integration Patterns, Messaging Channels (point-to-point and publish-subscribe channel). https://www.enterpriseintegrationpatterns.com/patterns/messaging/MessagingChannelsIntro.html
  8. Microsoft, Publisher-Subscriber pattern (asynchronous and eventually consistent; ordering; correlation; when not to use it). https://learn.microsoft.com/azure/architecture/patterns/publisher-subscriber
  9. Apache Kafka, Introduction (definition of event streaming). https://kafka.apache.org/intro
  10. Microsoft, Compare messaging services (Event Grid, Event Hubs, Service Bus). https://learn.microsoft.com/en-us/azure/service-bus-messaging/compare-messaging-services
  11. Amazon Web Services, Fanout Amazon SNS notifications to Amazon SQS queues for asynchronous processing. https://docs.aws.amazon.com/sns/latest/dg/sns-sqs-as-subscriber.html
  12. Microsoft, Anti-Corruption Layer pattern. https://learn.microsoft.com/en-us/azure/architecture/patterns/anti-corruption-layer
  13. Microsoft, Strangler Fig pattern. https://learn.microsoft.com/en-us/azure/architecture/patterns/strangler-fig
  14. CloudEvents, Specification v1.0 (source + id). https://github.com/cloudevents/spec/blob/main/cloudevents/spec.md
  15. Enterprise Integration Patterns, Dead Letter Channel (a message that cannot or should not be delivered is set aside on a separate channel, with its original channel recorded). https://www.enterpriseintegrationpatterns.com/patterns/messaging/DeadLetterChannel.html
  16. Banco de México, MI SPEI: transferencias ("Liquidado" status; CEP as proof of the credit). https://www.banxico.org.mx/servicios/mi-spei_-transferencias-ban.html
  17. IETF, RFC 9110 HTTP Semantics, §15.3.3 "202 Accepted". https://www.rfc-editor.org/rfc/rfc9110.html#section-15.3.3
  18. IETF, RFC 4130 MIME-Based Secure Peer-to-Peer Business Data Interchange Using HTTP (AS2), signed receipt (MDN) example. https://www.rfc-editor.org/rfc/rfc4130.html
  19. SAT (Mexico's tax authority), Resolución Miscelánea Fiscal 2026, rule 2.7.1.6 (issuing a CFDI without sending it to a certification provider through "Genera tu factura" or "Factura SAT Móvil"), based on art. 29 of the CFF (Federal Tax Code). https://www.sat.gob.mx/minisitio/NormatividadRMFyRGCE/documentos2026/rmf/rmf/RMF_2026-DOF-28122025.pdf