A deployment went out at 3 PM Cape Town time. The bug showed up at 6 PM. The developer who wrote the code was asleep in London. The operator who found it couldn't roll back. The on-call engineer in Tbilisi had never seen the module. It took four hours to fix a fifteen-minute problem.
That incident taught us more about distributed teams than any process document ever did. Not because the team was bad — because the handoff was.
We build payment systems with teams spread across time zones. We do this because the alternative — hiring only within commuting distance of a single office — means building a worse team. The people who understand both SWIFT message formats and modern event architectures are not concentrated in one city.
But payment systems punish distributed teams in ways that a CMS or an e-commerce platform doesn't. Bank cut-off times are absolute. A reconciliation discrepancy at 5 PM can't wait until tomorrow. And when money stops moving, the person answering the phone from the operations floor doesn't care which timezone your developer is in.
Here's what we've learned, incident by incident.
In a co-located team, context moves through air. Overheard conversations, whiteboard scribbles, someone leaning over and saying "by the way, I changed how the fee calculation works." None of that exists when your team is distributed.
We lost a full day once because a developer in one timezone refactored the bank statement parser, pushed clean code with passing tests, and logged off. The next morning, a developer in another timezone pulled the changes and started building a feature on top of the old parser interface. They didn't check Slack. The PR description said "refactored parser" but didn't mention the interface change. Two people, eight hours of wasted work, zero malice.
Now, every developer writes a handoff note before logging off. Not optional. Not "if there's something important." Every day. Five lines:
Five minutes to write. Saves hours. We treat it as a deliverable — if your day doesn't end with a handoff, your day isn't done.
We had a team that wanted to deploy continuously. Small changes, fast feedback, modern engineering. Sounds right.
Except the client's banks process settlement files at 4 PM CET. The aggregator sends batch confirmations at 6 PM. The reconciliation runs at 7 PM. Deploy a change to the matching engine at 3:30 PM, and if something breaks, you've corrupted the day's reconciliation. And the recon person doesn't know it until 8 PM, when the developer who deployed is offline.
We learned to define deployment windows in the client's banking timezone. Not our timezone. Not "business hours." The specific hours when the payment pipeline is quiet. For most of our clients, that's 10 PM to 6 AM in the local banking jurisdiction, or Saturday mornings.
This means someone on the team is deploying at an odd hour in their local time. We rotate it fairly and compensate it properly. The alternative — deploying whenever it's convenient for the developer — is optimising for developer comfort at the expense of operational safety.
The "follow the sun" model for on-call sounds elegant. Whoever is in the active timezone handles incidents. Coverage follows daylight around the globe.
We tried it. It failed for a specific reason: the person on call had no context about what happened during the previous shift.
A payment batch was stuck. The on-call engineer in timezone B picked it up. They checked the logs, saw a timeout, retried the batch. It failed again. They escalated. Forty minutes later, someone reached the developer in timezone A who had deployed a config change two hours before the end of their shift. The config change hadn't propagated to all nodes. A two-minute fix took almost an hour because the on-call engineer was solving the wrong problem.
Now we require a two-hour overlap between on-call shifts. The outgoing on-call briefs the incoming one: what was deployed, what's in flight, anything unusual. Structured, not "check Slack." Fifteen minutes, specific to operations state.
For payment systems specifically, the on-call person needs more than SSH access and a runbook. They need:
If your on-call person needs permission to act, your incident response is as slow as your approval chain.
With three or more time zones, there is no meeting time that doesn't punish someone. 9 AM in Dubai is 6 AM in London. 2 PM in Cape Town is 10 PM in Auckland. Someone is always joining a call when they should be eating dinner or sleeping.
The instinct is to schedule more meetings to compensate for the distance. "We're not in the same room, so let's video-call more." This is exactly wrong. More meetings means more interruption, more timezone tax, and more time spent talking about work instead of doing it.
We restrict synchronous meetings to three cases:
Everything else is written. Decisions, code reviews, architecture discussions, status updates — all async. The written record is the bonus: six months from now, you can find why a decision was made. You can't do that with a meeting that nobody recorded.
A PR submitted at 5 PM in one timezone won't be reviewed until 9 AM in another. That's 16 hours. If the feedback requires changes, and the changes require another review, you're at 48 hours for a non-trivial PR.
This is fine if you plan for it. It's a disaster if you don't.
We structure sprint work so that no developer is blocked waiting for a review. Practically:
A few things we tried early on that we've since abandoned:
Daily standups across all time zones. Replaced with async written updates. Nobody misses the 7 AM video call.
Splitting the pipeline by team timezone. "London handles ingestion, Dubai handles execution, Tbilisi handles reconciliation." This created exactly the wrong boundaries — the person who understood how a transaction was ingested was in a different timezone from the person debugging why it didn't reconcile. We split by feature now, not by pipeline stage.
Assuming async means slow. Async done well is faster than sync done badly. A well-written Slack message with context, options, and a recommendation gets a decision in hours. A meeting to discuss the same thing takes 30 minutes of schedule coordination, 15 minutes of preamble, and lands on "let's take this offline."
The distributed model works. Not because we've eliminated the friction — because we've moved the friction to places where it's manageable and kept it away from places where it's catastrophic. A delayed code review is manageable friction. A delayed incident response is catastrophic. Design your process around that distinction.
Zenlime builds payment systems with distributed teams across multiple time zones. Start a conversation.