# System Availability with Active-Active Architecture in Banking

> 5 Active-Active architecture designs using GKE gateways for the banking sector: pros, cons, obstacles, and recommendations from a real customer project.

Source: https://www.xalt.de/en/blog/system-availability-with-active-active-architecture/

TEAM XALT Atlassian Platinum Partner · 15 January 2025 · 13 min

An uninterrupted service is essential in the banking sector, where strict compliance requirements, security restrictions, and limited access to certain Google Cloud services lead to a complexity not found in less regulated environments. An Active-Active architecture provides a reliable solution for high availability and fault tolerance, but its implementation in Google Cloud for banks brings its own unique challenges. An effective Active-Active architecture can help overcome these challenges.

In this blog, we will examine five “Active-Active” designs that we implemented for a banking client, the obstacles we faced, and the lessons we learned. If you face similar challenges regarding high availability, these insights will help you shape your path more effectively.

In today's digital landscape, the need for an Active-Active architecture for banks is crucial to avoid downtime and strengthen customer retention.

Implementing an Active-Active architecture enables companies to synchronize their critical systems across multiple locations, thereby increasing resilience.

## Why an Active-Active Architecture in Banking?

In the banking sector, reliability, security, and compliance are non-negotiable – and if these challenges are not addressed, companies can be exposed to significant financial and reputational risks. An Active-Active architecture is not just an option but essential to ensure smooth operations and secure business continuity in an increasingly demanding landscape. The benefits of an Active-Active architecture are far-reaching and can be critical to a company's success.

Here are the reasons why it is important:

- **Continuous Availability**: Downtime is not only inconvenient for customers but also leads to financial losses and undermines trust. Active-Active setups eliminate single points of failure by distributing workloads across multiple regions, thereby ensuring uninterrupted service.
- **Fault Tolerance**: Failures, hardware errors, and disasters are inevitable, but by operating simultaneously in multiple regions, Active-Active architectures keep systems operational and mitigate the impact of such events.
- **Regulatory Compliance**: Strict regulations regarding data retention and disaster recovery must be adhered to, and Active-Active setups provide the cross-regional replication required to meet these requirements.
- **Low Latency**: Customers expect immediate responsiveness. By processing workloads closer to the user, Active-Active setups reduce latency and provide a superior user experience.
- **Customer Trust**: Reliability is the foundation of customer trust. Failing to meet expectations can irreparably damage relationships and brand reputation.

Implementing an Active-Active architecture is a strategic step that enables banks to optimise their operational processes and respond quickly to unexpected challenges.

While Google Cloud offers powerful tools for Active-Active architectures, the unique constraints in banking – such as restricted access to certain services – require tailored solutions. A well-planned Active-Active architecture can help minimise risks and enhance operational efficiency.

The risks of not implementing an Active-Active architecture are significant and can jeopardise a company's operational resilience.

### Why a Lack of Active-Active Architecture Could Harm Your Business

![Active-Active Architektur - Regional deployments](https://cdn.sanity.io/images/c475o02b/production/151d2f43e99fd9e2d0c79726b686e3211cde2b94-1112x899.jpg?w=1504&q=75&fit=max&auto=format)

In a non-active-active configuration, such as a regional deployment, a single region is responsible for the entire workload. While this approach is simpler and more cost-effective to implement, it has significant drawbacks. Availability is inherently limited, as any failure – whether due to hardware errors, natural disasters, or network issues – can lead to prolonged downtime. Without redundancy across regions, it is practically impossible to achieve a **99.99% availability target**, leaving critical systems unprotected and extending recovery times. This directly impacts customer trust, increases the risk of financial losses, and makes companies vulnerable in a highly competitive market.

## 5 Active-Active Architecture Designs

The Active-Active architecture for the banking sector is designed to ensure high availability, fault tolerance, and regulatory compliance. It utilises a combination of multi-regional deployments, synchronous and asynchronous replication strategies, and Split-Horizon DNS to efficiently manage both internal and external traffic. External clients are routed to public endpoints for seamless customer access, while internal clients are routed to private endpoints to optimise backend communication. This approach ensures a balance between performance, consistency, and failover capabilities, addressing the unique challenges of a highly regulated and security-critical environment.

The architecture for external traffic uses the **Google Kubernetes Engine (GKE) Multi-Cluster Gateway Service** of the Google Cloud Platform to ensure seamless, reliable, and secure access to customer-facing applications. This design provides a robust and scalable solution for managing external traffic across multiple active regions.

The internal data traffic architecture in an Active-Active configuration requires robust designs to ensure reliable service-to-service communication across multiple regions. The following section presents five approaches that leverage Google Kubernetes Engine (GKE) gateways and various load balancing methods to manage internal data traffic.

### 1. GKE Internal Global Multi-Cluster Gateway

**Overview**: A centralised internal gateway provided by GKE to manage service traffic across multiple clusters worldwide.

![Active-Active Architektur: GKE Internal Global Multi-Cluster Gateway](https://cdn.sanity.io/images/c475o02b/production/2873f57271076b0da7bd94f68870ceb00116bb8b-965x1035.jpg?w=1504&q=75&fit=max&auto=format)

- **Advantages:** Simplifies traffic management with a single control plane for routing across clusters.
- **Disadvantages:** At the time of writing, the GKE Internal Global Multi-Cluster Gateway had not yet been fully implemented in Google Cloud. According to the Google team, its release is planned for 2025.
- **Use case:** Suitable for organisations that prioritise simplicity and unified management in a global Active-Active configuration.

### 2. DNS-based global load balancing with regional GKE gateways

**Overview**: DNS resolves internal data traffic to regional GKE gateways and distributes requests to the nearest cluster.

![Active-Active Architektur: DNS-Based Global Load Balancing with Regional GKE Gateways](https://cdn.sanity.io/images/c475o02b/production/29c15ed104b1caea12cb71c915619ba1b5bb7d47-906x1103.jpg?w=1504&q=75&fit=max&auto=format)

Investing in a high-performance Active-Active architecture is of great importance for banks to remain competitive.

- **Advantages:** Offers flexibility in routing logic and combines well with Split-Horizon DNS for internal traffic resolution.
- **Challenges**: DNS changes depend on the TTL, which can lead to delays in failover during outages. The extended failover time undermines the concept of an Active-Active configuration, as traffic cannot be rerouted between regions quickly enough to ensure seamless operation.
- **Use case**: Ideal for workloads that require light traffic distribution with regional autonomy.

### 3. Cross-regional internal L7 load balancing with regional GKE gateways

![Active-Active Architektur: Cross-Regional Internal L7 Load Balancer with Regional GKE Gateways](https://cdn.sanity.io/images/c475o02b/production/03b408810ec134222abfb02301fb5389e8cb41e4-965x1035.jpg?w=1504&q=75&fit=max&auto=format)

**Overview**: A Layer 7 load balancer (application layer) manages traffic across regions and forwards requests to regional GKE gateways.

- **Advantages:** Provides advanced routing capabilities based on HTTP/HTTPS headers, paths, or other metadata.
- **Challenges**: Increased complexity and potential latency when routing traffic between regions. Furthermore, internal TLS certificate management in the customer environment was not compatible with cross-regional L7 load balancers.
- **Use case**: Best suited for microservice architectures with different routing requirements or complex application-level logic.

### 4. Cross-regional internal L4 load balancing with regional GKE gateways

**Overview**: A Layer 4 load balancer (transport layer) processes cross-regional traffic and distributes it to regional GKE gateways.

![Active-Active Architektur: Cross-Regional Internal L4 Load Balancer with Regional GKE Gateways](https://cdn.sanity.io/images/c475o02b/production/63dba99b4a08dd55adfefa434b56034bf6b27c39-965x1035.jpg?w=1504&q=75&fit=max&auto=format)

- **Advantages:** Simplifies network-level routing with lower latency compared to L7 load balancing.
- **Challenges:** Limited application-oriented routing features. The customer uses the RFC 6598 network segment (100.64.0.0/10) for its GCP cloud, but at the time of writing this document, GCP only supports RFC 1918 network segments (192.168.0.0/16, 10.0.0.0/8, 172.16.0.0/12).
- **Use case**: Effective for performance-critical applications where simplicity and speed are paramount.

### 5. Regional internal L4 load balancing with regional GKE gateways

**Overview**: Each region uses its own internal L4 load balancer to manage traffic forwarding between regions in an active-active configuration.

![Active-Active Architektur: Regional Internal L4 Load Balancer with Regional GKE Gateways](https://cdn.sanity.io/images/c475o02b/production/c93bb6821a66b973684060c25a1d51241613a9e2-965x1035.jpg?w=1504&q=75&fit=max&auto=format)

- **Advantages**: Low latency for regional communication while supporting cross-regional traffic as part of the active-active architecture.
- **Disadvantages**: Requires additional mechanisms for cross-regional failover and coordination
- **Challenge**: None – this design was successfully implemented and adapted to the environment's constraints.
- **Use case**: Ideal for scenarios requiring seamless intra- and inter-regional communication within a compliant and regulated environment.

## Key considerations for setting up an active-active architecture

Implementing an active-active architecture in a regulated environment such as banking requires careful planning and consideration of several critical factors. Here are the key considerations that influenced the design and implementation of the architecture:

#### 1. Compliance and security

- **Regulatory requirements**: Ensure that the architecture complies with industry regulations, including data residency, encryption standards, and disaster recovery mandates.
- **Network security**: Use firewalls, secure endpoints, and encryption (TLS) to protect data in transit and at rest.
- **Access control**: Implement role-based access control (RBAC) and enforce strict identity and access management (IAM) policies.

#### 2. Latency and performance

- **Proximity-based traffic routing**: Minimise latency by routing traffic to the nearest region using DNS or load balancers.
- **Cross-region communication**: Optimise data synchronisation between regions to balance latency and consistency.
- **Service response times**: Continuously monitor and optimise services to meet performance SLAs.

#### 3. Data consistency and synchronisation

- **Synchronous vs. asynchronous replication**: Choose the appropriate replication strategy based on data criticality and latency tolerance.
- **Conflict resolution**: Design mechanisms for handling conflicts in asynchronous replication scenarios.
- **Data partitioning**: Consider geographic segmentation of data to reduce synchronisation overhead.

#### 4. Scalability and fault tolerance

- **Load balancing**: Use a combination of global and regional load balancing modules to distribute traffic efficiently.
- **Auto-scaling**: Enable auto-scaling for GKE clusters and services to dynamically handle traffic spikes.
- **Failover mechanisms**: Implement robust failover processes to maintain service availability during regional outages.

#### 5. Infrastructure constraints

- **Service availability**: Take into account GCP services that may not be available or may be restricted in the banking environment.
- **Network configuration**: Align network segmentation with GCP-supported IP ranges (e.g., RFC-1918) and ensure compatibility with the architecture.
- **Resource quotas**: Monitor and manage cloud resource quotas to avoid capacity-related disruptions.

#### 6. Operational complexity

- **Deployment pipelines**: Automate deployment pipelines to ensure consistent infrastructure management across multiple regions.
- **Monitoring and observability**: Use centralized monitoring tools to gain insights into traffic patterns, service status, and anomalies.
- **Configuration management**: Ensure consistent configuration across all regions using tools such as Terraform or Config Connector.

#### 7. Cost optimisation

- **Traffic distribution**: Distribute traffic intelligently to minimise costs associated with cross-region data transfer.
- **Right-sizing resources**: Regularly review resource allocation to optimise cloud expenditure.
- **Billing transparency**: Leverage GCP billing tools to monitor and forecast costs associated with the active-active setup.

By considering these factors, the active-active architecture can meet the high standards of availability, performance, and compliance required in the banking sector.

### Final recommendations

Implementing an active-active architecture in a regulated and constrained environment such as the banking sector requires strategic decision-making and a clear understanding of the constraints. Based on our experience, the final recommendations are as follows:

1. **Assessing the need for an active-active configuration:**Not all applications require a **99.99% availability target**, and an active-active configuration is not always necessary. Many applications can function effectively with slightly lower availability targets, where a single-region configuration with robust failover and disaster recovery mechanisms can provide sufficient resilience.

  - For example, a **local corporate canteen web application** used by employees during business hours to view menus or order lunch does not need to be 99.99% available. In such cases, a single-region deployment with operating hours aligned to business times is both practical and cost-effective. This approach allows critical resources to be focused on systems that genuinely require high availability.
    - When deciding on an architecture, consider the criticality of the application, user expectations, and cost implications. For non-critical workloads, the additional complexity and effort of an active-active configuration may not be justified, allowing resources to be focused on other priorities. Always align the architecture with the specific availability requirements and business goals of the application.
2. **Understand the environment early**: Conduct a detailed assessment of compliance requirements, network constraints, and service availability to identify potential obstacles before designing the architecture.
3. **Prioritise simplicity where possible**: While multi- and cross-region setups are critical, an excess of technology can lead to unnecessary complexity. Opt for simpler solutions if they meet performance and compliance requirements. This approach aligns with the concept of **satisficing**, introduced by Herbert A. Simon in *Models of Man: Social and Rational*, which suggests finding a solution that meets adequacy criteria rather than pursuing an optimal but overly complex one.
4. **Leverage regional independence**: Strategically minimise cross-region dependencies where latency, performance, or compliance are paramount, while ensuring sufficient cross-region coordination to guarantee failover, consistency, and the benefits of an Active-Active configuration.
5. **Plan for scalability and resilience**: Prepare for traffic spikes and unexpected failures by using auto-scaling and robust failover mechanisms. Ensure that monitoring and observability are central to the architecture.
6. **Optimise costs carefully**: Balance performance requirements and cost efficiency by right-sizing resources and minimising cross-region data transfers. Regularly review and optimise cloud spend.
7. **Use a hybrid approach when necessary**: Combine multiple designs to meet different requirements, e.g., regional load balancing for local workloads and cross-region solutions for critical global services.

Implementing an Active-Active architecture enables banks to respond effectively to market changes while minimising risks.

### Conclusion

Developing an Active-Active architecture in Google Cloud for the banking sector is both challenging and rewarding. The constrained environment presented significant limitations but also led to innovative solutions tailored to the customer's specific needs. The successful implementation of the fifth design, utilising regional internal L4 load balancers, demonstrates the value of a thoughtful and adaptable approach.

By addressing key aspects such as compliance, performance, and scalability, this architecture provides the resilience and high availability required in the banking sector. The insights and experiences shared here can serve as a guide for other organisations facing similar challenges, enabling them to build robust and future-proof systems.

To remain competitive and relevant, banks must accelerate their transformation and continuously evolve their architectures to integrate new technologies and meet ever-evolving regulatory requirements. The time to act is now.

### Further reading

For those who wish to delve deeper into Active-Active architectures, cloud solutions, and the specific challenges of the banking sector, here are some recommended resources:

- Google architecture framework, the global deployment archetype: [Google Cloud global deployment archetype | Cloud Architecture Center](https://cloud.google.com/architecture/deployment-archetypes/global)
- Documentation on the GKE gateway: [About the Gateway API | GKE networking | Google Cloud](https://cloud.google.com/kubernetes-engine/docs/concepts/gateway-api)
- How to setup regional load balancers in GCP: [Set up a regional internal proxy Network Load Balancer with VM instance group backends | Load Balancing | Google Cloud](https://cloud.google.com/load-balancing/docs/tcp/set-up-int-tcp-proxy-migs)
- “Fundamentals of Software Architecture: An Engineering Approach” book by Mark Richards and Neal Ford

## Ready for high availability in your banking systems?

We help you design an Active-Active architecture in Google Cloud that balances compliance, security, and scalability.

[Talk to us](https://www.xalt.de/en/contact/)

## More articles

### [Greater Transparency in IT Asset Management: How IT Leaders Uncover Hidden Costs](https://www.xalt.de/en/blog/it-asset-management-transparency-hidden-costs/)

Most enterprises already have enough asset data – what they lack is visibility. How a lightweight tool turns administrative data into real decision-making insight, and uncovered nearly €50,000 a year in savings for one customer.

*Cloud*

### [Atlassian Cloud Migration: From Server and Data Center to the Cloud](https://www.xalt.de/en/blog/atlassian-cloud-migration-from-server-and-data-center-to-the-cloud/)

Why moving to Atlassian Cloud matters now, what the cloud offers and how a migration succeeds from assessment to go-live.

*Atlassian Cloud*

### [How to Use AI Safely and Efficiently in the Enterprise with Azure and OpenAI](https://www.xalt.de/en/blog/azure-openai-using-ai-safely-and-efficiently/)

Learn how to use artificial intelligence securely, compliantly, and efficiently within your organization using Azure and OpenAI – an overview.

*Cloud*

---

*This version is for AI agents. Every page of this site is available as Markdown: append `index.md` to its path. Index of all pages: [llms.txt](https://www.xalt.de/llms.txt)*
