12. Architecting Hybrid Container-Based Workloads

Author

Senthil Kumar

👈 Back to: 📝 Blog | 💼 LinkedIn | ✍️ Medium


How to Architect a Hybrid, Container-Based Solution on AWS

  • ❓ Key Question of this chapter: When half your workloads stay on-premises and half move to AWS, how do you connect the two, keep containers portable, migrate a database with near-zero downtime, and stay resilient without rewriting applications?
  • Workflow followed in all chapters: requirements → architectural drivers → candidate services → trade-offs → decision → justified rejections.
  • What changes is that the answer now spans two environments joined by a network.

Running example used throughout this chapter

An enterprise runs containerised applications and PostgreSQL databases in its own data centres. Data-centre contracts are expiring in waves, so it will move half the workloads to AWS now and keep the rest on-premises until their contracts end. This is a genuinely hybrid end state, not a big-bang migration.

The customer’s constraints:

  • Dedicated, consistent, low-latency connectivity between the data centre and AWS for high-volume traffic.
  • Keep containers private with no inbound internet, but allow outbound access, for example to pull updates.
  • Same orchestration tooling on both sides to simplify operations.
  • Lift-and-shift PostgreSQL to AWS with no application code rewrite.
  • High uptime, resilience and fault tolerance, with no single points of failure.
  • On-prem apps must store and read data in AWS over NFS without refactoring.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    subgraph DCn["Data centre (contracts<br>expiring in waves)"]
        C1["Containerised apps"]
        DB1["PostgreSQL"]
    end
    DCn <-->|"dedicated, low-latency link"| AWS

    subgraph AWS["AWS (half the workloads)"]
        C2["Containers<br>private, outbound only"]
        DB2["PostgreSQL<br>lift-and-shift, no rewrite"]
        NFS["Data over NFS<br>no refactor"]
    end

    C1 -.->|"same orchestration tooling"| C2
    DB1 -.->|"no app code rewrite"| DB2


12.1 The Method: nine steps across two environments

A hybrid design decomposes into the same reasoning chain as a cloud-only one, but every answer must hold in both environments and across the network underneath them.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    N["1. Connect<br>DC to AWS"] --> C["2. Run containers<br>privately"]
    C --> D["3. Migrate the<br>database"]
    D --> S["4. Share storage<br>across the boundary"]
    S --> O["5. Operate, observe,<br>scale, protect"]

Step Question the architect answers Winning choice (this case) Section
1 What are the real architectural drivers? — #sec-drivers
2 Data Center (DC) ↔︎ Cloud: How do the two environments connect? AWS Direct Connect with VPN failover #sec-connect
3 Common Compute: What runs the containers privately? ECS on EC2 launch type + NAT per AZ #sec-compute
4 Cloud DB: How does PostgreSQL move resiliently? RDS Multi-AZ + DMS, read replicas #sec-db
5 Data: Where does shared data live? S3 File Gateway + S3 #sec-storage
6 Orchestration: How is it operated the same way on both sides? ECS Anywhere, Systems Manager, AWS Backup #sec-operate
7 Monitoring: How do we know it is healthy? CloudWatch across DC and AWS #sec-observability
8 Scaling: What scales, and how? Cluster + service auto scaling #sec-scale
9 Disaster Recovery: How much disaster recovery do we buy? Cheapest tier that meets RTO/RPO #sec-dr

Architect’s takeaway: In hybrid design, the network is the foundation. Get connectivity wrong and every other decision inherits its latency, cost and failure modes.

Note - What is RTO and RPO metrics in DR measures:

  • RTO = Recovery Time Objective = How long can the system remain unavailable?
  • RPO = Recovery Point Objective = How much recent data can be lost?

12.2 Turn requirements into drivers

  • ❓ How does “half here, half there, for years” change the design?

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    R1["Half on-prem, half AWS,<br>for a long time"] --> D1["True hybrid, not big-bang"]
    R2["High-volume,<br>consistent traffic"] --> D2["Dedicated connectivity"]
    R3["Private containers,<br>outbound only"] --> D3["Private subnets + NAT"]
    R4["Same tooling both sides"] --> D4["Consistent orchestration"]
    R5["Lift-and-shift PostgreSQL,<br>no rewrite"] --> D5["Managed compatible DB"]
    R6["High uptime,<br>no SPOF"] --> D6["Multi-AZ + redundant links"]
    R7["NFS access to AWS data"] --> D7["Hybrid file storage"]

Customer requirement Architectural driver Implication
Half on-prem for years Hybrid, not migration Invest in durable connectivity and shared tooling
High-volume, consistent traffic Dedicated connectivity Direct Connect, not internet VPN as the primary
Private containers, outbound only Controlled egress Private subnets + NAT / VPC endpoints
Same tooling both sides Operational consistency ECS Anywhere + Systems Manager agents on-prem
No application rewrite Compatibility over optimality Managed engine that speaks the same wire protocol
No single points of failure Redundancy at every layer Multi-AZ, multi-link, multi-NAT
Keep NFS Protocol preservation Hybrid file gateway, not a client refactor

⚠️ Plan the address space before anything else. The most common hybrid blocker is overlapping CIDR ranges between the data centre and the VPC. Routing cannot resolve the overlap, and the fix, re-IP or NAT the overlap, is painful once workloads are live. Agree non-overlapping ranges at design time.

Architect’s takeaway: “Half and half, for a long time” is the driver that shapes everything. It rules out a big-bang migration and forces you to invest in durable connectivity and consistent tooling rather than treating on-prem as throwaway.


12.3 Connect the two environments

  • ❓ Dedicated line, encrypted tunnel, or a central hub, and when does each win?

12.3.1 The connectivity decision

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    Q0["Connect DC to AWS"] --> Q1{"Need consistent,<br>dedicated, high throughput?"}
    Q1 -->|Yes| DX["AWS Direct Connect - CHOSEN<br>private, off the public internet"]
    Q1 -->|"No / quick + cheap"| VPN["Site-to-Site VPN<br>IPsec over internet"]
    DX --> Q2{"Many VPCs / accounts /<br>Regions to interconnect?"}
    Q2 -->|Yes| TGW["Add Transit Gateway<br>+ Direct Connect Gateway"]
    Q2 -->|No| DONE["Direct Connect<br>+ VPN failover"]

Chosen: AWS Direct Connect. The customer needs consistent, dedicated throughput for high-volume DC↔︎AWS traffic. Direct Connect keeps traffic off the public internet, which removes the latency spikes and bottlenecks associated with internet-based connectivity.

🔴 Private is not the same as encrypted. Direct Connect traffic is not encrypted by default. It is a private circuit, not an encrypted one. If the compliance posture requires encryption in transit across the link, add MACsec on supported dedicated connections for layer-2 encryption, or run an IPsec Site-to-Site VPN over the Direct Connect public virtual interface. Do not let “off the public internet” be read as “encrypted.”

⏱️ Direct Connect has a lead time, and here it is on the critical path. A dedicated connection requires a cross-connect at a Direct Connect location, including LOA-CFA, partner coordination and physical patching, and typically takes weeks to months. Since this customer’s timeline is driven by expiring data-centre contracts, order the circuit early, and consider a hosted connection from an AWS Partner, provisioned in less time and available in sub-1 Gbps increments, to bridge the gap.

12.3.2 Direct Connect topology

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    DC["Data Center"] <--> DX1["Direct Connect<br>Location A"]
    BO["Branch Office"] <--> DX2["Direct Connect<br>Location B"]
    DX1 <-->|"SiteLink between<br>DX locations"| DX2
    DX1 --> R["Any AWS Region"]
    DX2 --> R

  • SiteLink lets you send data between Direct Connect locations, creating private links across global offices and data centres.
  • A Direct Connect Gateway lets one connection reach multiple VPCs across accounts and Regions. Without it, a private virtual interface maps to a single VPC.
  • Each dedicated connection is a single link, so AWS recommends a second connection for redundancy. Without a backup, either a second Direct Connect connection or an IPsec VPN, VPC traffic is dropped on failure.

12.3.3 The alternatives and when they would win

Site-to-Site VPN (virtual private gateway + customer gateway, with redundancy)

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    subgraph AWS["Amazon VPC"]
        VGW["Virtual Private Gateway<br>AWS-side concentrator"]
        E1["EC2 in AZ 1"]
        E2["EC2 in AZ 2"]
        VGW --- E1
        VGW --- E2
    end
    subgraph DCn["Customer Network"]
        CGW1["Customer Gateway 1"]
        CGW2["Customer Gateway 2"]
        SRV["Servers"]
        CGW1 --- SRV
        CGW2 --- SRV
    end
    VGW <-->|"IPsec VPN"| CGW1
    VGW <-->|"IPsec VPN redundant"| CGW2

Option How it works Prefer when
Direct Connect (chosen) Private, dedicated fibre to AWS, off the public internet You need consistent throughput and latency for high volume
Site-to-Site VPN IPsec tunnel over the internet through a virtual private gateway You want it fast and cheap, or as failover for Direct Connect
Client VPN Managed remote-access VPN for individual users A remote workforce needs to reach VPC or on-prem resources

⚠️ VPN failover is survival, not capacity. Each Site-to-Site VPN tunnel is capped at roughly 1.25 Gbps. Failing a 10 Gbps Direct Connect connection over to a VPN is a capacity cliff, not a transparent switch. Expect degraded service. For genuinely high-bandwidth production, AWS’s resiliency models recommend two Direct Connect connections at two separate locations, with the VPN as a third line of defence.

Client VPN covers three access shapes worth remembering:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    CE["Client VPN endpoint"] --> S1["Scenario 1<br>access a single VPC"]
    CE --> S2["Scenario 2<br>access on-prem only"]
    CE --> S3["Scenario 3<br>access VPC + client-to-client"]

12.3.4 Scaling connectivity: AWS Transit Gateway

Direct Connect answers one DC to one VPC. Once you have many VPCs, accounts or Regions, point-to-point peering explodes into an unmanageable mesh. Transit Gateway is the hub-and-spoke fix.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    subgraph Without["Without Transit<br>Gateway - mesh"]
        A1["VPC A"] --- A2["VPC B"]
        A2 --- A3["VPC C"]
        A1 --- A3
        A1 --- CG1["Customer GW"]
        A2 --- CG1
        A3 --- CG1
    end

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    subgraph With["With Transit Gateway - hub"]
        TGW(("AWS Transit<br>Gateway"))
        V1["VPC A"] --- TGW
        V2["VPC B"] --- TGW
        V3["VPC C"] --- TGW
        CG["Customer GW / VPN"] --- TGW
        DXG["Direct Connect Gateway"] --- TGW
    end

  • Each connection is made once to the hub, and route tables control which networks can communicate.
  • Inter-Region peering links transit gateways over the AWS backbone. That traffic is automatically encrypted and does not traverse the public internet, unlike the Direct Connect link itself.
  • Transit Gateway pairs naturally with a Direct Connect Gateway for multi-Region and multi-VPC estates.

For VPC, subnet and routing fundamentals, see 04. Networking.

Architect’s takeaway: Choose Direct Connect for consistent, high-volume hybrid traffic, but order it early, encrypt it explicitly, and size the fallback honestly. Keep a VPN as failover so a single link is never a SPOF (single point of failure), and reach for Transit Gateway + Direct Connect Gateway when the topology grows beyond a couple of VPCs.


12.4 Run the containers, privately

  • ❓ ECS or EKS? EC2 or Fargate? And how do private containers still reach the internet outbound?

12.4.1 Orchestrator decision: ECS or EKS?

This customer’s containers are plain Docker on EC2, so ECS, which offers an AWS-native model with less to operate and no separate control-plane concept to learn, is the right orchestrator. EKS earns its keep in a different, common hybrid scenario: an existing on-premises Kubernetes cluster.

If the workload you’re migrating is already a hybrid Kubernetes cluster, and the ask is to move part of it to the cloud with minimum effort, keep native Kubernetes APIs and tooling, and cut operational overhead for running the control plane, the answer is Amazon EKS, not ECS. EKS runs the same Kubernetes you already operate, so manifests, Helm charts and kubectl muscle memory carry over unchanged, while AWS manages the control plane’s availability and scaling. Fargate with ECS is the wrong fit here because it does not run Kubernetes and therefore fails the “keep native Kubernetes features” requirement outright.

Starting point Orchestrator Why
Plain Docker containers, no Kubernetes investment ECS (this chapter’s case) Simpler AWS-native model, less to learn, fits the EC2 launch-type requirement below
Existing on-prem Kubernetes cluster, hybrid migration EKS Preserves native Kubernetes APIs and tooling; AWS manages the control plane

12.4.2 Launch-type decision: why EC2, not Fargate

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    W["Container hosting on AWS"] --> Q1{"Need a custom AMI?"}
    Q1 -->|No| Q2{"Need SSH to the<br>underlying hosts?"}
    Q1 -->|Yes| EC2["EC2 launch type - CHOSEN"]
    Q2 -->|Yes| EC2
    Q2 -->|No| FAR["AWS Fargate<br>serverless, no host access"]

Chosen: Amazon ECS with the EC2 launch type. The customer requires a custom AMI and SSH access to the underlying instances to keep management operations similar across on-prem and cloud. Fargate supports neither, so EC2 is the correct launch type here, a case where the serverless option is rejected on purpose.

ECS on EC2 architecture

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    IMG["Container image"] --> REG["Container registry<br>ECR or Docker Hub"]
    TD["Task definition"] --> CL
    SVC["Service<br>task def + service description"] --> CL
    subgraph CL["ECS cluster in VPC"]
        subgraph AZ1["AZ 1"]
            I1["EC2 instance<br>ECS agent + tasks"]
        end
        subgraph AZ2["AZ 2"]
            I2["EC2 instance<br>ECS agent + tasks"]
        end
    end
    REG --> I1
    REG --> I2

  • The ECS agent runs on each EC2 instance and lets ECS orchestrate the nodes you own and manage.
  • In contrast, with Fargate, each task gets its own elastic network interface and there are no EC2 instances to manage. This is valuable when you do not need a custom AMI or host access.
EC2 launch type (chosen) Fargate launch type
You manage The EC2 cluster, including AMI, patching and SSH Nothing below the task
Custom AMI ✅ Yes ❌ No
SSH to host ✅ Yes ❌ No
Best when Custom AMI, host control, consistent tooling “Just run my container,” no infrastructure operations

See 03. Compute for the wider EC2 and container landscape.

12.4.3 Private subnets, outbound egress, and the NAT trap

The containers must be private with no inbound internet, but need outbound connectivity for updates. A NAT gateway provides exactly that.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    subgraph AZ1["AZ 1"]
        E1["ECS tasks<br>private subnet"] --> N1["NAT gateway 1<br>public subnet"]
    end
    subgraph AZ2["AZ 2"]
        E2["ECS tasks<br>private subnet"] --> N2["NAT gateway 2<br>public subnet"]
    end
    N1 --> IGW["Internet gateway"]
    N2 --> IGW
    IGW --> NET["Internet<br>downloads and updates"]
    NET -.->|"no unsolicited inbound"| E1

  • A NAT device replaces the private instance’s source IP with its own and translates responses back, so instances initiate outbound connections but cannot receive unsolicited inbound ones.
  • Because a VPC gives full control over the virtual network, security groups, which are instance-level and stateful, and network ACLs, which are subnet-level and stateless, are the firewall rules that enforce “private, outbound-only.” The NAT gateway handles address translation rather than access control.
  • A managed NAT gateway is recommended over a self-managed NAT instance for better availability, bandwidth and lower administration.
  • A public NAT gateway uses an Elastic IP and routes to an internet gateway. A private NAT gateway routes to other VPCs or on-premises networks through a Transit Gateway or Virtual Private Gateway and does not use an Elastic IP.

🔴 One NAT gateway is a single point of failure. A NAT gateway is scoped to one Availability Zone. If you put one NAT gateway in one public subnet and route every private subnet to it, an AZ failure takes all container egress down. That directly contradicts the customer’s “no single points of failure” requirement. Deploy one NAT gateway per AZ and give each AZ’s private subnet its own route table pointing to its local NAT gateway. This also avoids cross-AZ data transfer on every outbound byte.

💰 Don’t pay NAT charges for traffic that never needed the internet. NAT gateways bill per hour and per GB processed. ECS image pulls from ECR, the SSM Agent, CloudWatch Logs and S3 access go through NAT by default. Replace those paths with VPC endpoints: a gateway endpoint for S3, which has no hourly charge, and interface endpoints for ECR API/DKR, SSM and CloudWatch Logs. This reduces cost and keeps AWS-service traffic off the internet path.

Architect’s takeaway: “Serverless by default” is not a law. When the driver is a custom AMI + host SSH + consistent tooling, ECS on EC2 is the right call. Pair private subnets with one managed NAT gateway per AZ for resilient outbound-only access, and use VPC endpoints for AWS-service traffic.


12.5 Migrate the database, resiliently

  • ❓ How do you lift-and-shift PostgreSQL with near-zero downtime and no single point of failure?

Chosen: Amazon RDS for PostgreSQL, Multi-AZ, migrated with AWS DMS. A managed engine keeps the app code unchanged, Multi-AZ delivers the high availability the customer demands, and DMS moves the data while the source stays live.

For RDS fundamentals, backups and where databases sit in your VPC, see 06. Databases.

12.5.1 High availability: Multi-AZ DB instance with one standby

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    APP["Application"] --> PRI["Primary DB<br>AZ 1"]
    PRI -->|"synchronous<br>replication"| STB["Standby DB<br>AZ 2"]
    PRI -.->|"auto-failover on failure<br>same endpoint"| STB

  • RDS keeps a synchronous standby in another AZ and fails over automatically. The database endpoint does not change, so applications reconnect transparently.
  • The standby is not readable. It exists purely for availability.
  • A single RDS instance does not self-heal. High availability requires Multi-AZ.

12.5.2 Even higher availability: Multi-AZ DB cluster with two readable standbys

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    W["Writer endpoint"] --> P["Primary writer<br>AZ 1"]
    P -->|"replicate"| S1["Readable standby<br>AZ 2"]
    P -->|"replicate"| S2["Readable standby<br>AZ 3"]
    R["Reader endpoint"] --> S1
    R --> S2

  • This is a different deployment type from the Multi-AZ DB instance above. In the console and documentation it is the Multi-AZ DB cluster, available for PostgreSQL and MySQL.
  • It spans three AZs and includes two readable standbys, with failover typically under approximately 35 seconds, up to approximately 2× faster commit latency, and additional read capacity.

12.5.3 Scaling reads: read replicas

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    APPW["App servers<br>read and write"] -->|"read + write"| PRI["Primary DB"]
    PRI -->|"asynchronous<br>replication"| RR["Read replica<br>read-only"]
    BI["BI / reporting<br>app server"] -->|"read-only queries"| RR

  • Read replicas offload read-heavy and CPU-intensive work, such as BI reports, from the primary through asynchronous replication, raising aggregate read throughput.
  • Replicas can be promoted to standalone instances if needed.

Don’t confuse the three: Multi-AZ standby = availability, with synchronous replication, automatic failover and a non-readable standby. Multi-AZ DB cluster = availability + some read capacity. Read replica = performance and scale, using asynchronous replication with possible replica lag. They solve different problems and are often used together.

12.5.4 Choosing a scaling axis

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    NEED{"What is the bottleneck?"} -->|"More CPU + storage"| VUP["Scale up vertically<br>bigger instance"]
    NEED -->|"More read CPU only"| RR2["Add read replicas"]
    NEED -->|"More storage only"| AS["RDS Storage Auto Scaling"]
    NEED -->|"More IOPS/throughput"| ST["Switch storage type<br>Provisioned IOPS"]

  • Storage types: General Purpose SSD for broad and bursty workloads, Provisioned IOPS for I/O-intensive databases requiring consistent low latency, and Magnetic, which is a legacy choice to avoid for new work.
  • RDS Storage Auto Scaling grows storage automatically toward a configured maximum with virtually zero downtime.

12.5.5 The migration itself: AWS DMS

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    SRC["Source DB<br>on-prem PostgreSQL"] --> SE["Source endpoint"]
    SE --> RT["Replication task<br>DMS replication instance"]
    RT --> TE["Target endpoint"]
    TE --> TGT["Amazon RDS<br>for PostgreSQL"]

  • DMS is a replication server in the cloud. Define source and target endpoints, run a task, and it copies data while the source stays fully operational, enabling near-zero downtime with a short cutover window.
  • This is a homogeneous migration, PostgreSQL to RDS PostgreSQL, so no AWS SCT schema conversion is required. Heterogeneous migrations, such as Oracle to PostgreSQL, need SCT.

⚠️ DMS migrates data, not your whole schema. By default, DMS creates only the objects needed to move rows, including tables with primary keys. It does not migrate secondary indexes, sequences, default values, foreign keys, constraints, views, triggers or stored procedures. For a homogeneous PostgreSQL migration, the standard pattern is pg_dump/pg_restore for the schema, followed by DMS for bulk data and ongoing change capture. CDC from PostgreSQL requires logical replication on the source, such as wal_level = logical, or rds.logical_replication on an RDS source. Configure it before cutover because it requires a restart.

Architect’s takeaway: For a no-rewrite lift-and-shift, RDS + DMS is the default. The managed engine keeps the application unchanged, Multi-AZ provides availability with endpoint-stable failover, read replicas absorb reporting load, and DMS moves data while the source remains live, provided the schema is moved separately and logical replication is enabled in advance.


12.6 Share storage across the boundary

  • ❓ On-prem apps must keep writing over NFS, yet the data has to live in AWS with low local latency. Which storage service?

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    Q["On-prem apps write files over NFS,<br>data must live in AWS"] --> Q1{"Keep NFS + low-latency<br>local writes?"}
    Q1 -->|"Pure cloud NFS"| EFS["Amazon EFS<br>NFS + lifecycle, but adds<br>network latency to AWS"]
    Q1 -->|"Hybrid local cache"| SGW["S3 File Gateway - CHOSEN<br>local NFS cache, async to S3"]
    SGW --> S3["Amazon S3<br>lifecycle, CRR, versioning"]

Chosen: AWS Storage Gateway (S3 File Gateway) + Amazon S3. The customer must keep the NFS protocol for unmodified on-prem apps and store the files in AWS. File Gateway gives local NFS/SMB access with a local cache for low-latency writes and asynchronously transfers data to S3.

S3 File Gateway architecture

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    APP["On-prem apps<br>NFS or SMB"] --> GW["S3 File Gateway<br>VM appliance + local cache"]
    GW -->|"async, encrypted in transit"| S3["Amazon S3"]
    S3 --> LC["Lifecycle to Glacier /<br>Intelligent-Tiering"]
    S3 --> CRR["Cross-Region Replication"]
    CLOUD["AWS apps / analytics"] --> S3

Why it fits every requirement at once:

  • No refactoring: apps keep talking NFS v3/v4.1 or SMB.
  • Low latency: transparent local caching buffers apps from network congestion.
  • Cost management: data lands in S3, so lifecycle policies move cold files to cheaper tiers.
  • Future-proof: once fully migrated, apps can read natively from S3 with no additional data move.

⚠️ Two operational caveats. First, upload is asynchronous, so there is a non-zero RPO. A file written locally is not instantly durable in S3. Size the local cache disk for the expected write burst and monitor upload lag. Second, if anything writes to the bucket outside the gateway, the gateway will not see it until RefreshCache is run. Treat a File Gateway bucket as gateway-owned, or automate the refresh.

File Gateway is one of three Storage Gateway types. Volume Gateway presents block storage over iSCSI, backed by EBS snapshots, and integrates with AWS Backup for volume recovery. Tape Gateway presents a virtual tape library for existing backup software. Choose the gateway type based on the protocol the on-prem application already speaks. The application uses NFS/SMB here, so File Gateway is the appropriate choice.

A CRR fact worth stating explicitly: the source and destination buckets for Cross-Region Replication do not have to be in the same AWS account, and SSE through KMS is supported for replicated objects. This is useful when the DR account is separated from the production account.

The wider storage menu:

Service Type Prefer when
S3 Object Data lakes, unlimited scale, lifecycle and tiering
S3 File Gateway Hybrid file → S3 On-prem NFS/SMB with cloud-backed storage
EFS Managed NFS Multi-instance shared file system inside AWS
FSx Managed file Specific file-system software requirements
EBS Block A disk attached to an EC2 instance

A note on EBS sizing: with the older gp2, IOPS scaled with volume size, so capacity was often over-provisioned to obtain performance. gp3 decouples them, offering a baseline of 3,000 IOPS regardless of volume size, with IOPS and throughput provisioned independently. Prefer gp3 for new volumes.

More on storage classes, durability and lifecycle: 05. Storage.

Architect’s takeaway: When the requirement is “keep NFS but store in AWS,” S3 File Gateway is the purpose-built answer. It provides a local cache for latency, S3 for durability and lifecycle, and a clean path to native S3 later. Be explicit that the upload is asynchronous.


12.7 Operate one way across both environments

  • ❓ The customer wants the same tooling on-prem and in AWS. What makes that possible, and where does it stop?

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    GOAL["Consistent operations<br>across DC and AWS"] --> ECSA["Amazon ECS Anywhere<br>run ECS-managed containers<br>on your own hardware"]
    GOAL --> SSM["AWS Systems Manager<br>Run Command, Patch,<br>Parameter Store, Maintenance Windows"]
    GOAL --> BKP["AWS Backup<br>centralised, policy-based<br>protection across AWS and on-prem"]

  • ECS Anywhere runs and manages containers on customer-managed infrastructure using the same ECS APIs, scheduling and monitoring. It requires the ECS and SSM agents on local servers, plus outbound connectivity to AWS.
  • AWS Systems Manager provides one control plane: Run Command for running scripts across fleets without SSH, Patch Manager, Maintenance Windows, and Parameter Store for configuration and KMS-encrypted SecureStrings. It requires the SSM Agent on managed nodes.
  • AWS Backup centralises and automates policy-based data protection across EC2, EBS, RDS/Aurora, DynamoDB, EFS/FSx, Storage Gateway volumes and on-prem VMware, providing one place for compliance.

⚠️ “Same tooling” has limits. ECS Anywhere external instances register with the EXTERNAL launch type and do not support awsvpc task networking, Elastic Load Balancing target-group registration, or Fargate. On-prem tasks are managed through the same control plane but are not operationally identical to cloud tasks. A local load-balancing approach is still required. The on-prem side also depends on outbound connectivity to AWS. If the link drops, running containers continue operating, but orchestration is unavailable.

Architect’s takeaway: Hybrid operational consistency comes from three pillars: ECS Anywhere for common container tooling, Systems Manager for common operations and automation, and AWS Backup for common protection policies. This gives a consistent control plane, not identical capability.


12.8 Observe both sides through one pane

  • ❓ Half the estate has no AWS API in front of it. How do you monitor it the same way?

In a hybrid estate, the link itself is the component most likely to hurt you and least likely to be watched. Instrument it first.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    DXm["Direct Connect<br>+ VPN tunnels"] --> CW["Amazon CloudWatch<br>metrics, logs, alarms"]
    ECSm["ECS cluster + services"] --> CW
    RDSm["RDS Multi-AZ<br>+ read replicas"] --> CW
    SGWm["S3 File Gateway"] --> CW
    DMSm["DMS replication task"] --> CW
    ONP["On-prem servers<br>CloudWatch agent via SSM"] --> CW
    CW --> AL["Alarms to SNS<br>on-call notification"]

Signals worth alarming on

Layer Metric Why it matters
Direct Connect ConnectionState The earliest, cleanest “the dedicated link is gone” signal
Direct Connect ConnectionBpsEgress/Ingress, ConnectionErrorCount Saturation and physical errors before users notice
Site-to-Site VPN TunnelState Confirms the failover path is available before it is needed
ECS Container Insights, CPUReservation, MemoryReservation, running vs desired task count Desired ≠ running means the cluster cannot place tasks
RDS ReplicaLag, FreeStorageSpace, DatabaseConnections, failover events Replica lag silently makes reports wrong
S3 File Gateway CachePercentDirty, CloudBytesUploaded, FilesFailingUpload A rising dirty cache means asynchronous upload is falling behind and the RPO is growing
DMS CDCLatencySource, CDCLatencyTarget Indicates whether the database is ready for cutover

Practices

  • Install the CloudWatch agent on-premises through Systems Manager hybrid activations so data-centre servers publish the same metrics and logs as EC2. This is what makes “one pane of glass” real rather than aspirational.
  • Alarm on TunnelState for the standby VPN, not only the primary Direct Connect link. An untested failover path is not a failover path.
  • Treat desired-vs-running task count as the container liveness check on both sides.

See 07. Monitoring for CloudWatch fundamentals.

Architect’s takeaway: In hybrid, monitor the seam. Connection state, VPN tunnel state and gateway upload lag are the metrics that tell you whether the two halves are still operating as one system.


12.9 Scale the moving parts

  • ❓ With ECS on EC2, what exactly scales, and how?

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
    subgraph ECS["Two scaling layers"]
        L1["Cluster Auto Scaling<br>EC2 instances"] -->|"capacity provider +<br>CloudWatch metrics"| ASG["Auto Scaling group"]
        L2["Service Auto Scaling<br>task count"] -->|"Application Auto Scaling"| TASK["Desired tasks up or down"]
    end
    CW["CloudWatch alarms"] --> L1
    CW --> L2

  • Cluster auto scaling grows or shrinks the EC2 fleet through an Auto Scaling group capacity provider with managed scaling.
  • Service auto scaling grows or shrinks the number of tasks through Application Auto Scaling, driven by CloudWatch alarms.
    • Target tracking, the recommended approach, works like a thermostat: choose a metric target and AWS maintains it.
    • Step scaling applies adjustments according to the size of the alarm breach.
  • Do not forget the database: use RDS Storage Auto Scaling for storage growth.

Architect’s takeaway: With the EC2 launch type, you own two scaling dimensions: cluster capacity through hosts and service capacity through tasks. Prefer target-tracking policies. This is operational work that Fargate would remove, and it belongs in the EC2-versus-Fargate trade-off.


12.10 Design for disaster recovery

  • ❓ How much resilience does an enterprise workload need, and what does each level cost?

DR strategies sit on a spectrum from cheap and slow to expensive and immediate. The choice is an RTO/RPO versus cost and complexity trade-off.

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    BR["Backup and Restore<br>lowest cost<br>highest RTO"] --> PL["Pilot Light<br>core infra on,<br>rest off"]
    PL --> WS["Warm Standby<br>scaled-down but<br>always running"]
    WS --> MS["Multi-site Active/Active<br>highest cost<br>near-zero RTO"]

Strategy What’s running in the recovery Region RTO/RPO Cost & complexity
Backup & restore Nothing; redeploy from backups + IaC Hours Lowest
Pilot light Core services, such as database replication and storage, are on; app servers are off Lower Low–medium
Warm standby Scaled-down but fully functional copy always on Minutes Medium
Multi-site active/active Full capacity serving traffic in all Regions Near zero Highest
  • Active/passive, including backup and restore, pilot light and warm standby, recovers into a passive site. Active/active serves traffic from all Regions.
  • IaC through CloudFormation or CDK is essential for fast, repeatable redeployment. Also back up code, configuration and AMIs, and automate redeployment through CodePipeline.

Active-active versus active-passive topologies are covered further in 08. Optimization.

Make the network itself resilient by backing Direct Connect with a Site-to-Site VPN:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    DCn["Data Center"] -->|"primary dedicated"| DX["AWS Direct Connect"]
    DCn -.->|"failover IPsec"| VPN["Site-to-Site VPN"]
    DX --> AWS["AWS VPC"]
    VPN -.-> AWS

Architect’s takeaway: Pick the cheapest DR strategy that still meets the required RTO/RPO. Most enterprises land on warm standby unless near-zero downtime justifies active/active. Never leave the network link as a SPOF: back Direct Connect with a VPN, while remembering the bandwidth cliff described in #sec-connect.


12.11 The reference architecture

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    subgraph DCn["On-Premises"]
        OC["Containers<br>ECS Anywhere"]
        ODB["PostgreSQL"]
        OFS["NFS-mounted File Share"]
        SGW["S3 File Gateway<br>local cache"]
        OFS --> SGW
    end
     
    DCn <==>|"Direct Connect<br>plus VPN failover"| AWS
     
    subgraph AWS["AWS Cloud VPC"]
        subgraph A1["AZ 1"]
            ECS1["ECS on EC2<br>private subnet"] --> NAT1["NAT gateway 1"]
        end
        subgraph A2["AZ 2"]
            ECS2["ECS on EC2<br>private subnet"] --> NAT2["NAT gateway 2"]
        end
        VPE["VPC endpoints<br>ECR, SSM, S3, Logs"]
        RDS["RDS PostgreSQL<br>Multi-AZ + read replica"]
        S3["Amazon S3<br>Intelligent-Tiering"]
        ECS1 --> RDS
        ECS2 --> RDS
        ECS1 -.-> VPE
        ECS2 -.-> VPE
        SGW -.->|async| S3
    end
     
    ODB -.->|"DMS data + pg_dump schema"| RDS
    OPS["ECS Anywhere · Systems<br>Manager · AWS Backup"] -.-> DCn
    OPS -.-> AWS
    CW["CloudWatch: connection state, tunnel state,<br>task counts, replica lag, cache lag"] -.-> AWS

How each driver is satisfied

Driver Where it is met
Dedicated, consistent connectivity Direct Connect, with MACsec or IPsec if encryption is required, plus VPN failover
Private containers, outbound only Private subnets + NAT gateway per AZ + VPC endpoints
Custom AMI + host access ECS EC2 launch type
Same tooling both sides ECS Anywhere, Systems Manager, AWS Backup
No-rewrite DB migration RDS PostgreSQL + DMS, with schema moved through pg_dump
High availability RDS Multi-AZ, redundant links, multi-AZ NAT, multi-AZ ECS
NFS access to AWS data S3 File Gateway → S3
Observability CloudWatch across both sides through the SSM-installed agent
Cost efficiency VPC endpoints instead of NAT egress, S3 Intelligent-Tiering
Resilience / DR Warm standby, or the selected DR tier, plus IaC

12.12 When not to use these choices

Symptom Reconsider Better fit
No custom AMI or host access required EC2 launch type Fargate to remove cluster operations
Low-volume or temporary connectivity Direct Connect Site-to-Site VPN, which is cheaper and quicker with no circuit lead time
Migration deadline shorter than Direct Connect provisioning Dedicated connection Hosted connection through a partner, or VPN first
Only one VPC and one account Transit Gateway Direct connection or simple peering
Shared file system inside AWS only S3 File Gateway Amazon EFS
Heterogeneous database engines One-step DMS DMS + AWS SCT for schema conversion
On-prem tasks require ELB or awsvpc networking ECS Anywhere Local orchestration, or move the workload to AWS
Near-zero RTO is mandated and funded Warm standby Multi-site active/active
Read scaling is needed rather than availability Multi-AZ standby Read replicas, and vice versa

Architect’s takeaway: Every choice here changes with differences in volume, tooling requirements, engine compatibility, provisioning time or RTO budget. Best fit means balancing technical, operational and cost factors across both environments.


12.13 Architect’s cheat sheet

The cloud decision-making for the requirement described at the beginning of the blog:

%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
    ROOT["Hybrid container<br>workload on AWS"]

    ROOT --> N["Network"]
    N --> N1["Direct Connect = consistent,<br>private, high-volume"]
    N --> N2["DX is NOT encrypted<br>by default"]
    N --> N3["DX has weeks-to-months<br>lead time"]
    N --> N4["Transit Gateway + DX<br>Gateway = many VPCs"]

    ROOT --> C["Compute"]
    C --> C1["ECS on EC2 =<br>custom AMI + SSH"]
    C --> C2["Fargate = no host,<br>no custom AMI"]
    C --> C3["One NAT gateway PER AZ"]
    C --> C4["VPC endpoints beat<br>NAT for AWS traffic"]

    ROOT --> D["Database"]
    D --> D1["RDS PostgreSQL<br>= lift-and-shift"]
    D --> D2["Multi-AZ = availability;<br>endpoint stable"]
    D --> D3["Read replica = read<br>scaling, async"]
    D --> D4["DMS moves data,<br>not full schema"]

    ROOT --> S["Storage + Ops"]
    S --> S1["S3 File Gateway =<br>NFS local, S3 behind"]
    S --> S2["ECS Anywhere + SSM + Backup"]
    S --> S3["Watch connection state<br>and VPN tunnel state"]
    S --> S4["DR: cheapest tier<br>that meets RTO/RPO"]

Network:

  • In hybrid design, the network is the foundation. Decide connectivity first.
  • Agree non-overlapping CIDR ranges between the data centre and VPC before anything else.
  • Use Direct Connect for consistent, dedicated, high-volume DataCenter(DC)↔︎AWS traffic that stays off the public internet.
  • Direct Connect is private, not encrypted. Add MACsec or an IPsec VPN over Direct Connect if compliance requires encryption.
  • Direct Connect takes weeks to months to provision. Use a hosted connection or VPN to bridge a deadline.
  • Always add a VPN failover for Direct Connect. A single connection is a SPOF, and without backup, VPC traffic drops.
  • VPN failover is survival, not capacity: approximately 1.25 Gbps per tunnel. Two Direct Connect connections at two locations provide the stronger high-availability design.
  • Site-to-Site VPN is suitable for quick, lower-cost connectivity or failover. Client VPN is for remote workforce access.
  • Transit Gateway converts an unmanageable VPC mesh into a hub.
  • A Direct Connect Gateway extends connectivity to multiple VPCs and Regions.

Compute:

  • Use ECS on EC2 when you need a custom AMI or SSH. Fargate supports neither AMI not SSH.
  • Private subnets with a managed NAT gateway provide outbound-only internet access without unsolicited inbound connections.
  • A NAT gateway is AZ-scoped. Deploy one per AZ, or the architecture recreates the SPOF it was designed to remove.
  • Use VPC endpoints, including an S3 gateway endpoint and ECR, SSM and Logs interface endpoints, to reduce NAT data-processing cost and keep AWS-service traffic off the internet.

Database:

  • RDS with a compatible managed engine supports lift-and-shift with no application rewrite.
  • Multi-AZ DB instance = availability, with synchronous replication, a non-readable standby, automatic failover and an unchanged endpoint.
  • Multi-AZ DB cluster, available for PostgreSQL and MySQL, has two readable standbys, approximately 35-second failover and faster commits.
  • Read replica = performance, with asynchronous replication, read-only access and possible replica lag. Do not confuse it with a standby.
  • Scale RDS along the appropriate axis: vertical capacity for CPU and storage, replicas for read CPU, Storage Auto Scaling for storage, or Provisioned IOPS for I/O performance.
  • DMS migrates data, not the full schema. It does not migrate secondary indexes, sequences, foreign keys or procedures. Use pg_dump/pg_restore for the schema.
  • Change Data Capture (CDC) from PostgreSQL requires logical replication on the source before cutover.

Storage:

  • S3 File Gateway keeps NFS/SMB access local with a cache while storing data asynchronously in S3, avoiding application refactoring.
  • File Gateway upload is asynchronous, creating a real RPO, and out-of-band bucket writes require RefreshCache.
  • Know the storage menu: S3 for object storage, EFS for cloud NFS, FSx for specific file systems and EBS for block storage. Prefer gp3, which decouples IOPS from volume size.

Ops:

  • ECS Anywhere runs ECS-managed containers on customer-owned hardware, but does not support awsvpc, ELB target groups or Fargate.
  • Systems Manager, including Run Command, Patch and Parameter Store, together with AWS Backup, centralises hybrid operations and protection.
  • Install the CloudWatch agent on-premises through SSM hybrid activations. That is what makes one pane of glass real.

On-Prem ↔︎ Cloud:

  • Monitor the seam: Direct Connect ConnectionState, VPN TunnelState, gateway CachePercentDirty, and DMS CDC latency.
  • ECS on EC2 has two scaling layers: cluster capacity through hosts and service capacity through tasks. Prefer target tracking.
  • DR spectrum: backup and restore → pilot light → warm standby → active/active. Cost rises and RTO falls as you move right. Use IaC for redeployment.
  • “Serverless by default” is not a law. ECS on EC2 here proves that fit beats fashion.

Sources

Inspired from course notes, “Architecting Solutions on AWS,” Week 3: Designing a Hybrid Solution for Container-Based Workloads, using an enterprise container and PostgreSQL migration use case.