%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph OnPrem["On-premises today<br>one data centre, one power feed"]
FE1["Frontend<br>HTML/CSS/JS"] --> BE1["Backend<br>Apache + Java"]
BE1 --> DB1[("MySQL")]
HAD["Hadoop cluster<br>+ visualisation tools"]
end
OnPrem -->|"decouple + migrate"| Cloud
subgraph Cloud["Goal: on AWS"]
FE2["Frontend tier"]
BE2["Backend tier"]
DB2[("Database tier")]
EMR2["EMR analytics<br>tier"]
end
14. Capstone Project: Decoupling a Three-Tier App and Migrating Hadoop Analytics with Amazon EMR
👈 Back to: 📝 Blog | 💼 LinkedIn | ✍️ Medium
Capstone: One Customer, Two Workloads, One Migration
- ❓ Key question of this chapter:
- When a customer hands you two on-premises workloads at once:
- a three-tier web app and a Hadoop analytics estate and
- tells you to decouple one and re-platform the other onto Amazon EMR
- how do you design both without one decision quietly breaking the other?
- When a customer hands you two on-premises workloads at once:
- Workflow followed in all chapters: requirements → architectural drivers → candidate services → trade-offs → decision → justified rejections.
- What’s new is that this chapter: Learnings from 3 previous chapters applied at once
- it reuses the decoupling pattern from Chapter 10,
- the ingest→store→query→visualise pipeline from Chapter 11, and
- the lift-and-shift database pattern from Chapter 12
Running example used throughout this chapter
A customer runs two workloads on physical servers in one data centre:
- A three-tier web application: a static frontend (HTML/CSS/JavaScript), a backend (Apache + a Java application), and a MySQL database — serving a public dynamic website. Today all three tiers live on the same physical servers.
- A Hadoop analytics workload: a massive on-premises dataset processed by a physical Hadoop cluster, with visualisation tools on top.
The problem the customer actually has: everything shares one data centre and one power supply. A single power outage takes down the website and the analytics estate at the same time.
Goal: migrate both workloads to AWS.
Decouple the web app’s three layers so no single failure or single server takes down the whole site.
Host the analytics workload in the cloud too, with an explicit requirement to run it on an Amazon EMR cluster.
Ingestion, storage and visualisation choices for the analytics side are ours to make.
In scope: re-platforming the three tiers and the analytics pipeline onto AWS.
Out of scope: rewriting the Java application’s business logic or the Hadoop jobs’ code.
14.1 The Method: one estate, two tracks
Both workloads answer the same four questions * what talks to what, where does compute run, * where does state live, * how is it observed and secured*
But the analytics track has one constraint fixed in advance: * it must use Amazon EMR. * That single fixed point shapes every other analytics decision in this chapter, the same way “short-staffed” shaped Chapter 11.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
A["1. Decouple<br>the app tiers"] --> B["2. Migrate the<br>database"]
B --> C["3. Ingest analytics<br>data to AWS"]
C --> D["4. Process with<br>Amazon EMR"]
D --> E["5. Query and<br>visualise"]
| Step | Question the architect answers | Winning service(s) | Section |
|---|---|---|---|
| 1 | What are the real architectural drivers? | — | 14.2 |
| 2 | Client Tier: How does the frontend become independent? | Amazon S3 + CloudFront | 14.3 |
| 3 | Backend Tier: Where does the Java backend run, decoupled? | Amazon ECS on Fargate + ALB | 14.4 |
| 4 | Database Tier: How does MySQL move, without a rewrite? | Amazon RDS for MySQL + AWS DMS | 14.5 |
| 5 | Ingestion: How does the Hadoop dataset get into AWS? | AWS DataSync (+ Snowball for the bulk seed) | 14.6 |
| 6 | Transformation/Processing: How is it processed — on the mandated EMR? | Amazon EMR, transient, reading from S3 via EMRFS | 14.7 |
| 7 | Querying: How do we query it with SQL? | AWS Glue Data Catalog + Amazon Athena | 14.8 |
| 8 | Visualization: How do business users see it? | Amazon QuickSight | 14.9 |
| 9 | How is every hop secured? | IAM + KMS + Security Groups + WAF | 14.10 |
| 10 | Monitoring: How do we see what’s happening, on both tracks? | Amazon CloudWatch | 14.11 |
| 11 | Optimization: How do we cut over safely, and control cost? | DMS CDC cutover, Spot task nodes, S3 lifecycle | 14.12 |
Architect’s takeaway: Treat this as two pipelines sharing one VPC and one security model, not one giant diagram. Every decision below is deliberately small and local — that is what makes a two-workload migration reviewable.
14.2 Turn requirements into architectural drivers
- ❓ “Decouple the app, EMR for analytics, and it’s up to you for the rest” — what does that actually constrain?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
R1["One data centre =<br>one blast radius"] --> D1["Decouple every tier;<br>no shared failure domain"]
R2["Public dynamic website"] --> D2["Independent, scalable<br>frontend + backend"]
R3["MySQL, no rewrite wanted"] --> D3["Engine-compatible<br>managed database"]
R4["Hadoop workload,<br>EMR mandated"] --> D4["S3-backed, transient<br>EMR cluster"]
R5["Ingestion/storage/viz<br>= architect's choice"] --> D5["Reuse proven<br>ingest-store-query-visualise"]
| Customer requirement | Architectural driver | Implication |
|---|---|---|
| Single data centre, single power domain | Eliminate the shared failure domain | Each tier becomes independently deployable and independently recoverable |
| Public, dynamic three-tier website | Decoupled, scalable app layers | Frontend, backend and database each get their own AWS service, not one bigger server |
| MySQL, minimal appetite for code changes | Engine compatibility over optimality | A managed MySQL-compatible database, moved with a migration service, not a rewrite |
| Hadoop workload, EMR is mandatory | Fixed compute engine for analytics | Every ingestion/storage decision must feed EMR cleanly — no service that bypasses it |
| Free choice of ingestion, storage, visualisation | Reuse a proven pipeline shape | Ingest → store → process (EMR) → query → visualise, as in Chapter 11, with EMR replacing the ETL step |
The refactor-vs-lift-and-shift decision, made explicit. This chapter deliberately takes a middle path, and states why for each tier:
- Frontend — refactor (extract to S3 + CloudFront). The HTML/CSS/JS has no server-side logic; there is nothing to “rewrite”, only to relocate. This is the cheapest possible refactor with the biggest payoff: the frontend tier disappears as a maintenance burden entirely.
- Backend — moderate refactor (containerise, don’t rewrite). The Java application is packaged into a container image, not rewritten. This decouples it from a specific physical server and lets it scale independently, without touching business logic.
- Database — lift-and-shift (same engine, migrated). MySQL becomes Amazon RDS for MySQL. Wire protocol and SQL dialect are unchanged, so the application’s data-access code needs no changes at all.
- Analytics — mandated re-platform (EMR, not a rewrite of the jobs). The customer’s Hadoop/Spark jobs keep running on EMR’s Hadoop-compatible runtime; what changes is where the data lives (S3 instead of on-cluster HDFS), not the job code.
This mirrors the assignment’s own framing: refactor where it is nearly free (frontend), refactor lightly where it buys resilience without rewriting logic (backend, analytics platform), and lift-and-shift where compatibility matters more than elegance (database engine).
Architect’s takeaway: “Decouple” and “EMR” are both non-negotiable constraints in this brief — the design freedom is entirely in how you satisfy them, not whether.
14.3 Decouple the frontend
- ❓ The frontend is static HTML/CSS/JS living on the same server as the backend. What removes it as a shared point of failure?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
Q["Static HTML/CSS/JS,<br>today on the app server"] --> Q1{"Any server-side<br>logic in these files?"}
Q1 -->|No| S3["Amazon S3 + CloudFront - CHOSEN<br>static hosting + global cache"]
Q1 -->|"Yes / dynamic pages"| BEOPT["Keep with the<br>backend tier instead"]
Chosen: Amazon S3 for storage, Amazon CloudFront in front of it. Static assets have no reason to live on a compute instance at all. Moving them to S3 removes the frontend as a dependency of the backend’s uptime, and CloudFront caches them at edge locations close to users — lower latency, and far less load reaching the origin.
| Requirement | How S3 + CloudFront meets it |
|---|---|
| No shared failure domain with the backend | S3 is independently durable (11 nines) and served without any EC2/container instance |
| Public dynamic website, fast for users | CloudFront edge caching cuts latency versus serving every request from one Region |
| DDoS / basic web protection | AWS WAF attaches to CloudFront in front of the origin |
Architect’s takeaway: The frontend tier is the easiest win in this whole capstone — it costs nothing in refactor effort and removes an entire tier from the backend’s blast radius.
14.4 Decouple the backend compute
- ❓ The Java app runs on Apache on a physical server. How does it run on AWS without a full rewrite, while still being independently scalable?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
W["Java app + Apache,<br>today on a physical server"] --> Q1{"Willing to<br>containerise?"}
Q1 -->|"Yes, no code rewrite"| ECS["Amazon ECS on Fargate - CHOSEN<br>serverless containers"]
Q1 -->|"No, keep as VMs"| EC2["EC2 Auto Scaling Group<br>viable, more to patch/manage"]
Chosen: Amazon ECS on AWS Fargate, behind an Application Load Balancer. The Java application and Apache layer are packaged as a container image — a packaging change, not a code change — and Fargate removes the underlying EC2 fleet entirely, so there is no host patching to inherit from the old physical servers.
Backend tier architecture
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
CF["CloudFront<br>frontend"] -->|"dynamic requests"| ALB["Application<br>Load Balancer"]
ALB --> T1["ECS Fargate task<br>AZ 1"]
ALB --> T2["ECS Fargate task<br>AZ 2"]
ASG["ECS Service<br>Auto Scaling"] -.-> T1
ASG -.-> T2
| ECS on Fargate (chosen) | EC2 Auto Scaling Group | |
|---|---|---|
| You manage | Just the container and its task definition | AMIs, patching, capacity |
| Rewrite required | No — container the existing app | No — same VM-based app |
| Independently scalable from DB/frontend | Yes | Yes |
| Fit here | Team wants less host management going forward | Team wants to keep VM-level control |
Architect’s takeaway: “Decoupled” does not require a rewrite. Packaging the existing Java app into a container and running it on Fargate behind an ALB satisfies the decoupling requirement with the smallest possible code change.
14.5 Decouple and migrate the database
- ❓ How does MySQL move off the shared physical server with zero application code changes and no single point of failure?
Chosen: Amazon RDS for MySQL, Multi-AZ, migrated with AWS DMS — the same lift-and-shift pattern used for PostgreSQL in 12.5, applied here to MySQL.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
APP["ECS Fargate<br>backend"] --> PRI["RDS for MySQL<br>Primary, AZ 1"]
PRI -->|"synchronous<br>replication"| STB["Standby<br>AZ 2"]
SRC["On-prem MySQL"] --> DMS["AWS DMS<br>replication task"]
DMS -->|"initial load +<br>ongoing changes"| PRI
| Requirement | How RDS + DMS meets it |
|---|---|
| No application rewrite | Same wire protocol and SQL dialect as on-prem MySQL |
| No single point of failure | Multi-AZ standby with automatic failover; endpoint does not change |
| Decoupled from the backend server | Database now lives on its own managed service, reachable only from the VPC |
| Near-zero-downtime migration | DMS replicates while the source stays live; a short cutover window at the end |
Same caveat as Chapter 12 applies here. DMS moves rows reliably, but by default only creates the objects needed to move data — secondary indexes, triggers, stored procedures and views need a separate schema export (e.g.
mysqldumpfor schema only, then DMS for bulk load and change data capture). Confirm binary logging is enabled on the source for CDC before starting the task.
Architect’s takeaway: For a database tier with “keep the app code as-is” as a hard requirement, an engine-compatible managed service plus a replication-based migration tool is the standard answer — it is exactly what makes this a lift-and-shift rather than a rewrite.
14.6 Ingest the analytics data into AWS
- ❓ The Hadoop workload’s dataset is “massive” and on-premises. How does it get into AWS, and does it need to move all at once?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
Q["Massive on-prem dataset<br>on HDFS"] --> Q1{"Network bandwidth<br>enough for a timely transfer?"}
Q1 -->|"No — too big / too slow"| SNOW["AWS Snowball Edge - CHOSEN<br>for the one-time bulk seed"]
Q1 -->|"Yes, ongoing feed"| DS["AWS DataSync - CHOSEN<br>for ongoing incremental sync"]
SNOW --> S3["Amazon S3<br>raw zone"]
DS --> S3
Chosen: a two-phase ingestion. A one-time bulk transfer with AWS Snowball Edge seeds the initial “massive” dataset without saturating a link that has not been sized for it, followed by AWS DataSync for the ongoing, incremental feed of new files landing on-premises after the seed.
| Option | Fit here |
|---|---|
| AWS Snowball Edge (chosen, initial seed) | Physical device shipped to the data centre; avoids transferring the entire historical dataset over the network in one go |
| AWS DataSync (chosen, ongoing) | Automated, scheduled, checksum-verified sync of new/changed files from an on-prem NFS/HDFS-adjacent share to S3 |
A one-off distcp / manual copy |
Rejected — no verification, no scheduling, no resumability at this data volume |
Why S3, not on-cluster HDFS, as the landing zone. The single biggest architectural decision in this chapter’s analytics track is landing data in Amazon S3 rather than on HDFS attached to the EMR cluster. It decouples storage from compute — the exact same benefit Chapter 11 found for the clickstream lake — and it is also what makes an EMR cluster safe to terminate between jobs (see 14.7).
Architect’s takeaway: Bulk-seed a large historical dataset with Snowball, then keep it current with DataSync. Either way, the destination is S3, because that is what lets the EMR layer above it be transient instead of permanent.
14.7 Process the workload with Amazon EMR
- ❓ The customer requires an EMR cluster. How do you run it in a way that’s actually cost-efficient and resilient, not just “technically EMR”?
14.7.1 EMR cluster shape: transient, not persistent
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
W["Run existing Hadoop/Spark<br>jobs on EMR"] --> Q1{"Does the cluster need<br>to run 24/7?"}
Q1 -->|"No — scheduled batch jobs"| TR["Transient cluster - CHOSEN<br>spin up, run, terminate"]
Q1 -->|"Yes — continuous processing"| PE["Persistent (long-running)<br>cluster"]
Chosen: a transient EMR cluster, launched to run each batch job (or a scheduled job window) and terminated afterwards. Because the data lives in S3, not on-cluster HDFS (14.6), terminating the cluster loses no data — the cluster is compute, not storage.
| Transient cluster (chosen) | Persistent cluster | |
|---|---|---|
| Pay for compute | Only while a job runs | 24/7, whether processing or idle |
| Data survives termination | Yes — it lives in S3 (EMRFS) | Only if data is also kept off-cluster |
| Fit here | Scheduled/batch Hadoop analytics | Continuous, always-on processing needs |
14.7.2 EMR node types and how they read S3
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
S3R["S3 raw zone"] -->|"EMRFS"| MST["Master node<br>cluster + job coordination"]
MST --> CORE["Core nodes<br>HDFS cache + compute"]
MST --> TASK["Task nodes<br>compute only, Spot-friendly"]
CORE --> S3C["S3 curated zone<br>job output"]
TASK --> S3C
- EMRFS lets Hadoop/Spark address
s3://paths as if they were HDFS, so the existing job code runs unmodified against S3. - Master node coordinates the cluster; core nodes hold a local HDFS cache and run tasks; task nodes run compute only and hold no HDFS data — which is exactly why task nodes are the safe place to use Spot Instances (see 14.12).
14.7.3 Why not EMR Serverless or EMR on EKS here
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
Q["Compute layer for<br>the Hadoop workload"] --> Q1{"Requirement says<br>'spin up an EMR cluster'?"}
Q1 -->|Yes| EMRC["EMR on EC2 - CHOSEN<br>satisfies the explicit ask"]
Q1 -->|"No constraint"| SRV["EMR Serverless<br>no cluster to manage"]
| Rejected option | Reason |
|---|---|
| EMR Serverless | Removes cluster management entirely — a good default in general, but the brief explicitly asks for a cluster, and the team’s existing operational runbooks assume one |
| EMR on EKS | Adds a Kubernetes control plane the team has no existing investment in; this workload has no Kubernetes requirement, unlike the Chapter 12 hybrid scenario |
Architect’s takeaway: EMR’s job here is to run existing Hadoop/Spark code unchanged against data that lives in S3. Making the cluster transient and reading via EMRFS is what turns “we were told to use EMR” into a design that is also cost-efficient and resilient — not just compliant with the requirement.
14.8 Catalog and query the lake
- ❓ Once EMR has written curated output to S3, how does anyone run SQL against it without standing up a database?
Chosen: AWS Glue Data Catalog for schema, Amazon Athena for ad hoc SQL — the same pattern as 11.5, reused here on EMR’s output instead of a clickstream.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
S3C["S3 curated zone<br>EMR job output"] --> GDC["Glue Data Catalog<br>schema + partitions"]
GDC --> ATH["Amazon Athena<br>serverless SQL"]
S3C --> ATH
| Requirement | How it’s met |
|---|---|
| SQL access to EMR’s output, no extra servers | Athena queries the curated S3 zone directly; no database to provision |
| Schema shared between EMR and Athena | A Glue crawler (or an EMR step that registers partitions) keeps the Data Catalog in sync with what EMR writes |
Architect’s takeaway: EMR does the heavy processing; Glue + Athena is the lightweight, serverless way to make the result queryable without adding a second always-on cluster.
14.9 Visualise the insight
- ❓ How do the business users who currently use the on-prem visualisation tools see the same insights on AWS?
Chosen: Amazon QuickSight, reading through Athena — the same choice as 11.6.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ATH["Amazon Athena"] --> QS["Amazon QuickSight<br>dashboards + SPICE"]
QS --> U["Business users<br>replace on-prem viz tools"]
| QuickSight strength | Why it fits |
|---|---|
| Serverless, pay-per-session readers | No new licensing model to negotiate for existing report consumers |
| SPICE in-memory cache | Dashboards stay fast without re-querying Athena on every click |
| Sits directly on Athena | No extra data movement beyond what EMR already wrote to S3 |
Architect’s takeaway: Visualisation is the one place where “it’s up to you” genuinely means pick the simplest thing that replaces what the customer already has — QuickSight-on-Athena does that without inventing a new data path.
14.10 Secure both tracks
- ❓ Two workloads, one VPC — what security posture has to hold for both at once?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph Transit["In transit"]
T1["HTTPS via<br>CloudFront + ALB"]
T2["TLS: DMS, DataSync,<br>EMRFS to S3"]
end
subgraph Rest["At rest"]
R1["SSE-KMS on<br>every S3 bucket"]
R2["RDS storage<br>encryption"]
end
subgraph Identity["Identity"]
I1["ECS task role:<br>only its own resources"]
I2["EMR service +<br>instance roles, scoped to S3"]
end
| Requirement | Control |
|---|---|
| Public app tier protected | AWS WAF on CloudFront and the ALB |
| App tiers isolated from each other | Security groups per tier: ALB → ECS tasks → RDS, no direct internet-to-database path |
| Analytics data protected in transit and at rest | TLS on DataSync/DMS/EMRFS transfers; SSE-KMS on every S3 bucket (raw, curated) |
| Least privilege per workload | Separate IAM roles for the ECS task, the DMS replication instance, and the EMR cluster’s service and instance profiles — each scoped only to the resources it needs |
| Database reachable only from the backend | RDS in a private subnet, security group allowing inbound only from the ECS tasks’ security group |
For IAM roles, KMS and the shared responsibility model, see 02. Security.
Architect’s takeaway: Running two workloads in one account does not mean one security posture. Scope every role — ECS task, DMS instance, EMR cluster — to only the S3 prefixes and databases it touches.
14.11 Make both tracks observable
- ❓ A backend outage is loud. A stalled analytics pipeline is silent. How do you catch both?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ALB["ALB + ECS"] --> CW["Amazon CloudWatch<br>metrics, logs, alarms"]
RDS["RDS for MySQL"] --> CW
DS["DataSync tasks"] --> CW
EMR["EMR cluster + steps"] --> CW
ATH["Athena"] --> CW
CW --> AL["Alarms to SNS"]
| Track | Signal worth alarming on |
|---|---|
| App tier | ALB 5XXCount, ECS service task health, RDS FreeStorageSpace / CPUUtilization / replication lag |
| Analytics tier | DataSync task success/failure, EMR step failures, EMR cluster idle time (a cost signal, not just health), Athena ProcessedBytes |
Architect’s takeaway: Alarm on the app tier’s errors and the analytics tier’s absence of progress (a step that never completes, a DataSync task that stops running) — the same “pipelines fail silently” lesson from 11.8 applies directly to EMR steps.
14.12 Cut over safely, and control cost
- ❓ How do you switch traffic to AWS without a big-bang outage, and keep the EMR side affordable?
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
A["DMS replicates<br>MySQL, source stays live"] --> B["DataSync keeps S3<br>current with on-prem"]
B --> C["Test app + EMR jobs<br>against AWS copies"]
C --> D["Cutover: point DNS,<br>stop on-prem writes"]
- Database cutover: keep DMS replicating (CDC) until the switch, then a short, planned cutover window — the same near-zero-downtime pattern as 12.5.5.
- Analytics cutover: run the EMR pipeline against synced S3 data in parallel with the on-prem Hadoop cluster until outputs are verified to match, before decommissioning the physical cluster.
- Cost levers for EMR: use Spot Instances on task nodes only (they hold no HDFS data, so interruption just delays a job, it doesn’t lose it); terminate the cluster between scheduled runs; keep S3 storage on Intelligent-Tiering for the raw zone.
- Cost levers for the app tier: ECS Service Auto Scaling to match ALB traffic; RDS storage auto scaling instead of over-provisioning up front.
For broader cost and resilience patterns, see 08. Optimization.
Architect’s takeaway: Both cutovers follow the same shape — replicate in parallel, verify, then switch — and both cost profiles improve the same way: pay for compute only while it’s doing something (ECS auto scaling, transient EMR, Spot task nodes).
14.13 The reference architecture
Application tier on AWS
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
USERS["Internet users"] --> CF["CloudFront + WAF"]
CF -->|"static"| S3F["S3<br>frontend"]
CF -->|"dynamic"| ALB["Application<br>Load Balancer"]
ALB --> ECS["ECS on Fargate<br>Java backend"]
ECS --> RDS["RDS for MySQL<br>Multi-AZ"]
DMS["AWS DMS"] -->|"migrate"| RDS
Analytics tier on AWS
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ONPREM["On-prem HDFS<br>dataset"] -->|"bulk seed"| SNOW["Snowball Edge"]
ONPREM -->|"ongoing"| DS["AWS DataSync"]
SNOW --> S3R["S3 raw zone"]
DS --> S3R
S3R -->|"EMRFS"| EMR["Amazon EMR<br>transient cluster"]
EMR --> S3C["S3 curated zone"]
S3C --> GDC["Glue Data Catalog"]
GDC --> ATH["Amazon Athena"]
S3C --> ATH
ATH --> QS["Amazon QuickSight"]
How each requirement is satisfied
| Requirement | Where it is met |
|---|---|
| Decouple frontend, backend, database | S3+CloudFront / ECS on Fargate+ALB / RDS for MySQL — three independent services, no shared server |
| No app rewrite | Backend is containerised, not rewritten; database engine is unchanged |
| Host both workloads in the cloud | Application tier and analytics tier both run entirely on AWS, in one VPC |
| Analytics must run on Amazon EMR | Transient EMR cluster, reading/writing S3 via EMRFS |
| Architect’s choice of ingest/store/visualise | Snowball + DataSync ingest; S3 store; Glue + Athena query; QuickSight visualise |
| No single point of failure | Multi-AZ RDS, multi-AZ ECS tasks, CloudFront edge distribution, EMR data safe in S3 independent of cluster life |
14.14 When not to build it this way
An honest architect states the boundaries of their own recommendation.
| Symptom | Why this design struggles | Better fit |
|---|---|---|
| Analytics needs sub-second, always-on processing | Transient EMR is batch-shaped | A persistent EMR cluster, or a streaming service as in 11. Serverless Data Analytics |
| No constraint to use a cluster at all | EMR (on EC2) still carries cluster concepts to manage | EMR Serverless, once the “spin up a cluster” requirement is gone |
| Backend needs a custom AMI or host-level SSH | Fargate gives no host access | ECS on EC2, as in 12.4 |
| Team already has deep Kubernetes investment | ECS/EMR-on-EC2 underuses that skill | EKS for the app tier, EMR on EKS for analytics |
| Heavy, predictable, always-on BI workload at scale | Athena’s per-scan pricing stops being the cheapest option | Amazon Redshift |
| Database needs read scaling beyond one primary | A single RDS instance | Add RDS read replicas, as in 12.5.3 |
Architect’s takeaway: This design is optimal for batch-shaped Hadoop analytics with a mandated EMR cluster, alongside a public three-tier app that needs decoupling but not a rewrite. Change the traffic pattern, the team’s skills, or the “must use EMR” constraint, and the winning services change with it.
14.15 Architect’s cheat sheet
The cloud decision-making for the requirement mentioned in the beginning of the blog:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ROOT["Capstone:<br>app + EMR analytics"]
ROOT --> APP["App tier"]
APP --> A1["Frontend to S3<br>+ CloudFront"]
APP --> A2["Backend containerised<br>on ECS Fargate"]
APP --> A3["Database lift-and-shift<br>to RDS + DMS"]
ROOT --> AN["Analytics tier"]
AN --> N1["Snowball for bulk seed,<br>DataSync ongoing"]
AN --> N2["S3 = the landing zone,<br>not on-cluster HDFS"]
AN --> N3["EMR = transient,<br>reads via EMRFS"]
AN --> N4["Glue + Athena +<br>QuickSight on top"]
ROOT --> X["Cross-cutting"]
X --> X1["Scope every IAM role<br>to its own resources"]
X --> X2["Encrypt in transit<br>and at rest"]
X --> X3["Alarm on errors AND<br>on silent pipeline stalls"]
X --> X4["Replicate in parallel,<br>verify, then cut over"]
- Two workloads sharing one data centre and one power feed is a single blast radius — decoupling removes that shared failure domain.
- Refactor where it’s cheap (static frontend → S3), refactor lightly where it buys resilience (containerise the backend, don’t rewrite it), lift-and-shift where compatibility matters most (MySQL → RDS for MySQL).
- Frontend → S3 + CloudFront: no server-side logic to preserve, so there is nothing to “rewrite” — only to relocate.
- Backend → ECS on Fargate + ALB: containerising is a packaging change, not a code change.
- Database → RDS for MySQL + DMS: same wire protocol and SQL dialect keeps the application’s data-access code untouched.
- DMS moves data, not the whole schema — export secondary indexes, triggers and stored procedures separately.
- A mandated EMR cluster is a fixed point that shapes every ingestion and storage decision around it — check it does not get silently bypassed.
- Land analytics data in S3, not on-cluster HDFS — this single choice is what makes the EMR cluster safe to terminate.
- Transient EMR clusters cost less than persistent ones when jobs are scheduled/batch, precisely because the data survives termination in S3.
- EMRFS lets existing Hadoop/Spark code address
s3://paths unmodified — no job rewrite needed. - Core nodes cache HDFS data; task nodes are compute-only — task nodes are the safe place for Spot Instances.
- Reject EMR Serverless only when the requirement explicitly asks for a managed cluster; otherwise it is usually the simpler default.
- Snowball Edge for the one-time bulk historical seed; DataSync for the ongoing incremental feed — pick based on data volume vs. link bandwidth.
- Glue Data Catalog + Athena turn EMR’s S3 output into ad hoc SQL without a second always-on database.
- QuickSight on Athena replaces on-prem visualisation tools without inventing a new data path.
- Scope every IAM role separately — ECS task, DMS replication instance, EMR service/instance profile — to only what it needs.
- Encrypt in transit and at rest on both tracks: TLS everywhere, SSE-KMS on every bucket, RDS storage encryption.
- App-tier failures are loud (5XX errors); analytics-tier failures are silent (a stalled EMR step, a paused DataSync task) — alarm on both kinds.
- Cut over both tracks the same way: replicate in parallel with the source live, verify outputs match, then switch — never a big-bang cutover.
- Cost control on both tracks follows one rule: pay for compute only while it is doing something — ECS auto scaling, transient EMR, Spot task nodes.
- Draw one diagram per workload, not one diagram for everything — it is easier to review, and easier to get right.
- “Technically satisfies the requirement” (an EMR cluster) is not the same as “well-architected” — make the cluster transient and S3-backed, or you have complied with the letter of the ask while paying for the worst version of it.
Sources
Inspired from Capstone Project Question