%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
Block["Each building block<br/>of a System"] --> Q{"Made concrete<br/>in the cloud"}
Q --> M["💰 What does it cost<br/>when idle vs busy?"]
Q --> L["⏱️ What latency does<br/>this choice bake in?"]
Q --> R["💥 If it fails,<br/>how far does it spread?"]
M --> Dec["The service choice<br/>is the answer to<br/>all three at once"]
L --> Dec
R --> Dec
Architecting Solutions in AWS: Requirements to Justified Trade-offs.
Being a Technical Architect means composing systems from building blocks. Being a Cloud Architect is lower-level systems thinking: Leveraging real cloud services to build each block, and at what cost?
👈 Back to: 📝 Blog | 💼 LinkedIn | ✍️ Medium
❓ 1. Why this blog and the book
Knowing what a service does is not the same as knowing when to choose it and what you must be prepared to defend once you have.
Technical Architect → Cloud Architect
In the previous blog post on being a Technical Architect, we composed systems from building blocks. Each block was a clean item on a whiteboard.
In Cloud, for each block one needs to worry about:
- What is the cost of building a block in a particular way in cloud?
- What is the associated latency for that chosen path?
- What is the blast radius if things fail?
This is a companion to Foundational Questions. Visual Answers. Mastering AWS Fundamentals. That book builds the fundamentals, chapter by chapter: security, compute, networking, storage, databases, monitoring, optimization. This book assumes those fundamentals and asks a harder question: given a real customer’s requirements, how do you combine those building blocks into one architecture — and justify the services you rejected?
🧭 2. Flow of Decision-making for a Cloud Architect

- Functional drivers = what the system must do.
- For example: “store relational data”, “run SQL analytics”, “serve a REST API”, “process a stream”, “run containerized microservices”.
- These filter the candidate set; if the requirement is “transactional relational writes”, a queue and an object store are simply not candidates.
- Non-functional drivers = how well it must do it.
- Scale, latency, cost, resilience, ops burden.
- These rank the survivors — among the relational options (RDS, Aurora, self-managed on EC2), the NFRs pick the winner.
📝 3. What this book covers
Every chapter follows the same workflow: requirements/problem understanding → architectural drivers → collect candidate services → compare trade-offs → finalize decision → justify rejections.
🧩 10. Architecting a Serverless Web Backend
- Problem: A monolithic, on-prem order service must become an auto-scaling, decoupled, serverless backend in the cloud
Before an Architect’s Intervention:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
U[Customers] --> W[Web tier]
W --> O[Monolithic order service]
O --> M[(MySQL)]
O -.-> I[Inventory service]
O -.-> A[Accounting service]
O -.-> F[Fulfillment service]
subgraph InScope["In scope for Migration"]
W
O
M
end
subgraph OutOfScope["Out of scope"]
I
A
F
end
The Naive First Draft: “Make it serverless”: swap the monolith for one Lambda that does everything.
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
FE["Frontend clients"] --> APIGW["API Gateway"]
APIGW --> L["One Lambda<br>all order logic"]
L --> DDB["DynamoDB"]
L -->|"sync calls"| DS["Inventory, Accounting,<br>Fulfillment"]
Why the First Draft Fails → What it needs:
- One Lambda for every route is a monolith in disguise; one bug takes all routes down → split functions by responsibility.
- Synchronous downstream calls: one slow service fails the whole order → decouple with SQS, fan out with SNS.
- A sale-day burst hits Lambda and DynamoDB head-on → SQS absorbs it, Lambda drains at its own pace.
- Poison messages retry, then silently expire → dead-letter queue makes failures visible.
After an Architect’s Intervention:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
FE["Frontend clients"] --> APIGW["API Gateway<br>auth, validation, throttling"]
APIGW --> SQS["SQS<br>buffer / decouple"]
SQS --> LAM["Lambda<br>order processing"]
LAM --> DDB["DynamoDB<br>orders"]
LAM -.->|"failures"| DLQ["Dead-letter queue"]
DDB --> STR["DynamoDB Streams"]
STR --> LAM2["Lambda<br>publisher"]
LAM2 --> SNS["SNS topic"]
SNS --> D1["Fulfillment"]
SNS --> D2["Accounting"]
SNS --> D3["Inventory"]
APIGW --> CW["CloudWatch + X-Ray<br>across all services"]
LAM --> CW
DDB --> CW
SNS --> CW
📊 11. Serverless Data Analytics
- Problem: A Static S3-hosted menus website needs a clickstream pipeline to be used along with it.
- This should turn a one-off request into a recurring ingest → store → query → visualise data pipeline.
- Such a Data Analytics pipeline gives insights to restaurant owners on analytics from customer preferences
Before an Architect’s Intervention:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
C[Customer scans QR code] --> S3M["S3 static menu<br>HTML pages"]
S3M --> ORD["'Order this item'<br>feature"]
ORD -.-> PAY["Third-party payment API"]
S3M --> CLICK["Clickstream:<br>views + navigation"]
CLICK --> PIPE["Analytics pipeline<br>ingest, store, query, visualise"]
subgraph InScope["In scope<br>(newly added)"]
CLICK
PIPE
end
subgraph OutOfScope["Out of scope<br>(not touching for edit)"]
ORD
PAY
end
The Naive First Draft: “Just save every click somewhere and query it.”
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
B["Menu in browser"] -->|"one PUT per click"| S3["S3 bucket<br>raw JSON"]
S3 --> ATH["Athena<br>scans raw JSON"]
ATH --> QS["QuickSight<br>queries on every view"]
Why the First Draft Fails → What it needs:
- Browser writes straight to S3: exposed bucket, no throttling → API Gateway as the front door.
- One write per click: millions of tiny files and per-request charges → Firehose buffers into batched files.
- Athena bills by bytes scanned; raw JSON drags every byte of every row → Parquet + Glue catalog + time partitions.
- Every dashboard view re-scans S3 → QuickSight SPICE caches the dataset in memory.
After an Architect’s Intervention:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
USER["Customer<br/>scans QR code"]
subgraph SITE["Static Menu Website"]
EDGE["CloudFront distribution<br/><br/>protected by AWS WAF"]
MENU[("S3 private origin<br/><br/>HTML, CSS and JavaScript")]
BROWSER["Menu running in browser<br/>JS emits click events"]
EDGE -->|"Fetch static assets"| MENU
EDGE -->|"Serve menu application"| BROWSER
end
USER -->|"Request menu"| EDGE
subgraph PIPELINE["Serverless Clickstream Analytics Pipeline"]
AG["API Gateway<br/><br/>HTTPS, throttling and abuse control"]
KDF["Amazon Data Firehose<br/><br/>Direct PUT, buffering,<br/>JSON to Parquet"]
subgraph LAKE["S3 Data Lake"]
DATA["Parquet data<br/>Hive-style prefixes"]
ERR["Error output prefix"]
end
REPLICA[("S3 replica bucket<br/><br/>Second Region")]
GDC["Glue Data Catalog<br/><br/>Table schema and partitions"]
ATH["Amazon Athena<br/><br/>SQL on S3"]
SPICE[("QuickSight SPICE<br/><br/>In-memory dataset")]
QS["QuickSight Dashboards"]
BROWSER -->|"Click events"| AG
AG --> KDF
GDC -.->|"Schema for<br/>Parquet conversion"| KDF
KDF --> DATA
KDF -.->|"Failed records"| ERR
DATA -->|"Asynchronous CRR<br/>with versioning"| REPLICA
GDC -.->|"Table metadata"| ATH
DATA --> ATH
ATH -->|"Dataset ingestion<br/>and scheduled refresh"| SPICE
SPICE --> QS
end
subgraph OBS["Cross-Cutting Observability"]
CW["Amazon CloudWatch<br/><br/>Metrics, logs and alarms"]
SNS["Amazon SNS<br/>Operational notifications"]
CW --> SNS
end
AG -.->|"Count, 4XX, 5XX, latency"| CW
KDF -.->|"Delivery success,<br/>freshness and throttling"| CW
REPLICA -.->|"Replication lag"| CW
ATH -.->|"Failures and bytes scanned"| CW
SPICE -.->|"Ingestion or refresh failures"| CW
IAC["CloudFormation<br/>provisions the architecture"] -.-> SITE
IAC -.-> PIPELINE
IAC -.-> OBS
🔗 12. Hybrid, Container-Based Workloads
- Problem: Half a company’s workloads stay on-premises, half move to AWS.
- Need: Connecting the On-prem Data Center with AWS infra, keeping containers portable, and migrating a database with near-zero downtime.
The Naive First Draft: “Spin up a new cluster in AWS, connect over the internet, dump-and-restore the database.”
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
DC["Data centre<br>containers, PostgreSQL, NFS"] -->|"public internet"| K["New cluster in AWS<br>own tooling, public subnets"]
DC -->|"weekend dump + restore"| DB["PostgreSQL on AWS"]
DC -->|"rewrite apps for S3 API"| S3["Amazon S3"]
Why the First Draft Fails → What it needs:
- The internet link brings latency spikes to DC↔︎AWS traffic that will run for years → Direct Connect + VPN failover.
- Different tooling on each side means two ops models → ECS in AWS + ECS Anywhere on-prem.
- Containers in public subnets widen the attack surface → private subnets, NAT gateways, VPC endpoints.
- Dump-and-restore means hours of downtime → DMS for data,
pg_dumpfor schema only. - Rewriting NFS apps for S3 is a refactor the brief rules out → S3 File Gateway keeps NFS.
After an Architect’s Intervention (simple version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph DCn["Data centre (contracts<br>expiring in waves)"]
direction TB
C1["Containerised apps"]
DB1["PostgreSQL"]
C1 <--> DB1
end
DCn <-->|"dedicated, low-latency link"| AWS
subgraph AWS["AWS (half the workloads)"]
C2["Containers<br>private, outbound only"]
DB2["PostgreSQL<br>lift-and-shift, no rewrite"]
NFS["Data over NFS<br>no refactor"]
end
C1 -.->|"same orchestration tooling"| C2
DB1 -.->|"no app code rewrite"| DB2
After an Architect’s Intervention (detailed version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph DCn["On-Premises"]
OC["Containers<br>ECS Anywhere"]
ODB["PostgreSQL"]
OFS["NFS-mounted File Share"]
SGW["S3 File Gateway<br>local cache"]
OFS --> SGW
end
DCn <==>|"Direct Connect<br>plus VPN failover"| AWS
subgraph AWS["AWS Cloud VPC"]
subgraph A1["AZ 1"]
ECS1["ECS on EC2<br>private subnet"] --> NAT1["NAT gateway 1"]
end
subgraph A2["AZ 2"]
ECS2["ECS on EC2<br>private subnet"] --> NAT2["NAT gateway 2"]
end
VPE["VPC endpoints<br>ECR, SSM, S3, Logs"]
RDS["RDS PostgreSQL<br>Multi-AZ + read replica"]
S3["Amazon S3<br>Intelligent-Tiering"]
ECS1 --> RDS
ECS2 --> RDS
ECS1 -.-> VPE
ECS2 -.-> VPE
SGW -.->|async| S3
end
ODB -.->|"DMS data + pg_dump schema"| RDS
OPS["ECS Anywhere · Systems<br>Manager · AWS Backup"] -.-> DCn
OPS -.-> AWS
CW["CloudWatch: connection state, tunnel state,<br>task counts, replica lag, cache lag"] -.-> AWS
🏛️ 13. Account Governance and Multi-Account Management
- Problem: A single, unmanaged AWS account has to become a governed multi-account structure, with automatic provisioning, single sign-on, central logging and guardrails.
The Naive First Draft: “Keep one account; separate clients with tags and IAM policies.”
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ROOT["One account<br>root + long-lived IAM users"] --> T["Tags + IAM policies<br>per client"]
T --> A["Client A"]
T --> B["Client B"]
T --> C["Client C"]
Why the First Draft Fails → What it needs:
- One misconfigured policy reaches every client → one account per client, grouped in Organizations OUs.
- Cost per client depends on perfect tagging → the account is the free, default unit of cost allocation.
- Quotas and API limits are per account, so one noisy client throttles all → separate accounts, separate quotas.
- Shared root and long-lived IAM users → IAM Identity Center (SSO) + Control Tower guardrails.
After an Architect’s Intervention (simple version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph Before["Single account (today)"]
ROOT["root + unmanaged IAM users"] --> W1["Client A website"]
ROOT --> W2["Client B website"]
ROOT --> W3["Client C website"]
end
Before -->|"re-architect"| After
subgraph After["Multi-account (goal)"]
SSO["Single sign-on"] --> ACCA["Client A account"]
SSO --> ACCB["Client B account"]
SSO --> ACCC["Client C account"]
ACCA --> LOG["Dedicated logging account"]
ACCB --> LOG
ACCC --> LOG
GR["Guardrails +<br>enforced tagging"] -.-> ACCA
GR -.-> ACCB
GR -.-> ACCC
end
After an Architect’s Intervention (detailed version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
subgraph Org["AWS Organization -<br>all features enabled"]
MGMT["Management account<br>Control Tower, org CloudTrail, billing<br>NO workloads"]
subgraph SecOU["Security OU"]
LOGACC["Log Archive account<br>central CloudTrail and Config"]
TOOL["Security Tooling account<br>delegated admin: GuardDuty,<br>Config, Security Hub"]
end
SS["Shared Services account<br>delegated IAM Identity Center admin"]
subgraph WL["Workloads OU"]
subgraph CustA["Customer A OU"]
DEVA["Dev"]
PRDA["Prod"]
end
end
end
USER["Workforce"] -->|SSO portal| SS
SS -->|permission sets| DEVA
SS -->|permission sets| PRDA
MGMT -->|SCP guardrails| WL
DEVA -->|logs| LOGACC
PRDA -->|logs| LOGACC
TOOL -->|findings| SEChub["Security Hub"]
🎯 14. Capstone: Decoupled App + EMR Analytics
Problem:
- One customer, migrate two workloads to cloud at once:
- decoupling a three-tier app and
- re-platforming a Hadoop estate onto Amazon EMR
- This capstone effort reuses the patterns from Chapters 10-12 together.
Before an Architect’s Intervention:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph OnPrem["On-premises today<br>one data centre, one power feed"]
FE1["Frontend<br>HTML/CSS/JS"] --> BE1["Backend<br>Apache + Java"]
BE1 --> DB1[("MySQL")]
HAD["Hadoop cluster<br>+ visualisation tools"]
end
The Naive First Draft: “Lift and shift as-is.”
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
U["Users"] --> EC2["One EC2<br>Apache + Java + static files"]
EC2 --> RDS[("RDS MySQL<br>single-AZ")]
HAD["Always-on Hadoop<br>cluster on EC2 + HDFS"]
Why the First Draft Fails → What it needs:
- One EC2 is a single point of failure → ALB + ECS on Fargate across AZs.
- A single-AZ database moves the same risk one tier down → RDS Multi-AZ standby.
- The backend serves static files it shouldn’t → CloudFront + WAF + S3 for the frontend.
- An always-on Hadoop cluster pays for idle compute, with data locked in HDFS → transient EMR reading S3 via EMRFS.
After an Architect’s Intervention (simplified version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
subgraph OnPrem["On-premises today<br>one data centre, one power feed"]
FE1["Frontend<br>HTML/CSS/JS"] --> BE1["Backend<br>Apache + Java"]
BE1 --> DB1[("MySQL")]
HAD["Hadoop cluster<br>+ visualisation tools"]
end
OnPrem -->|"decouple + migrate"| Cloud
subgraph Cloud["Goal: on AWS"]
FE2["Frontend tier<br>CloudFront-WAF-S3"]
BE2["Backend tier<br>ALB-<br>ECS_on_Fargate"]
DB2[("Database tier")]
EMR2["EMR analytics<br>tier"]
end
After an Architect’s Intervention (detailed version):
Application tier on AWS
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
USERS["Internet users"] --> CF["CloudFront + WAF"]
CF -->|"static"| S3F["S3<br>frontend"]
CF -->|"dynamic"| ALB["Application<br>Load Balancer"]
ALB --> ECS["ECS on Fargate<br>Java backend"]
ECS --> RDS[("RDS for MySQL<br>Multi-AZ")]
OnPremMySQ[("On-prem MySQL")] --> DMS["AWS DMS"] -->|"migrate<br>Full load or<br>Changed Data"| RDS
Analytics tier on AWS
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
ONPREM["On-prem HDFS<br>dataset"] -->|"bulk seed"| SNOW["Snowball Edge"]
ONPREM -->|"ongoing"| DS["AWS DataSync"]
SNOW --> S3R["S3 raw zone"]
DS --> S3R
S3R -->|"EMRFS"| EMR["Amazon EMR<br>transient cluster"]
EMR --> S3C["S3 curated zone"]
S3C --> GDC["Glue Data Catalog"]
GDC --> ATH["Amazon Athena"]
S3C --> ATH
ATH --> QS["Amazon QuickSight"]
🤖 15. Agentic AI Harness: Desktop to AWS
- Problem: An enterprise agent harness (xAgents) already runs as a desktop app: a Rust core running a custom GOAP-inspired ReAct loop over local SQLite, LanceDB, MCP tools and multiple LLMs.
- Need: Org-wide agents in AWS now (and Azure in future), running the same Rust codebase, with one governance model, enterprise identity and central observability — while personal agents stay on the desktop.
- Unlike Chapters 10–14, the hard problem is not scale or migration — it is portability without a rewrite.
What xAgents is today:
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
UI["React + TS UI"] <--> IPC["Tauri IPC"] <--> CORE["Rust Core<br/>GOAP ReAct Loop"]
CORE --> LLM["LLM Gateway<br/>Bedrock · Azure · Ollama"]
CORE --> TOOLS["Tools<br/>File · CSV · Excel · PDF · MCP"]
CORE --> DB[("SQLite<br/>State + Audit")]
The Naive First Draft: “Fork the core for the cloud, hard-wire it to cloud services, sync every desktop’s data up, and call it from a web UI.”
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
USER["User"] --> DESKAPP["Desktop app<br>React FE"] --> DESK["Desktop core<br>with SQLite, LanceDB"]
USER --> WEB["Web UI"]
DESK -->|"sync all data up"| S3[("S3 / cloud DB")]
WEB -->|"sync API call,<br>waits for whole run"| FORK["Forked cloud core<br>with Postgres, OpenSearch, Gateway"]
FORK --> S3
Why the First Draft Fails → What it needs:
- A forked cloud core means two codebases that drift apart; every fix and every skills-governance rule must be made twice → one Rust codebase; desktop and cloud differ only in the adapters plugged into it.
- The forked core is hard-wired to its stores (today the AgentLoop calls SQLite, LanceDB and a
matchon the LLM provider directly), so moving to cloud means editing the loop → extract provider interfaces (Persistence, Memory, VectorStore, Registry, Tool, LLM), wrap the existing desktop implementations behind them, and add cloud implementations alongside. - Syncing all desktop data up breaks the rule that sensitive data stays on-device unless explicitly routed → desktop runs keep their data locally; org-wide agents use cloud persistence; data moves from desktop to cloud only when a user routes it.
- A synchronous API call that waits for the whole multi-step run outlasts API Gateway’s ~30 s integration timeout → submit the run, return a run ID, stream progress.
- A web UI calling the core directly has no enterprise auth → reuse the desktop’s Azure Entra ID tenant, validated by an API Gateway JWT authorizer in front of the core.
After an Architect’s Intervention (over-simplified version): Core - Provider Architecture
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart TD
CORE["SAME xAgents Core<br>owns behaviour"]
CORE -->|"provider interfaces"| B["Adapters own integrations"]
subgraph Desktop["Desktop"]
D1[("SQLite + LanceDB")]
D2["Local MCP · Ollama"]
end
subgraph AWS["AWS Cloud - now<br>(EKS)"]
A1[("RDS PostgreSQL ·<br>Bedrock KB")]
A2["AgentCore Gateway ·<br>Bedrock"]
end
subgraph Azure["Azure - future<br>(AKS)"]
Z1[("Azure PostgreSQL ·<br>Azure AI Search")]
Z2["MCP servers ·<br>Azure OpenAI"]
end
B --> Desktop
B --> AWS
B -.-> Azure
classDef core fill:#FFE0B2,stroke:#FB8C00,stroke-width:2px
class CORE core
After an Architect’s Intervention (medium complexity):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
U["Enterprise users"] --> D["Desktop app<br>Tauri + React"]
U --> W["Web UI<br>org-wide agents"]
CB["ONE Rust<br>codebase in Git"] -.-> DCORE
CB -.-> CCORE
CB -.-> ZCORE
D --> DCORE["Rust Core<br>AgentLoop"]
DCORE --> SQL[("SQLite")]
DCORE --> LAN[("LanceDB")]
W --> CCORE["Same Rust Core<br>in AWS"]
subgraph Current["Current state<br>(minimal refactoring only,<br>to keep one codebase for cloud)"]
D
DCORE
SQL
LAN
end
subgraph InScope["In scope<br>(newly added)"]
W
CCORE
ZCORE["Same Rust Core<br>in Azure (future)"]
AWSRDS[("Amazon RDS PostgreSQL")]
AWSOPEN[("Amazon OpenSearch")]
end
CCORE --> AWSRDS
CCORE --> AWSOPEN
CCORE --> MCP["External MCP tools"]
DCORE --> MCP["External MCP tools"]
After an Architect’s Intervention (detailed version):
%%{init: {'theme': 'base', 'themeVariables': {'primaryColor': '#E3F2FD', 'primaryBorderColor': '#1E88E5', 'lineColor': '#424242', 'fontSize': '14px'}}}%%
flowchart LR
USER["Enterprise user"]
ENTRA["Azure Entra ID<br>SSO, MFA, groups"]
subgraph DESKTOP["Desktop harness - personal agents"]
TAURI["Tauri + React UI"] --> DCORE["xAgents Core"]
DCORE --> DSTORE[("SQLite + LanceDB")]
end
subgraph CLOUD["AWS harness - org-wide agents"]
AMP["Amplify<br>React web UI"] --> APIGW["API Gateway<br>JWT authorizer,<br>async + WebSocket"]
APIGW -->|"VPC Link"| ALB["Internal<br>Application<br>Load Balancer"]
subgraph EKS["Amazon EKS"]
CCORE["SAME xAgents Core"]
end
ALB --> CCORE
CCORE --> RDS[("RDS PostgreSQL")]
CCORE --> KB[("Bedrock KB")]
CCORE --> GW["AgentCore Gateway<br>+ Policy"]
CCORE --> REG["AWS Agent Registry"]
CCORE --> GRD["Bedrock Guardrails"] --> LLM["Bedrock ·<br>Azure OpenAI"]
OBS["AgentCore Observability<br>+ CloudWatch"]
end
USER --> TAURI
USER --> AMP
TAURI <-.->|"sign-in"| ENTRA
AMP <-.->|"sign-in"| ENTRA
GW --> ENT["External Tools<br>SAP, SharePoint,<br>M365, APIs"]
DCORE --> ENT
DCORE -.->|"pull approved<br>agents and skills"| REG
CCORE -.->|"OpenTelemetry step traces"|OBS
DCORE -.->|"OpenTelemetry step traces"|OBS
classDef core fill:#FFE0B2,stroke:#FB8C00,stroke-width:2px
class DCORE,CCORE core
🏁 4. Conclusion
How the functional and non-functional drivers dictated the architecture decisions across the chapters:
| Chapter | Functional driver (produces candidates) | Non-functional driver (picks winner) |
|---|---|---|
| 10. Serverless backend | “Serve a REST API + persist orders” → Lambda/Fargate + DynamoDB/RDS | Spiky load + zero-idle cost → Lambda + SQS + DynamoDB |
| 11. Data analytics | “Ingest a stream + run SQL for dashboards” → Kinesis/Firehose, Athena/Redshift | Occasional queries + cost → Firehose + S3 + Athena (read-time) over a warehouse |
| 12. Containers | “Run containerized microservices” → ECS/EKS, Fargate/EC2 | Ops burden + scale → ECS + Fargate unless the K8s ecosystem is genuinely needed |
| 13. Multi-account | “Isolate clients/environments” → single account, per-client accounts, per-env accounts | Blast radius + billing clarity → Organizations + OUs over one shared account |
| 14. Capstone (monolith) | “Transactional relational store” → RDS/Aurora/EC2-hosted DB | Resilience + cost → RDS Multi-AZ over single-AZ or self-managed |
| 15. Agentic Harness in both Desktop & AWS Cloud | “Run a long, tool-calling agent loop for many users” → AgentCore Runtime, EKS, ECS Fargate, Lambda | Portability + one codebase + AWS-now/Azure-later → unchanged core on EKS behind provider interfaces (Kubernetes carries over to AKS); AgentCore Runtime if AWS-only |
The repeated message across the chapters:
- Start with the smallest sufficient set of blocks and add one piece of complexity per pain point (a failure to isolate, a burst to absorb, idle resources you should not pay for).
- Blast radius isn’t always a resource sitting inside an account, sometimes the blast radius is the account, and the account boundary is the fix.
- When the same system must run in more than one place, the first decision is not a service but a boundary: what the platform is (behaviour) versus what it uses (infrastructure). Adopt managed services that provide capabilities; reject the ones that would take over the behaviour (Ch. 15).
- None of these decisions are made in isolation on a real system: a single build typically decouples its tiers (Ch. 14), picks an ingest/store/query path for its data (Ch. 11), and sits inside an account structure that contains the blast radius of all of it (Ch. 13) — at the same time, not in sequence.
📜 5. Source & Attribution
The intuition and examples in Chapters 10–14 are inspired by the Architecting Solutions in AWS Coursera course.
Chapter 15 applies the same method to a an imaginary Agentic harness app - xAgents, with AWS service facts checked against AWS documentation in October 2026.
