AWS SAA Study Guide

One-page study guides for the AWS Certified Solutions Architect – Associate (SAA-C03) exam. Pick a topic on the left. Official exam page β†—

πŸ”‘ IAM: One-Page Study Guide
Identity (Who) + Policy (What's allowed) β†’ access to a Resource IAM is global (not region-specific) and free.
πŸ‘€ Identities
TypeWhat it isCredentials
User A person or app (long-term identity) Password (console) and/or Access Keys (CLI/API)
Group A bucket of Users, for bulk permission assignment None: can't log in as a Group, can't nest groups
Role A temporary identity "assumed" by a User, AWS service, or federated identity None stored: temporary credentials issued via STS, auto-expire
Rule of thumb: Users/apps needing standing access β†’ User. Anything temporary or service-to-service (e.g. EC2 β†’ S3) β†’ Role. Role is always the more secure/best-practice choice when it fits.
Service-linked role: a distinct Role subtype. Some AWS services (e.g. Auto Scaling, RDS) create and manage their own Role automatically the first time you use a feature that needs it. You don't create it, and you usually can't edit its permissions directly. It's pre-linked to that one service's use case and AWS controls when it can be deleted (typically only once nothing is using it). A common wrong-answer distractor: it's still an IAM Role, just not one you author yourself.
Cross-account Role + External ID (frequently tested!) A Role's trust policy can name another AWS account (not just a service) as the principal allowed to assume it. This is how one account (e.g. a monitoring/audit vendor, or a security team's account) gets scoped, temporary access into another without a User ever existing in the target account. When the trusting account and the assuming account belong to different organizations (a third-party SaaS vendor, not your own AWS Organization), add an External ID (a shared secret string the third party must supply in its AssumeRole call) to the trust policy's condition. This defeats the "confused deputy" problem: without it, the third party could be tricked into assuming a role on behalf of one of its other customers into your account, because the trust policy alone can't tell which of the vendor's own customers is actually asking. An External ID is unnecessary for a Role assumed purely within your own Organization (IAM Identity Center Permission Sets, member-account access); it specifically targets the third-party-vendor scenario.
An Access Key has no permissions of its own. It's just a credential pair (Access Key ID + Secret Access Key) that proves "I am this User". The actual permissions come entirely from whatever Policies are attached to that User, directly or via a Group. Two keys belonging to the same User are equally powerful; rotating or deactivating a key changes what can authenticate, not what the User is allowed to do.
Two ways to get credentials, and they're not interchangeable. (1) Long-term: a User's Access Key ID + Secret Access Key, valid until rotated/deleted, permissions = that User's attached Policies. (2) Temporary: a Role is assumed via STS (AssumeRole), which hands back a short-lived Access Key ID + Secret Access Key + session token that auto-expires, permissions = that Role's attached Policies. Same credential shape (a key pair), completely different lifecycle and source of permissions. This is the exam distinction, not "Users have keys, Roles don't."
IAM Role (temporary)Access Key (User, long-term)
Where credentials live Nowhere persistent: for an AWS service (EC2 instance profile, Lambda execution role) they're fetched from instance metadata / injected at invoke time and held only in memory; never written to disk Typically ~/.aws/credentials on disk, an environment variable, or (worst case) hardcoded in app config/source, a real, repeatedly-exploited leak surface: committed to a public repo, left on a compromised laptop, pulled via an IMDSv1 SSRF bug
Lifetime Short-lived, auto-expires (as little as 15 min, default 1 hr, up to 12 hr for most AssumeRole calls); a leaked one is only useful for a limited window Indefinite: stays valid until someone manually deactivates or deletes it, which is exactly why forgotten/leaked keys are such a common real-world breach vector
Rotation None needed: a brand-new set is issued every time the Role is assumed Manual, and easy to forget (this is AWS's official best-practice topic #4, below)
Setup cost More upfront config: a trust policy (who/what may assume it) plus the permission policies themselves One click to generate: simple, which is also why it's overused where a Role would be safer
Best for AWS services acting on your behalf (EC2, Lambda, ECS), cross-account access, federated/SSO human logins, anywhere the caller can be handed a credential rather than typing one in Genuinely long-term, non-interactive use where nothing can assume a Role for you (some legacy tools/third-party integrations with no role-assumption support)
The "no keys stored" win is strongest for AWS services, not for a human at a CLI. An EC2 instance profile or Lambda execution role genuinely never has any static credential anywhere. That's the classic exam scenario above. A person running the CLI on their own laptop still needs to authenticate the AssumeRole call itself somehow: either with a User's long-term Access Key sitting in ~/.aws/credentials as a source_profile (which only partially avoids the problem: the day-to-day CLI credentials are temporary, but a static key still exists behind the scenes) or, cleaner, via IAM Identity Center / SSO (see below), which needs no long-term key at all. Don't over-claim "Roles mean nothing is ever in ~/.aws" for the human case. It depends on how the Role is being assumed.
πŸ“œ Policies
JSON documents with statements: Allow or Deny on specific Actions + Resources.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Allow",
      "Action": "s3:PutObject",
      "Resource": "arn:aws:s3:::my-bucket/*"
    }
  ]
}
Version: the policy language version, always 2012-10-17 (the current/only version; don't overthink it). Statement: an array of one or more permission blocks, each with an Effect (Allow/Deny), one or more Action(s) (the API action(s) being granted, e.g. s3:PutObject), and one or more Resource(s) (the ARN(s) it applies to, wildcards allowed).
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Effect": "Deny",
      "Action": "s3:DeleteObject",
      "Resource": "arn:aws:s3:::my-bucket/*"
    }
  ]
}
Same shape, just Effect: Deny. This blocks s3:DeleteObject on the bucket no matter what any other Allow statement says, on this policy or any other attached to the identity. This is the "Explicit Deny always wins" rule below in practice.
Identity-based Attached to a User/Group/Role
Resource-based Attached to the resource itself (e.g. S3 bucket policy)
AWS Managed Created/maintained by AWS, can't edit
Customer Managed You create it, reusable across users/groups/roles in your account
Inline Embedded 1:1 in a single User/Group/Role; deleted when that entity is deleted
AWS recommends Customer Managed over Inline in most cases.
IAM Policy Simulator tests what a policy actually allows/denies before you attach it; IAM Access Analyzer continuously scans your account and flags unintended external/public access.
βš–οΈ Evaluation Logic (frequently tested!)
1 Default = implicit deny (nothing is allowed unless stated)
2 Explicit Allow overrides the implicit deny
3 Explicit Deny always wins: overrides any Allow, no matter where it comes from
Request made Explicit DENY in any policy? Yes 🚫 DENY No Explicit ALLOW in any policy? Yes βœ… ALLOW No 🚫 DENY (implicit default) Explicit Deny is checked first and short-circuits everything else. This is why "Deny always wins" regardless of how many Allow statements exist.
Advanced: A Permissions Boundary sets the maximum permissions an entity (User/Role) can have, no matter what its identity policies allow. Org SCPs work the same way but at the account/OU level. More relevant at Pro level, but good to recognize.
πŸ–₯️ Roles + EC2 (classic exam scenario)
EC2 instance needs an Instance Profile containing exactly one Role.
App on the instance retrieves temporary credentials automatically from instance metadata, no keys to manage.
This is always preferred over storing a User's Access Keys on the instance.
EC2 Instance Instance Profile IAM Role (1 per profile) AWS STS AssumeRole Temporary Credentials auto-expire, no keys stored No Access Keys ever touch the instance; the app fetches short-lived credentials from instance metadata and AWS handles the rest.
Terminology note: the temporary credentials STS hands back are technically the same shape as a long-term Access Key: an Access Key ID + Secret Access Key, plus a session token. "No Access Keys" above means no long-term User Access Key is created, stored, or exposed on the instance, not that the credential format is different. See the "Two ways to get credentials" note under Identities above.
The launching User and the running instance are two separate principals: a classic gotcha. Creating the EC2 instance only needs ec2:RunInstances (and related) permissions on your User/Role. Once it's running, anything the instance itself does (an app calling S3, a script calling DynamoDB, or you SSHing in and running the AWS CLI from inside it) authenticates as the instance profile's Role, not as the User who launched it. Your own permissions don't carry over onto the box at all. This is why a full-admin User can launch an instance whose app then gets AccessDenied calling S3; the Role attached to the instance simply wasn't granted s3:*, regardless of what the launching User could do.
You can bridge that gap with a User's Access Key, but it's an anti-pattern. It's technically possible to put a User's long-term Access Key ID + Secret Access Key on the instance (env vars, aws configure, a config file) so the instance authenticates as that User instead of via its Role. This is exactly the "storing a User's Access Keys on the instance" practice called out as never-preferred above: it reintroduces a static, leakable, manually-rotated credential on a box that could have had none at all. The exam-correct answer for "an EC2 instance needs AWS access" is always an Instance Profile Role; a hardcoded Access Key on the instance is the wrong-answer distractor.
πŸ” Authentication Methods
Access viaCredential
Console Username + Password (+ MFA recommended)
CLI / API (User) Access Key ID + Secret Access Key (long-term; avoid where possible)
CLI / API (Role) Temporary credentials via STS (Security Token Service): short-lived, auto-expire
Every action you take (console click, CLI command, SDK call) is an API call. The Console is just a UI wrapper around the same AWS APIs the CLI/SDK use directly. This is why IAM policies grant/deny specific Actions (e.g. s3:PutObject) rather than "console access" vs "CLI access". The permission is on the underlying API action, not the interface used to trigger it. It's also why CloudTrail can log every single one, regardless of which interface made the call.
🏒 IAM Identity Center (formerly AWS SSO)
Solves a different problem than everything above: one login for a workforce across many AWS accounts (and even non-AWS apps), instead of a separate IAM User in every account.
What it replacesCreating individual IAM Users in each member account of an Organization; instead, people sign in once and get a portal listing every account/role they're allowed into
Permission SetsReusable templates of permissions assigned to a user/group per-account. Under the hood these provision IAM Roles in the target account, so it's built on the same Role mechanics as above, just centrally managed
Identity sourceIts own built-in directory, or federate from an external IdP (Microsoft AD, Okta, Azure AD, etc.) via SAML
Exam signal: a question about one person needing access to multiple AWS accounts without a separate User in each one is pointing at IAM Identity Center, not at creating more IAM Users or hand-rolling cross-account Roles yourself. It's an AWS Organizations-level feature, not something you turn on inside a single account in isolation.
πŸ›οΈ AWS Organizations & Service Control Policies (SCPs) (frequently tested!)
AWS Organizations groups multiple AWS accounts (e.g. one per team, environment, or department) under one management account, arranged into Organizational Units (OUs), folders you can nest and apply policy to as a group instead of account-by-account.
A Service Control Policy (SCP) is a guardrail attached to the whole Organization, an OU, or a single account. It sets the maximum available permissions for every principal in scope, including that account's own root user/admin. An SCP never grants anything by itself (it has no effect on a management account); it only restricts what IAM policies inside the account are allowed to permit.
IAM PolicyService Control Policy (SCP)
Attached toA User, Group, or RoleAn AWS account, an OU, or the whole Organization
Can it grant access?Yes, an Allow here actually grants permissionNo, it only sets a ceiling; an Allow in an SCP does nothing on its own without a matching IAM policy Allow
Affects the account's root user?No, Users/Roles onlyYes, even the account root/admin cannot exceed an SCP's ceiling
Typical useDay-to-day least-privilege access for a specific identityOrg-wide guardrails: block a whole category of action everywhere it applies, regardless of any individual permissions
Final permission = the intersection of both. A principal can only actually do something if an IAM policy allows it and no applicable SCP blocks it. An SCP Allow plus an IAM policy Deny still results in Deny (same "explicit Deny always wins" rule as above, just evaluated across two separate policy layers instead of one).
Common exam SCP patterns: deny cloudtrail:StopLogging org-wide so no admin in any account can quietly disable audit logging; deny ec2:RunInstances above a certain instance size in a sandbox/dev OU to cap accidental cost; deny actions outside an approved AWS Region.
Exam pattern: "must apply even if the user/account has full admin/root privileges, with the least operational overhead across many accounts" β†’ an SCP at the OU level, not a per-account IAM policy repeated in every account. That's the "least overhead across an Organization" signal.
βœ… Best Practices: the ones that actually get tested
Lock away the root account: don't use it day-to-day; create an admin IAM User instead.
Enable MFA, especially for privileged/root accounts.
Least privilege: grant only what's needed.
Prefer Roles over long-term Access Keys wherever possible.
Use Groups to manage User permissions at scale, not per-user policies.
Rotate credentials regularly; remove unused ones.
Full official list: AWS's Security best practices in IAM guide: 14 topics, summarized below.
πŸ“˜ AWS's Official 14 IAM Best Practice Topics
1Require human users to use federation with an identity provider for temporary credentials
2Require workloads to use temporary credentials with IAM Roles
3Require multi-factor authentication (MFA)
4Update access keys when needed, for use cases that genuinely require long-term credentials
5Follow best practices to protect your root user credentials
6Apply least-privilege permissions
7Get started with AWS managed policies, then move toward least privilege
8Use IAM Access Analyzer to generate least-privilege policies based on access activity
9Regularly review and remove unused users, roles, permissions, policies, and credentials
10Use conditions in IAM policies to further restrict access
11Verify public and cross-account access to resources with IAM Access Analyzer
12Use IAM Access Analyzer to validate your IAM policies for secure and functional permissions
13Establish permissions guardrails across multiple accounts (Org SCPs/RCPs)
14Use permissions boundaries to delegate permissions management within an account
🧠 Quick Memory Hooks
User = person  Β·  Group = filing cabinet for people  Β·  Role = borrowed hat  Β·  Policy = rulebook
Deny always beats Allow.
No standing credentials = Role = best practice.
🌐 VPC & Networking: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
A VPC is your own isolated network inside a Region; everything else on this page is about controlling what can reach what, inside it and in/out of it. See the EC2 & Compute topic for the subnet/AZ/ENI layout diagram. This page covers the routing and filtering layer on top of it.
πŸ“ VPC Basics: CIDR sizing
A VPC is defined by a CIDR block (e.g. 10.0.0.0/16), the range of private IPs available inside it. Size: /16 to /28 (biggest to smallest).
Every subnet's CIDR must be a subset of its VPC's CIDR, and subnets in the same VPC can't overlap each other.
AWS reserves the first 4 and last 1 IP addresses in every subnet (network address, VPC router, DNS, future use, broadcast). A /24 subnet (256 addresses) really only gives you 251 usable.
A VPC can span multiple AZs (via multiple subnets) but always lives in exactly one Region.
πŸ—ΊοΈ Subnets & Route Tables
Every subnet is associated with exactly one route table (though one route table can serve many subnets). Unassociated subnets fall back to the VPC's main route table.
Every route table always has an implicit local route (traffic within the VPC's own CIDR) that can't be removed. This is what lets resources in different subnets of the same VPC reach each other by default.
A subnet counts as "public" purely because its route table has a route to an Internet Gateway; nothing else about the subnet changes.
Route table per AZ, for HA: a subnet lives in exactly one AZ, so routing each AZ's private subnet to that AZ's own NAT Gateway means giving each AZ its own route table. One shared route table pointing every private subnet at a single AZ's NAT Gateway would defeat the point of deploying one NAT Gateway per AZ.
πŸšͺ Internet Gateway vs. NAT Gateway vs. NAT Instance
Internet GatewayNAT GatewayNAT Instance
PurposeTwo-way internet access for a public subnetOutbound-only internet for a private subnetSame as NAT Gateway, but you manage it
Managed byAWS (attach to VPC, no config)AWS (provisioned, patched, scaled for you)You; it's just an EC2 instance running NAT software
AvailabilityN/A (not a bottleneck resource)Lives in one AZ; deploy one per AZ for HASingle instance; you handle HA yourself
BandwidthN/AScales automatically up to 100 GbpsCapped by the instance type you chose
Exam pattern: "instances in a private subnet need outbound internet only, minimal management" β†’ NAT Gateway, essentially always the right answer over a NAT Instance now.
Exam pattern (HA, frequently tested!): "design must survive an AZ failure" + a NAT Gateway in the design β†’ one NAT Gateway per AZ, each with its own private subnet's route table pointing to it. A single shared NAT Gateway is a single point of failure for every other AZ if its own AZ goes down.
πŸ›‘οΈ Security Groups vs. NACLs (frequently tested!)
🌍 Internet Subnet boundary: NACL (stateless) Instance boundary: Security Group (stateful) EC2 Instance Inbound Outbound return traffic auto-allowed by SG Every packet crosses the NACL first (subnet edge), then the Security Group (instance edge). Both must allow it.
Security GroupNACL
Applies toThe instance/ENI (instance-level firewall)The subnet (network-level firewall). Affects every instance in it
RulesAllow only (no explicit Deny)Allow and Deny rules, evaluated in rule-number order, lowest first
StateStateful: return traffic is automatically allowed, regardless of outbound rulesStateless: inbound and outbound must each be explicitly allowed, even for a reply
EvaluationAll rules evaluated together; any match winsRules evaluated in order until the first match; a Deny can be hit before an Allow further down
DefaultNew SG: all inbound denied, all outbound allowedDefault NACL: allows all in/out; a custom NACL denies all until you add rules
VPC Flow Logs: captures metadata (source/dest IP, port, protocol, ACCEPT/REJECT, not the packet contents) for traffic at the VPC, subnet, or ENI level, sent to CloudWatch Logs or S3. The go-to tool for "why is my traffic being blocked": diagnosing which layer (SG or NACL) is actually rejecting a connection.
Two different "defaults": don't mix them up. The table's "Default" row describes a new SG you create: all inbound denied, all outbound allowed. The default security group that AWS auto-creates with every VPC is different again: it also allows all outbound, but its inbound rule is a self-reference (allow all traffic from any other resource that has that same default SG attached), not a flat deny. In practice this means two freshly launched instances left on the default SG can already talk to each other on every port, which surprises people expecting "deny all inbound" to be universal.
A Security Group rule's source/destination can be another Security Group, not just a CIDR (frequently tested!) This is the standard way to wire up a multi-tier design: the app-tier SG's inbound rule allows traffic from the web-tier SG (not from a CIDR range), and the database-tier SG's inbound rule allows traffic from the app-tier SG. Each tier only ever accepts traffic from the specific tier directly in front of it. Because the reference is to the group, not to fixed IPs, it stays correct automatically as instances in that tier scale in/out under an ASG, no CIDR range to keep updating. This is the best-practice answer whenever a question describes least-privilege network access between tiers, over widening a CIDR-based rule.
πŸ”— VPC Peering
A private, 1:1 network connection between two VPCs (same or different accounts/Regions). Traffic stays on the AWS backbone, never touches the public internet.
Not transitive: if VPC A peers with B, and B peers with C, A cannot reach C through B; you'd need a direct A↔C peering (or a Transit Gateway, at Pro-level scale).
The two VPCs' CIDR blocks must not overlap. Peering can't be established (or route tables can't be added) if they do.
Each side must add a route to the other's CIDR in its own route table; peering existing alone doesn't route traffic.
🎯 VPC Endpoints
Lets resources in a private subnet reach an AWS service without going through an Internet Gateway or NAT; traffic never leaves the AWS network.
Gateway EndpointOnly for S3 and DynamoDB. Free. Works by adding a target in your route table, no ENI involved.
Interface EndpointFor most other AWS services, powered by AWS PrivateLink. Creates an ENI with a private IP in your subnet. Small hourly + data cost.
Exam pattern: "private subnet, needs to reach S3, no internet access at all" β†’ Gateway Endpoint (free) is the expected answer over a NAT Gateway (which costs more and routes through a public subnet's IGW anyway).
Scope limitation (frequently tested!) A VPC Endpoint (Gateway or Interface) only ever connects to an AWS service, or another VPC's own resource deliberately exposed as a PrivateLink endpoint service. It cannot reach an arbitrary third-party endpoint out on the public internet (an external payment processor's API, a SaaS vendor's REST API, etc.). For that, a private subnet still needs a NAT Gateway or NAT Instance routing out through an Internet Gateway. A VPC Endpoint is not a general substitute for NAT, only for the specific AWS services it supports.
πŸ”Œ Direct Connect & Site-to-Site VPN (frequently tested!)
Both connect an on-premises network to a VPC. The difference is what kind of link, not what it's used for.
Site-to-Site VPNDirect Connect
What it isAn encrypted IPsec tunnel over the public internetA private, dedicated physical network connection from your premises into AWS, never touches the public internet
Setup timeMinutes, fully software-configuredWeeks to months; needs a physical cross-connect at an AWS Direct Connect location (colo facility or partner)
PerformanceVariable, subject to public internet congestion/latencyConsistent, low-latency, high-throughput: a real dedicated line
Cost modelLower: pay for the VPN connection + data transferHigher fixed cost for the dedicated line, but often cheaper data transfer at scale
Typical useQuick to stand up, backup link, lower-volume/less latency-sensitive trafficConsistent large data volumes, latency-sensitive workloads, a strict "must not traverse the public internet" compliance requirement
Exam pattern: "must not traverse the public internet at all" or "consistent, dedicated bandwidth" β†’ Direct Connect. "Quick to set up, encrypted, lower/no upfront hardware" β†’ Site-to-Site VPN. The two are also commonly paired: a VPN as an automatic failover path if the Direct Connect link goes down.
🧠 Quick Memory Hooks
NACL = bouncer at the building door (subnet, stateless, checks you both ways). Security Group = bouncer at your apartment door (instance, stateful, remembers who let you in).
Peering is a one-hop friendship, not a chain.
Gateway Endpoint = free, S3/DynamoDB only. Interface Endpoint = everything else, small cost, uses PrivateLink.
πŸ–₯️ EC2 & Compute: One-Page Study Guide
Pick an instance type for the workload, a purchasing option for the billing commitment, and a storage type for the data, three independent choices. Security Groups/NACLs live in VPC & Networking; scaling policies live in ELB & Auto Scaling, both separate topics.
🏷️ Instance Type Families
FamilyOptimized forExample use case
General Purpose (M, T, A)Balanced compute/memory/networkWeb servers, small-to-medium apps, dev/test
Compute Optimized (C)High vCPU-to-memory ratioBatch processing, media transcoding, gaming servers, HPC
Memory Optimized (R, X, z)Large RAM per vCPUIn-memory caches, large databases, real-time big-data analytics
Storage Optimized (I, D, H)High, fast local disk I/ONoSQL databases, data warehousing, distributed file systems
Accelerated Computing (P, G, Inf, Trn)Hardware GPU/acceleratorML training/inference, graphics rendering, video encoding
Naming decode (m6g.large): m = family, 6 = generation, g = processor attribute (g=AWS Graviton/ARM, a=AMD, n=network optimized, no letter=Intel), large = size. Higher generation number β‰ˆ newer, better price/performance. The exam often expects "pick the newer generation" as the answer when everything else ties.
T-family only is burstable: earns CPU credits at baseline, spends them to burst above baseline; Unlimited mode lets it keep bursting past its credit balance for a small extra charge instead of throttling back to baseline.
Full official list of families, generations, and specs: AWS's Amazon EC2 instance types guide.
πŸ” Decoding an Instance Type: worked example
m5.largem = General Purpose family  Β·  5 = generation 5  Β·  large = instance size (2 vCPU / 8 GiB RAM)
r5.larger = Memory Optimized family  Β·  5 = generation 5  Β·  large = instance size (2 vCPU / 16 GiB RAM; same vCPU count as m5.large, double the RAM)
Common mix-up: the leading letter is the tell for the family, not the number: m5.large is General Purpose (the "m" is short for a balanced, "middle-of-the-road" mix of compute/memory), not memory-optimized. If a question is testing "high RAM per vCPU," the letter to look for is r (or x/z for even more extreme ratios). See the family table above.
πŸ—ΊοΈ Where EC2 Lives: VPC, Subnets & ENIs
Terminology: Region vs. Availability Zone (AZ). A Region is a large geographic area (e.g. us-east-1, eu-west-2) that's fully independent of every other Region: its own copy of most services, its own data, nothing replicates between Regions unless you set that up yourself. Each Region is made up of multiple Availability Zones, one or more physically separate data centers within that Region, far enough apart to survive an independent failure (power, fire, flooding) but close enough together for low-latency links between them. Why it's tested: "high availability" almost always means spread across multiple AZs in one Region; "disaster recovery" almost always means spread across multiple Regions.
Every EC2 instance launches into a VPC, which spans a whole Region, inside one specific Availability Zone (AZ), inside one specific subnet in that AZ. There's no such thing as an instance outside a VPC, and a subnet never crosses AZ boundaries.
Public subnetIts route table has a route to an Internet Gateway (IGW). Instances here can hold both a private IP and a public IP/Elastic IP, and reach (and be reached from) the internet directly
Private subnetNo route to an IGW; instances here only ever get a private IP, never a public one; outbound-only internet access needs a NAT Gateway sitting in a public subnet
🌍 Internet Internet Gateway (IGW) VPC Β· 10.0.0.0/16 (Region: us-east-1) Availability Zone A Public Subnet (10.0.1.0/24) Route: 0.0.0.0/0 β†’ IGW EC2 Instance eth0 (primary) 10.0.1.10 (private) eth1 (additional) 10.0.1.11 (private) + 3.10.55.2 (public/EIP) hot-attach/detach NAT Gateway Private Subnet (10.0.2.0/24) No route to IGW EC2 Instance eth0 (primary) 10.0.2.10 (private only, no public IP) 0.0.0.0/0 β†’ IGW outbound-only, via NAT A subnet decides what a private IP can reach; a public/Elastic IP association decides whether an ENI is also reachable from the internet.
Default VPC: AWS auto-creates one per Region on your behalf. It comes pre-wired with a public subnet in every AZ, an IGW already attached, and a route table already sending 0.0.0.0/0 β†’ IGW. Convenient for quickly launching a test instance with internet access, no networking setup required.
Custom VPC: you build the CIDR range, subnets, and route tables yourself. Internet connectivity is not automatic; you must create and attach an IGW, then add a route to it in the relevant subnet's route table before that subnet counts as "public." This is the standard, best-practice setup for real workloads (lets you control exactly which subnets are public vs. private).
Public IP by default: depends on which VPC you're in. Every subnet has an "auto-assign public IPv4" setting that decides whether a launched instance gets a public IP with no extra step. In the default VPC, every subnet has this switched on. Launch an instance with default settings and it gets a public IP automatically, alongside the default SG's permissive self-reference rule above, which is why a brand-new default-VPC instance is reachable from the internet (if you also allow the right port) with zero networking configuration. In a custom VPC, auto-assign public IPv4 is off by default on any subnet you create; you must enable it on the subnet (or attach an Elastic IP after launch) for an instance there to get a public IP at all.
ENI = think "virtual NIC." Every instance gets a primary ENI automatically at launch, fixed to the subnet you launched into (and so to that subnet's AZ). It always carries a private IP, and in most cases can't be detached while the instance is running. You can also create and hot-attach additional ENIs to a running instance for a second network presence (e.g. a separate management interface, or fast failover: detach the ENI from a failed instance and attach it to a standby, keeping the same IP). An ENI is "public" only in the sense that it sits in a public subnet and has a public IPv4/Elastic IP associated with it. The ENI object itself doesn't have a fixed public/private type.
Full detail on subnets, route tables, IGW/NAT, and security groups vs. NACLs lives in the VPC & Networking topic. This section is just the "where does my instance actually sit" mental model for Compute.
🏷️ Public IP vs. Private IP vs. Elastic IP
TypePersistenceCostTypical use case
Private IPv4Persists for the life of the ENI (survives stop/start)FreeInternal-only traffic: DB tier, backend services, anything reached only via a load balancer/NAT/VPN
Public IPv4 (auto-assigned)Not persistent: released and re-assigned to a new random address on every stop/startBilled hourly (all public IPv4 addresses, attached or not)Quick/throwaway dev-test instances where the IP itself doesn't matter
Elastic IPStatic: yours until you release it; remap between instances/ENIs on demandBilled hourly, same rate as any public IPv4 (see note), but still a soft-limited resource (5 per account by default), so release ones you're not usingProduction internet-facing endpoints needing a stable, memorable address: a DNS record, a partner's IP allowlist, or a fast-failover target
Pros/cons in one line each: Private IP: free and stable, but never internet-reachable on its own. Public IP (auto): zero setup, but changes every restart, so nothing should hardcode it. Elastic IP: the only one that's both stable and reachable, but it's a finite, deliberately-nudged-away-from resource (a soft limit of 5 per account by default); don't hold one you're not using.
Billing gotcha (frequently tested!): since Feb 2024, AWS bills every public IPv4 address (auto-assigned or Elastic), a small hourly rate, whether it's attached to a running resource or not. Before that change, an Elastic IP was only charged while unattached (to discourage hoarding) and an auto-assigned public IP was free. The "avoid unattached EIPs" instinct from older material is still good practice, but it's no longer the only cost driver. Minimizing how many public IPv4 addresses you provision at all is now the bigger lever.
Why the instance can't see its own public/Elastic IP: that address is never actually on the instance's ENI. The ENI only ever carries the private IP. The public IP is a 1:1 NAT mapping held at the Internet Gateway, translated on the way in/out; the instance only learns it by asking instance metadata or an external service.
πŸ”Œ ENI vs. ENA vs. EFA (don't mix these up)
ENIElastic Network Interface, the virtual NIC itself: its identity (private/public IPs, MAC address, security groups). What you attach/detach.
ENAElastic Network Adapter, the high-performance driver/hardware behind an ENI on modern instance types, enabling "enhanced networking" up to 100 Gbps. Almost every current-generation instance uses this by default.
EFAElastic Fabric Adapter, an ENA variant that adds an OS-bypass hardware interface for ultra-low-latency, tightly-coupled inter-node traffic. Built for HPC/distributed ML training (MPI-style workloads), typically paired with a Cluster Placement Group.
Mental model: ENI is the network card you can see and manage (IP/MAC/security groups). ENA is what makes that card fast. EFA is a specialized version of ENA for the narrow case of HPC nodes that need to talk to each other with minimal latency, not something you'd pick for a typical web app.
Why ENA/EFA are possible at all (the Nitro System): the underlying hardware/hypervisor platform for the next generation of EC2 instances. Virtually every current-generation instance type runs on it, offloading networking, storage, and management functions to dedicated Nitro cards instead of the host CPU. Performance win: with almost nothing left for a traditional hypervisor to do on the host CPU, virtually all of that CPU/RAM is handed to your instance instead. This is what delivers near-bare-metal performance and enables enhanced networking (ENA/EFA) up to 100 Gbps. Nitro Enclaves is the exam-relevant spin-off: an isolated, hardened compute environment carved out of an instance with no persistent storage, no interactive access, and no external networking, for processing highly sensitive data (PII, cryptographic keys) with a minimized attack surface.
πŸ’³ Purchasing Options (frequently tested!)
OptionCommitmentDiscount vs. On-DemandBest for
On-DemandNoneBaseline (0%)Short-term, spiky, unpredictable workloads; first-time/unknown-duration apps
Reserved Instances1 or 3 yearsUp to ~72%Steady-state, predictable usage on a specific instance family/region
Savings Plans1 or 3 years, $/hr spend commitmentUp to ~72%Same as Reserved but flexible across instance family/size/OS/region (Compute Savings Plans) or just size within a family (EC2 Instance Savings Plans)
Spot InstancesNone (can be reclaimed)Up to ~90%Fault-tolerant, flexible, interruptible workloads (batch, CI, stateless web tiers)
Dedicated HostsOn-Demand or 1/3-yr ReservationVariesCompliance/licensing needs a physical server mapped to you (BYOL, per-socket/core licensing)
Dedicated InstancesNone or ReservationVariesPhysical isolation from other accounts, but no control over host placement
Capacity ReservationsNone (no term); reserves capacity, not discountNone by itself (pair with RI/Savings Plan for a discount)Guarantee capacity is available in a specific AZ, e.g. for a known future launch/DR
Reserved: Standard vs. Convertible. Standard = bigger discount, can't change instance family. Convertible = smaller discount, can swap instance family/OS/tenancy during the term.
Spot gotcha: AWS can reclaim a Spot instance with a 2-minute warning (via CloudWatch event / instance metadata) when it needs the capacity back. Never use Spot for anything that can't tolerate sudden termination or can't checkpoint its work.
πŸ” Savings Plans vs. Reserved Instances: closer look
AttributeReserved InstancesSavings Plans
What the commitment attaches toA specific instance family + Region/AZ + OS + tenancy (Standard); Convertible can swap family/OS/tenancyA $/hr spend commitment; applies automatically to any matching usage, no instance-level binding
FlexibilityStandard: none. Convertible: family/OS/tenancy, but stays EC2-onlyCompute Savings Plans: any instance family/size/OS/Region, plus Fargate and Lambda. EC2 Instance Savings Plans: size/OS/tenancy flexible, but locked to one family + Region
Payment optionsAll/Partial/No UpfrontAll/Partial/No Upfront
Capacity guaranteeZonal RI reserves actual capacity in a specific AZ; Regional RI does notNone: never guarantees capacity, purely a billing discount
Resale if plans changeUnused RIs can be sold on the AWS Reserved Instance Marketplace to recoup costNo resale mechanism; you're committed for the term
Best forA known, fixed instance configuration you won't change, especially if you also need a capacity guarantee (Zonal)Steady-state spend across a mix that may shift over time (instance types, or even EC2 β†’ Fargate/Lambda)
Exam pattern: "steady-state usage, but the instance mix might change" or "want the discount to also cover Fargate/Lambda" β†’ Savings Plans (Compute). "Need a capacity guarantee in a specific AZ" or "might resell the commitment later" β†’ Reserved Instances.
πŸ’° You're Billed for What You Provision, Not What You Use (frequently tested!)
For most EC2-adjacent resources, AWS charges for the size/capacity you provisioned, regardless of how much of it your workload actually consumes. Idle capacity still costs money. This is exactly why "rightsizing" is a real, tested cost-optimization lever, not just a nice-to-have.
Ex. 1Launch an m5.xlarge (4 vCPU / 16 GiB) but your app only ever uses ~10% CPU and 2 GiB RAM β†’ you're still billed the full m5.xlarge hourly rate, not 10% of it.
Ex. 2Provision a 500 GB gp3 EBS volume but only store 50 GB of actual data on it β†’ you're billed for all 500 GB provisioned, every month, not the 50 GB used.
Tools that help catch this: AWS Compute Optimizer (recommends better-fitting instance types/sizes from actual utilization history) and Trusted Advisor (flags idle/over-provisioned resources). Downsizing an over-provisioned instance or shrinking an oversized volume is one of the most common "how do I reduce cost" exam answers.
πŸ’Ύ EBS vs. Instance Store
Terminology: "Ephemeral." AWS's own name for instance store is ephemeral storage. Ephemeral just means short-lived/temporary. It's the exam's go-to vocabulary word for "this data disappears the moment the instance stops or terminates". If a question describes something as ephemeral, it's pointing you at instance store, not EBS.
EBSInstance Store
PersistenceSurvives stop/start & instance termination (if not set to delete)Physically attached to the host, lost on stop or termination
PerformanceNetwork-attached (slightly higher latency)Physically local, highest possible IOPS/throughput
FlexibilityDetach/reattach to another instance, resize, snapshotFixed to the instance's lifecycle, can't detach
Use caseBoot volumes, databases, anything needing durabilityCache, buffer, scratch/temp data, data already replicated elsewhere
Rule of thumb: if data must survive the instance, it belongs on EBS (or S3/EFS), never instance store.
EBS lives in one AZ. A volume is created inside a single Availability Zone and is automatically, redundantly replicated within that AZ by AWS for durability. That replication is what protects against a single hardware failure. It's also why a volume can only attach to an instance in the same AZ, and why moving one to a different AZ (or Region) requires taking a snapshot first (snapshots live in S3, which is Region-wide) and restoring a new volume from it there.
What it looks like from inside the instance: once attached, an EBS volume just shows up as an ordinary local disk to the OS: a drive letter like D:\ on Windows, or a device like /dev/xvdf on Linux. The instance has no idea it's talking to network-attached storage under the hood.
πŸ“€ EBS Volume Types
gp3General purpose SSD: default choice; IOPS/throughput provisioned independently of size. Use for: boot volumes, dev/test, most everyday app and small-to-medium database workloads
gp2Older general purpose SSD: IOPS scales with volume size (3 IOPS/GB). Use for: legacy volumes not yet migrated to gp3; rarely the right pick for anything new
io1 / io2Provisioned IOPS SSD: highest performance, mission-critical low-latency DBs; io2 Block Express for the largest scale. Use for: large relational DBs (Oracle, SQL Server, SAP HANA) with sustained heavy IOPS and sub-millisecond latency needs
st1Throughput Optimized HDD: big sequential workloads; cannot be a boot volume. Use for: big data/Hadoop, data warehouses, log processing; anything reading large sequential chunks fast, not small random reads
sc1Cold HDD: lowest cost, infrequently accessed data; cannot be a boot volume. Use for: archive-style data, infrequently-accessed backups (cost matters more than speed)
Exam pattern: "cheapest option for infrequent access" β†’ sc1. "Boot volume" β†’ must be SSD (gp2/gp3/io1/io2), never st1/sc1. "Need consistent sub-millisecond latency for a DB" β†’ io1/io2.
EBS Multi-Attach: io1/io2 only. Normally an EBS volume attaches to exactly one instance at a time. io1/io2 can opt into Multi-Attach: the same volume attached to up to 16 instances at once, but only instances in the same AZ. It doesn't give you a shared filesystem for free. Without a cluster-aware filesystem (or your app coordinating writes itself), multiple instances writing to the same blocks will corrupt data. Used for cluster-aware apps designed for shared block storage.
πŸ“Έ EBS Snapshots
A snapshot is a point-in-time, incremental backup of an EBS volume, stored durably in S3 behind the scenes (you don't manage the bucket; it's not visible in your S3 console). Incremental means the first snapshot copies every used block, and every snapshot after that only stores the blocks that changed since the last one, but each snapshot still restores to a complete, standalone volume; deleting an older snapshot in the chain doesn't break the newer ones.
Unlike the volume itself, a snapshot is Region-wide, not AZ-locked. This is exactly what makes it the mechanism for moving a volume's data across AZs or Regions (see the AZ-locked note above): snapshot the volume, then restore a new volume from that snapshot in whichever AZ (same Region) you need, or copy the snapshot to another Region first if you need to cross Regions.
Backup / DRScheduled snapshots (e.g. via Amazon Data Lifecycle Manager or AWS Backup) protect against accidental deletion, corruption, or a bad deploy; restore a new volume from any snapshot in the chain
Move across AZ/RegionSnapshot in the source AZ β†’ restore a volume from it in the target AZ (same Region), or copy the snapshot to another Region first for cross-Region moves
Building an AMIAn AMI's block device mapping is literally a set of references to EBS snapshots. Creating a custom AMI creates the underlying snapshot(s) for you
Resize/retype a volumeRestore a snapshot into a larger volume, or onto a different volume type (e.g. gp2 β†’ gp3), instead of resizing the original in place
Frequently tested: a snapshot can be taken while the volume is in use, but for a fully consistent backup (especially a root/boot volume) best practice is to stop the instance first, or at minimum flush filesystem writes. An "in flight" write not yet flushed to disk at snapshot time can leave the restored volume in a slightly inconsistent state. A snapshot of an encrypted volume is automatically encrypted too, and copying a snapshot lets you encrypt an originally-unencrypted one along the way.
πŸ”’ EBS Encryption
An encrypted EBS volume protects data at rest on the volume, data in transit between the instance and the volume, and every snapshot made from it (plus any new volume restored from those snapshots), all using AES-256 under an AWS KMS key. It's fully transparent: the OS and your application see a normal disk, with encrypt/decrypt happening on the underlying host hardware, at negligible performance cost.
You can't flip encryption on an existing volume directly. The path is: snapshot the unencrypted volume β†’ copy that snapshot with encryption enabled (this is the one step where you can turn it on) β†’ create a new volume from the encrypted copy β†’ attach the new volume in place of the old one. Same trick works for changing which KMS key protects a volume.
Compliance / regulated dataPCI-DSS, HIPAA, GDPR-style requirements for encryption at rest, e.g. a volume holding payment records or health data
Org-wide "encrypt by default"Enable default encryption per Region so every new volume/snapshot is encrypted automatically, without anyone remembering to tick a box
Safer snapshot sharingAn encrypted snapshot can never be made public, a built-in guardrail against accidentally exposing sensitive data when sharing/copying snapshots across accounts
Key choice: the default AWS-managed key (aws/ebs) works with zero setup; a Customer Managed Key (CMK) via KMS costs a little more but gives you rotation control, a separate key per project/environment, and the ability to revoke access (or delete the key) to instantly make the data unreadable.
πŸ“¦ AMIs (Amazon Machine Images)
Terminology: AMI. An AMI is a template for launching an instance. Think of it like a "frozen disk image" or a phone's factory restore image, not a running server. Example: you configure one EC2 instance by hand (install Nginx, deploy your app code, apply OS patches), then create an AMI from it. That AMI now captures the whole machine at that moment. Launch 20 new instances from it and every one boots up already running Nginx with your app pre-installed, with no setup script needed.
An AMI is a template: OS + installed software + configuration + permissions + a block device mapping (which EBS snapshots become which volumes).
Golden AMI pattern: bake app/config/patches into a custom AMI ahead of time so new instances launch pre-configured, much faster than running a bootstrap script on every launch.
AMIs are region-specific. Copy an AMI to another region to launch instances there.
Sources: AWS-provided, AWS Marketplace, community, or your own (created from an existing instance/snapshot).
πŸ“ Placement Groups
A logical grouping that controls how AWS places instances on underlying hardware. Pick a strategy to optimize for either performance (low latency, high throughput) or fault isolation.
AttributeClusterSpreadPartition
GoalLowest latency, highest throughputMaximize instance isolationIsolate groups of instances
AZ scopeSingle AZMulti-AZ allowedMulti-AZ allowed (within a Region)
Hardware isolationLow (packed together)Highest (one instance per rack)Medium (isolated per-partition)
Instance limitsNone7 instances per AZUp to 7 partitions per AZ; scales to hundreds of instances overall
Risk of correlated failureHigherLowestLower
Typical use caseHPC, big data, tightly-coupled/high-speed appsCritical workloads that must not fail togetherLarge distributed systems (HDFS, Cassandra, Kafka); partition info exposed to the instance for rack-aware replication
Cluster 1 Availability Zone Packed together (lowest latency, shared failure risk) Spread Max 7 per AZ Distinct hardware each (isolated failures) Partition P1 P2 P3 Separate racks per partition (rack-aware apps) Cluster trades fault isolation for speed; Spread and Partition trade some speed back for isolation, at increasing scale.
Can't move instances inExisting instances can't easily be added to a group after the fact; launch new instances into it; relaunching is usually simpler than migrating
Enhanced networkingCluster benefits most when every instance type in the group supports it; mixing types limits the payoff
CostPlacement groups themselves are free; no extra charge to use one
πŸ”„ Instance Lifecycle: Stop vs. Terminate
States: pending β†’ running β†’ either stopping/stopped (Stop) or shutting-down/terminated (Terminate). Reboot isn't a state change at all: it's just an OS restart; instance ID, IPs, and volumes are untouched.
Stop only exists for EBS-backed instances (frequently tested!): an instance-store-backed instance has no EBS root volume to power down onto and preserve. It can only be terminated, never stopped. If a question offers "Stop" as an option for an instance-store-backed instance, that option is wrong.
StopTerminate
Instance IDKept (same instance, can be started again)Gone permanently
Private IPKept (survives stop/start)Released (the primary ENI is deleted with the instance)
Elastic IPStays associated (still billed per the note above)Released back to your account
Root EBS volumeKept, billed as ordinary EBS storage while stoppedDeleted by default, unless DeleteOnTermination was turned off
Instance store dataLost the moment power stopsLost (same reason)
RAM contentsLost, unless Hibernate is enabled, which saves RAM to the root EBS volume first (see Hibernate note below)Lost
Compute billingStops (no per-hour instance charge while stopped)Stops
A stopped instance is the state both Hibernate and Auto Recovery build on top of. See their own notes on this page. Both Stop and Terminate can be gated by a guardrail flag. See Termination & Stop Protection below.
🩺 Status Checks & Auto Recovery (classic troubleshooting scenario)
Every instance runs two independent automated checks every minute. Telling them apart is the whole exam question: they imply completely different fixes.
CheckWhat it's really testingTypical causeFix
System Status CheckThe underlying AWS hardware/hypervisor/network the instance runs onHost power loss, hardware failure, network connectivity issue at the host levelStop/Start the instance. This moves it onto different underlying hardware (a reboot does not help, since it doesn't change hardware)
Instance Status CheckThe instance itself, its OS and network configCorrupted file system, kernel panic, exhausted memory, misconfigured network settingsReboot the instance; the hardware is fine, just the OS needs to restart
EC2 Auto Recovery: a CloudWatch alarm (on the StatusCheckFailed_System metric) that automatically stops and starts the instance for you the moment a System Status Check fails, no human needed to notice and intervene. The recovered instance keeps the same instance ID, private IP, Elastic IP, and EBS volumes; it just lands on healthy hardware. This is EC2-level reliability, distinct from Auto Scaling (which replaces an unhealthy instance with a brand-new one rather than recovering the same one).
Instance Retirement: AWS-initiated, not something you trigger. When the underlying hardware an instance sits on is degrading or scheduled for decommission, AWS schedules a retirement date and notifies you in advance (Personal Health Dashboard/email), giving you a window to Stop/Start the instance yourself (moving it to healthy hardware, same as Auto Recovery does) before AWS does it for you at the retirement date.
πŸ” Termination & Stop Protection
Termination ProtectionA per-instance flag (DisableApiTermination) that blocks a Terminate request from the console/CLI/API until it's switched off, the classic guardrail against accidentally deleting a production instance (and its instance-store data, and any EBS volumes set to delete-on-termination)
Stop ProtectionThe same idea for Stop (DisableApiStop), useful for an instance where even a stop/start (e.g. losing its public IP, or interrupting a long-running process) would be disruptive
Both are just flags you toggle on the instance: no extra cost, no separate service. Neither is a substitute for IAM permissions (a user with ec2:TerminateInstances can still disable the flag first, then terminate). Think of it as a safety catch against accidental clicks, not a security control.
🌐 Instance Metadata & Networking (frequently tested!)
Every instance can query its own metadata (instance ID, AMI ID, IAM Role credentials, user data, etc.) from inside the instance at:
http://169.254.169.254/latest/meta-data/
IMDSv2 vs. IMDSv1: IMDSv1 answers plain GET requests to that address, a classic SSRF target (a vulnerable app can be tricked into fetching it and leaking the instance's Role credentials). IMDSv2 requires a session token (via a PUT request first) before it will answer, closing that SSRF path. AWS/the exam now treats IMDSv2 as the best-practice default.
User DataScript run once, automatically, on first boot, used to bootstrap/configure a new instance (install packages, pull config, join a fleet)
Elastic IPStatic public IPv4 you own and can remap between instances; billed hourly whether attached or not (since Feb 2024); release ones you're not using, not just unattached ones
ENIElastic Network Interface, a virtual NIC (private IPs, security groups, MAC); additional ENIs can be hot-attached/detached between instances for failover. Full breakdown (incl. ENA/EFA) in "Where EC2 Lives" above
HibernateSaves in-memory RAM state to the root EBS volume on stop, restores it on start; faster warm boot than a cold start's OS+app re-init. Constraints: root volume must be encrypted, RAM capped at ~150 GiB, and it's not supported on instance-store-backed instances (nothing to durably save the RAM contents to)
πŸ”Œ Connecting to an Instance (frequently tested!)
Four ways to get a shell on an instance. The exam usually asks "what's the most secure / only way to connect here," so the differences matter more than the mechanics.
MethodRequiresProsConsBest for
EC2 Instance Connect SG allows inbound 22 (or 3389), a supported AMI (Amazon Linux 2/2023 have it preinstalled), IAM permission to push the key No key pair to manage or store; pushes a short-lived, one-time SSH key via the API; works straight from the browser console Still needs an open inbound port and a real network path (public IP, or same-VPC/VPN reachability); doesn't help a fully private instance with no such path Quick one-off browser access to a reachable (public or in-VPC) instance without setting up a local SSH client
Session Manager (Systems Manager) SSM Agent running (preinstalled on most current AMIs), an instance IAM role with SSM permissions, and outbound HTTPS reachability to the SSM endpoints (public internet, NAT, or SSM VPC endpoints) Zero inbound ports open at all: no SG rule for 22/3389, no bastion host, no key pair; every session is logged (who, when, commands) via CloudTrail/S3/CloudWatch Logs for audit Needs the Agent + IAM role provisioned ahead of time; still needs an outbound path to the SSM service (a fully air-gapped private subnet needs SSM VPC endpoints added) The AWS-recommended default, especially for private-subnet instances; removes the need for a bastion host entirely
SSH client (traditional) A key pair, an SG allowing inbound 22 from your source, and network reachability (public IP directly, or a bastion host / VPN / Direct Connect into a private subnet) Familiar tooling, full terminal control, works the same regardless of AWS-specific agents or console access You're responsible for storing/rotating the private key yourself; some inbound port must be opened somewhere; a private instance needs an extra bastion/VPN hop Existing key-based workflows, automation/scripts, or environments not using Systems Manager
EC2 Serial Console Enabled at the account level, plus an OS-level console user/password configured on the instance in advance Works even when the instance has no working network path at all: locked out by a bad SG/NACL/route table change, broken sshd, corrupted network config Text-only, no file transfer, no IAM-role convenience, and useless in the emergency it's meant for if you didn't configure OS console access beforehand True last-resort troubleshooting of an instance that's unreachable over the network by every other method here
Exam callout: "most secure way to connect to a private instance, no bastion host" β†’ Session Manager, every time. It's the only option here needing no inbound rule whatsoever. Serial Console is the odd one out: the only method that still works when the instance's own network stack is the thing that's broken, which is exactly why it can't rely on SSH, SSM, or Instance Connect (all of which need a working network).
🧠 Quick Memory Hooks
Compute-optimized = CPU-heavy  Β·  RAM-optimized = memory  Β·  I = IOPS-heavy local disk
Spot = cheapest but can vanish with 2 min warning. Reserved/Savings Plans = commit for a discount. On-Demand = pay for flexibility.
If it must survive termination, it's not on Instance Store.
Session Manager = no open ports, no keys, full audit trail. Serial Console = the only way in when the network itself is the problem.
πŸͺ£ S3 & Storage: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
S3 is object storage, not a disk: think "a bucket of files with a key," not a filesystem with folders. Almost every exam question is really asking which storage class or which access control layer. Same "11 nines" (99.999999999%) durability across every class. What changes between classes is availability and retrieval cost/speed, not durability.
πŸ—ƒοΈ Storage Classes (frequently tested!)
ClassAccess patternRetrieval timeTypical use case
StandardFrequent accessMillisecondsActively-used data (website assets, app content)
Intelligent-TieringUnknown/changing access patternMillisecondsAuto-moves objects between tiers by observed access, no retrieval fees for tier changes
Standard-IAInfrequent, but need fast access when neededMillisecondsBackups, disaster recovery files, older but still-needed data
One Zone-IAInfrequent, re-creatable dataMillisecondsSame as Standard-IA but single-AZ; cheaper, less resilient (fine for easily-regenerated data)
Glacier Instant RetrievalArchive, rarely accessed, but needs instant accessMillisecondsQuarterly reports, medical images; archived but occasionally opened immediately
Glacier Flexible RetrievalArchiveMinutes to hoursBackups you don't need in a hurry
Glacier Deep ArchiveLong-term archive, almost never accessedStandard: ~12h Β· Bulk: up to 48hCompliance/regulatory retention (7-10 year records), cheapest storage class
Exam pattern: "cheapest, access once or twice a year, can wait hours" β†’ Glacier Deep Archive. "Unpredictable access pattern, don't want to manage tiering manually" β†’ Intelligent-Tiering.
♻️ Lifecycle Rules
A lifecycle rule automatically transitions objects to a cheaper class, and/or expires (deletes) them, after a set number of days. No manual intervention once configured.
Standard 30d Standard-IA 90d Glacier Flexible 180d Deep Archive 365d πŸ—‘ Example rule: every threshold above is fully configurable per rule, and you can skip tiers or stop at any point.
Applies to a whole bucket, or scoped by prefix/tag, e.g. only objects under logs/.
πŸ•“ Versioning
Once enabled (bucket-level), every overwrite or delete keeps the old version instead of destroying it. A "delete" just adds a delete marker on top.
Protects against accidental overwrite/deletion. Restore by removing the delete marker or fetching an older version ID.
Can't be fully turned off once enabled, only suspended. MFA Delete adds a second-factor requirement to permanently delete a version.
Lifecycle rules can target noncurrent versions separately, e.g. auto-expire old versions after 90 days to control storage cost growth.
πŸ”’ Encryption
SSE-S3Amazon-managed keys: encryption at rest with zero setup; the default for every new object/bucket
SSE-KMSKMS-managed keys; adds an audit trail (every use logged to CloudTrail) and granular per-key IAM control; costs a bit more (KMS API calls). Comes in two flavors: an AWS managed key (created/rotated automatically, zero setup, but you can't set its key policy) or a customer managed CMK (you create it, control its key policy/rotation, can disable or schedule deletion). "Full audit + keys managed by us" in a question points at a customer managed CMK specifically, not just "SSE-KMS" in general
SSE-CYou supply your own encryption key with every request; AWS never stores it
Client-sideYou encrypt before upload; S3 only ever sees ciphertext
In transit: enforce HTTPS-only access via a bucket policy condition on aws:SecureTransport.
πŸ”‘ Access Control
Bucket PolicyA resource-based JSON policy on the bucket itself, the standard way to grant/restrict access, including cross-account
IAM PolicyAttached to a User/Role instead of the bucket; same underlying permission model, different attachment point
ACLLegacy, object/bucket-level access list. AWS now recommends disabling ACLs entirely and using policies instead
Block Public Access is a bucket (and account-wide) setting that overrides any policy/ACL trying to make something public. It's on by default for new buckets, and the exam expects you to know it must be deliberately turned off before a bucket can be made public, no matter what the policy says.
🌐 CORS (Cross-Origin Resource Sharing) (frequently tested!)
CORS is a browser security rule, not an S3-specific concept: a webpage's JavaScript is blocked by the browser from calling a different origin (domain, scheme, or port) than the page itself was loaded from, unless that other origin's response explicitly says it's allowed.
If a webpage hosted anywhere makes a JavaScript fetch/XHR request straight to an S3 bucket (a different origin from the page), the browser blocks it by default. Fix: add a CORS configuration on the bucket itself: a small JSON/XML rule set naming the allowed origins, HTTP methods (GET/PUT/POST/etc.), and headers. This is a setting on the bucket's CORS configuration, separate from a bucket policy or IAM policy.
Exam pattern (the classic trap): a question describes a script/webpage making authenticated requests to S3 and failing, with a bucket policy and IAM permissions already correctly in place. The fix is still enabling CORS. A bucket policy governs authorization (is this caller allowed to do this), while CORS governs a completely separate browser-enforced check that happens regardless of whether the caller is authorized. Fixing one doesn't fix the other; versioning and encryption settings are unrelated distractors here too.
βš–οΈ S3 vs. EFS vs. EBS
S3EFSEBS
What it isObject storage (a bucket, not a drive)Managed NFS file systemA virtual block-storage disk
Attaches toNothing; accessed over HTTP(S) API from anywhereMany EC2 instances at once, across multiple AZsOne instance at a time (Multi-Attach io1/io2 excepted, same AZ only)
ScopeRegion-wideRegion-wide (multi-AZ)Single AZ
Typical useStatic assets, backups, data lake, hostingShared content/config across a fleet of instancesBoot volumes, databases; anything needing a real filesystem for one instance
Exam pattern: "multiple EC2 instances behind a load balancer must all see the same files/documents". This is a shared-filesystem need, so EBS (one instance at a time) is disqualified regardless of how the data got out of sync in the first place; the fix is EFS, not "copy the same files to every instance's own EBS volume" (which just re-breaks the moment anything changes).
🚚 Getting Data Into AWS: Transfer & Migration Options (frequently tested!)
Which tool is right depends entirely on data volume and available bandwidth. The exam tests whether you can match the scenario's numbers to the right option, not just recognize the service names.
ToolWhat it doesBest for
S3 Transfer AccelerationRoutes an upload through CloudFront's global edge locations over AWS's private backbone instead of the public internet path to the bucket's RegionFrequent, ongoing uploads from geographically distributed users/sites where the network path itself is the bottleneck, not a one-off bulk migration
AWS DataSyncAn online, automated, ongoing transfer/sync agent between on-premises storage (NFS/SMB) and S3/EFS/FSx; encrypts in transit, validates data, can run on a scheduleRepeated or ongoing transfers over an existing network link (works well alongside Direct Connect); not for a single one-time bulk migration with very limited bandwidth
AWS Snowball EdgeAWS ships you a physical device; you load it with data on-site, ship it back, AWS ingests it directly into S3; bypasses the network entirelyLarge one-time (or infrequent) bulk transfers (tens of TB to PB) where available bandwidth would otherwise take days/weeks
AWS Storage Gateway (File Gateway)An on-premises virtual appliance presenting an NFS/SMB share that's actually backed by S3: files written locally land in S3, with the most-recently-used data cached locally for low-latency accessKeeping an existing on-prem file-based workflow while transparently using S3 as the actual backing store, including S3 Lifecycle rules on the data it writes
Exam pattern: "as quickly as possible, minimal operational complexity, one-off, very large volume, limited/expensive bandwidth" β†’ Snowball Edge. This beats Transfer Acceleration whenever the numbers imply days/weeks over the wire (a classic giveaway: a large TB/PB figure paired with a modest daily bandwidth figure that doesn't divide down to a reasonable transfer window). "Ongoing/repeated programmatic sync between on-prem and AWS storage" β†’ DataSync. "Keep an existing file-share workflow, back it with S3, apply lifecycle rules" β†’ Storage Gateway File Gateway.
Amazon Athena: a serverless query service that runs standard SQL directly against data already sitting in S3 (CSV, JSON, Parquet, etc.): no database to provision, no ETL pipeline, pay only per query (per data scanned). The default answer whenever a question wants ad-hoc/on-demand analysis of S3-resident logs or data with minimal setup, over standing up Redshift (a provisioned data warehouse) or a Glue+EMR pipeline (built for heavier, repeated ETL processing).
🧠 Quick Memory Hooks
Durability never changes across classes: only availability and retrieval speed/cost do.
S3 = a bucket (HTTP API)  Β·  EFS = a shared network drive for many instances  Β·  EBS = a private disk for one instance.
Block Public Access beats everything else: a wide-open bucket policy still won't be public if this is on.
πŸ—„οΈ Databases: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Pick the database by the shape of the data and the access pattern, not familiarity: relational + complex queries β†’ RDS/Aurora; key-value at massive scale β†’ DynamoDB; sub-millisecond lookups β†’ ElastiCache.
πŸ›’οΈ RDS Basics
Amazon RDS = a managed relational database: AWS handles patching, backups, and failover setup. Engines: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, and Aurora.
"Managed" means AWS handles the undifferentiated heavy lifting (OS patching, engine patching, backups); you still choose instance size, storage, and when maintenance windows apply.
Automated backups + manual snapshots both available; point-in-time recovery restores to any second within the backup retention window.
RDS Proxy (frequently tested!) A fully-managed connection pooler that sits between the application and the database. Solves a specific problem: a highly-concurrent, short-lived-connection caller (classically Lambda, where every concurrent invocation can open its own DB connection) can exhaust a database's max-connections limit and slow it down with constant connect/disconnect overhead. RDS Proxy pools and reuses a small number of real DB connections behind many app-side connections. It also speeds up failover: the proxy holds the connection and transparently re-points it to the new primary, so the application doesn't need to detect and reconnect itself. Exam pattern: "Lambda functions are overwhelming the database with connections" β†’ RDS Proxy, not simply "increase the DB instance size."
πŸ”€ Multi-AZ vs. Read Replicas (frequently tested!)
These solve two different problems and are often combined, not alternatives to each other.
Multi-AZ: High Availability Primary DB (AZ-a) sync Standby DB (AZ-b) Not readable (auto-failover only) Read Replica: Read Scaling Primary DB async Read Replica Readable (can promote to standalone) Multi-AZ = disaster insurance (same Region, automatic failover). Read Replicas = a scaling lever (same/cross-Region, manual promotion).
Multi-AZRead Replica
SolvesHigh availability / disaster recoveryRead traffic scaling: offload reporting/analytics reads
ReplicationSynchronousAsynchronous (small replication lag)
Readable?No, standby is idle until failoverYes, that's the whole point
FailoverAutomatic (AWS-managed, DNS endpoint stays the same)Manual promotion to a standalone instance
LocationA second AZ in the same RegionSame Region or cross-Region, plus a narrow Aurora-specific case: an RDS MySQL instance can replicate into an Aurora MySQL cluster as a migration path (not general cross-engine replication, e.g. no MySQL→PostgreSQL)
πŸ’Ž Aurora
AWS's own MySQL/PostgreSQL-compatible engine, same wire protocol/drivers, but a re-architected storage layer that auto-replicates data 6 ways across 3 AZs and claims up to 5x MySQL / 3x PostgreSQL throughput.
Aurora ReplicasUp to 15, low replication lag (shared storage layer, not a full copy); can also fail over automatically, unlike a standard RDS read replica
Global DatabaseOne primary Region + up to 5 read-only secondary Regions, typically <1s lag; for globally-distributed reads or DR
Aurora ServerlessAuto-scales capacity up/down (even to zero) with demand, for spiky or unpredictable workloads instead of sizing an instance by hand
Exam pattern: when a question wants "the best relational performance/availability on AWS," Aurora is usually the intended answer over plain RDS MySQL/PostgreSQL.
⚑ DynamoDB (frequently tested!)
Fully-managed key-value / NoSQL store, single-digit millisecond latency at any scale, no servers to manage at all (serverless).
Partition Key (hash key)Determines which physical partition an item lives on; the core design decision for even data distribution and query performance. A high-cardinality key (lots of distinct values) spreads load evenly; a low-cardinality one creates a "hot partition"
Sort Key (range key), optionalPaired with the partition key to form a composite primary key. Items sharing a partition key are stored together, ordered by sort key, so a query can efficiently pull a range (e.g. all of one customer's orders, sorted by date)
On-Demand capacityPay per request, auto-scales instantly; for unpredictable/spiky traffic
Provisioned capacitySet Read/Write Capacity Units (RCU/WCU) ahead of time; cheaper at steady, predictable load. Pair with DynamoDB Auto Scaling (a target-utilization setting, e.g. 70%, conceptually identical to ASG Target Tracking) so capacity adjusts within a min/max band automatically instead of a human resizing it. This is the "most cost-effective" answer far more often than switching to On-Demand or adding DAX, when the workload is steady rather than spiky. A ThrottledRequests CloudWatch alarm signals capacity is set too low for either mode.
DAXDynamoDB Accelerator, an in-memory read-through/write-through cache in front of DynamoDB, microsecond latency for read-heavy workloads. Solves a read-latency problem, not a cost/throttling problem; a common exam distractor when the real fix is just right-sizing capacity (above)
DynamoDB StreamsAn ordered, 24-hour log of every item-level change (insert/update/delete) on a table, the trigger source for reacting to data changes (e.g. invoking a Lambda function per change), not just a Lambda poll-source detail
TTL (Time To Live)Auto-deletes an item once its designated timestamp attribute passes: free (no WCU consumed) background deletion, for data with a natural expiry (session data, temporary tokens) instead of a manual cleanup job. A completely different "TTL" from CloudFront's cache TTL (same term, unrelated concept)
Global TablesMulti-Region, multi-active replication: write to any Region, reads stay fast and local everywhere
Global Secondary Index (GSI)Local Secondary Index (LSI)
KeysA different partition key (and optionally a different sort key) from the base tableSame partition key as the base table, a different sort key
When it can be createdAny time (added to an existing table)Only at table creation time; cannot be added or removed later
CapacityIts own separate RCU/WCU, provisioned independently of the base tableShares the base table's capacity
ConsistencyEventually consistent reads onlyEventually or strongly consistent reads (your choice)
LimitUp to 20 per tableUp to 5 per table
Exam pattern: a question needing to query by a completely different attribute than the base table's partition key (e.g. table keyed by Song ID, but the app needs to query "all songs by this Artist") β†’ GSI, essentially always. An LSI can't change the partition key at all. Reach for LSI only when the scenario explicitly needs strong consistency on an alternate sort order for the same partition key, and the index was planned before the table was created.
🧊 ElastiCache
RedisMemcached
Data structuresRich (lists, sets, sorted sets, hashes, pub/sub)Simple key-value only
PersistenceOptional snapshotting/AOF (can survive a restart)None: pure in-memory, gone on restart
HAMulti-AZ with automatic failover, read replicasNone built-in; just multiple independent nodes
ScalingCluster mode shards data across nodesSimple horizontal scaling via multiple nodes
Exam pattern: "need persistence, replication, or pub/sub" β†’ Redis. "Simplest possible object cache, no HA needed" β†’ Memcached.
🧭 Choosing the Right Database
1Complex queries, joins, transactions, existing SQL app β†’ RDS / Aurora
2Massive scale, simple access patterns, need single-digit ms at any size β†’ DynamoDB
3Sub-millisecond lookups, session store, leaderboard, pub/sub β†’ ElastiCache
4Analytics over huge historical datasets, complex OLAP queries β†’ Redshift (data warehouse, out of scope of this page but good to recognize by name)
🧠 Quick Memory Hooks
Multi-AZ = insurance policy (sync, unreadable standby). Read Replica = extra hands (async, readable, scales reads).
Aurora = RDS's faster, AWS-native cousin.
DynamoDB = no servers, no limits, simple lookups. ElastiCache = everything in RAM, blink-fast.
GSI = a new door with a new key. LSI = the same door, a different sort order, and only if you said so on move-in day.
RDS Proxy = a bouncer for the database's front door, so Lambda's crowd of short visits doesn't overwhelm it.
βš–οΈ ELB & Auto Scaling: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
A Load Balancer spreads traffic across instances; Auto Scaling changes how many instances exist. Together: elastic capacity that survives both traffic spikes and instance failures.
πŸ›‘οΈ High Availability vs. Fault Tolerance (frequently tested!)
High Availability (HA): the system keeps running with minimal downtime when a component fails: it recovers automatically, but recovery can involve a brief, real interruption while failover happens. The goal is "downtime measured in seconds/minutes, not hours," not necessarily zero.
Fault Tolerance (FT): the system keeps running with zero service interruption when a component fails: end users never notice. This needs redundant capacity that's already active and already absorbing load before the failure, not a standby that has to be switched to.
High AvailabilityFault Tolerance
Downtime on failureBrief (a failover event), small but realNone: no interruption at all
How it's achievedA standby/replica takes over after detecting failureRedundant components are already active in parallel, each individually sized to absorb the loss of another
CostLower: standby capacity, not necessarily fully activeHigher: every "extra" unit of redundancy is running (and billed) all the time, not just on standby
AWS exampleRDS Multi-AZ: synchronous standby in another AZ, but there's still a real (~1-2 min) failover event on the primary's failureAn ASG spread across 3 AZs behind an ALB, sized so any one AZ can be lost and the remaining two still handle full load with zero dropped requests; no failover event, nothing to "switch to"
Relationship: Fault Tolerance is a stronger, more expensive goal that sits on top of High Availability, not a separate, unrelated concept. Every fault-tolerant system is highly available (it has no meaningful downtime at all), but not every highly-available system is fault tolerant (Multi-AZ RDS is HA, but that failover window means it isn't strictly FT). S3 and DynamoDB are fault tolerant by design: their multi-AZ redundancy is invisible to you, with no failover event to reason about.
Exam pattern: "recovers automatically, brief/acceptable interruption is fine" β†’ High Availability (Multi-AZ, Auto Recovery). "Must keep serving traffic with zero interruption even if a whole AZ goes down" β†’ Fault Tolerance (over-provisioned ASG/ELB across β‰₯3 AZs, or a managed service that's fault tolerant by default like S3/DynamoDB).
πŸ”€ Load Balancer Types (frequently tested!)
Elastic Load Balancing (ELB) is the umbrella AWS service, a fully managed load balancer that AWS scales, patches, and runs across multiple AZs for you (no servers of your own to manage). It comes in four types, split across the two networking layers the exam cares about: Application (Layer 7) and Network (Layer 4), plus Gateway (Layer 3) and a legacy fourth type.
Application (ALB)Network (NLB)Gateway (GWLB)Classic (CLB), legacy
Layer7 (HTTP/HTTPS)4 (TCP/UDP)3 (IP), transparent network gatewayBoth 4 and 7, but shallowly; predates ALB/NLB's split
RoutingPath/host-based rules, e.g. /api/* β†’ one target group, /images/* β†’ anotherJust forwards connections, no content awarenessPasses all traffic through to a fleet of virtual appliancesOne flat routing model: no path/host rules, no per-listener target groups
PerformanceVery fast, but not the fastestMillions of requests/sec, ultra-low latency, static IP supportN/A (not about speed)Lower throughput than ALB/NLB, no modern performance features
Typical useWeb apps, microservices, containersExtreme performance needs, TCP-only protocols, static IP requirementInserting a fleet of firewalls/IDS/IPS appliances transparently in front of trafficOnly ever the right answer for an old EC2-Classic-era app that predates VPC. AWS actively steers new designs toward ALB/NLB instead
Exam pattern: "route by URL path or hostname" β†’ ALB. "Extreme throughput, need a static IP" β†’ NLB. "Deploy third-party security appliances inline" β†’ GWLB. If "Classic Load Balancer" appears as an answer option at all, it's almost always the wrong one. The exam uses it to test whether you know it's the deprecated previous-generation type, not a live recommendation.
GWLB + GENEVE (frequently tested!): GWLB listens for every IP packet on every port (not specific listeners like ALB/NLB) and hands each one to a fleet of virtual appliances using the GENEVE protocol on port 6081: the appliance inspects/filters the packet and sends it back, transparently, before it continues to its real destination. Knowing "GENEVE, port 6081" by name is a recurring exact-recall question.
Target types: ALB can register EC2 instances, IP addresses (including on-prem/peered-VPC targets), and, uniquely, Lambda functions as targets (the request becomes the Lambda's event payload; no traditional health check needed since Lambda has no "instance" to check). NLB registers instances, IP addresses, or an Application Load Balancer (see below). GWLB registers virtual-appliance instances/IPs, not application targets.
Cross-Zone Load Balancing (frequently tested!): controls whether each load balancer node (one is provisioned per enabled AZ) spreads its share of traffic across every registered target in every AZ, or only the targets sitting in its own AZ. ALB has it on by default, free. NLB and GWLB have it off by default, and turning it on incurs cross-AZ data transfer charges for any traffic that has to cross an AZ boundary to reach a target. A classic "why is my traffic unevenly spread across AZs" gotcha.
Worked example: one load balancer spanning 2 AZs: AZ-A has 1 target, AZ-B has 3 targets. Cross-zone off: the LB node in AZ-A can only forward to AZ-A's 1 target, so that single instance absorbs 50% of all traffic (as much as the other 3 combined): an even split per node, not per target, so mismatched AZ instance counts create hot spots. Cross-zone on: every LB node spreads its traffic across all 4 targets regardless of AZ, so each instance gets roughly a 25% share instead.
Every ELB (ALB/NLB/GWLB) must be enabled in at least 2 AZs, a hard requirement, not just an HA best practice. Within each enabled AZ, an ELB can only use one subnet; size that subnet /27 or larger, with at least 8 free IP addresses, since AWS reserves capacity in it for the load balancer's own nodes to scale into; a subnet that's nearly full of other resources can silently block the ELB from scaling.
πŸ—οΈ ALB and NLB Deployments (frequently tested!)
Internet-facing vs. Internal (both ALB and NLB): chosen when you create the load balancer, not something you can quietly bolt on later.
Internet-facingInternal
DNS resolves toPublic IP(s)Private IP(s) only
SubnetsPublic subnets (nodes need a route to an Internet Gateway)Private subnets (no path to/from the public internet at all)
Typical useA public web app's front doorInternal microservice-to-microservice traffic, or a private API only other VPC resources call
Exam pattern: "expose this tier only to other services inside the VPC, never to the internet" β†’ Internal load balancer, not Internet-facing. A common wrong-answer trap is reaching for extra security groups instead of the simpler internal-scheme fix.
Listeners: a load balancer can run multiple listeners at once (e.g. one on port 80, one on port 443). Each listener checks for connection requests on its own protocol/port and decides what to do with them independently of the others.
ALB listener rules (content-aware): each listener evaluates a priority-ordered list of rules, each made of conditions (path, hostname, HTTP header, query string, source IP) and an action: forward to a target group, redirect (e.g. every HTTP request on port 80 β†’ the equivalent HTTPS URL on port 443), or fixed-response (return a static status code/body with no backend call at all, e.g. a 503 maintenance page). A default catch-all rule runs when nothing else matches. This is what makes ALB "Layer 7" in practice: the routing decision can depend on the actual request content, not just where it came from.
NLB listeners (connection-aware, not content-aware): a listener typically just forwards every connection straight to one target group: no path/host rules, no redirects, no fixed responses, because NLB operates below the layer where that content is visible. Target types: instance, IP, or even an ALB (put an NLB in front of an ALB to get NLB's static IP / extreme-throughput properties while keeping ALB's content-based routing behind it).
Exam pattern: "need path-based or host-based routing" or "need an HTTP→HTTPS redirect rule" → ALB listener rules. NLB cannot do either.
πŸ”’ Secure Listeners for ELB (frequently tested!)
A secure listener is an HTTPS (ALB) or TLS (NLB) listener: the load balancer terminates the TLS connection itself using an X.509 certificate, decrypts the request, then talks to the targets (in plain HTTP, or re-encrypted via HTTPS/TLS to the target group for end-to-end encryption). Terminating TLS at the load balancer offloads the CPU cost of encryption from every backend instance onto the load balancer instead.
AWS Certificate Manager (ACM): the standard way to get the certificate a secure listener needs: request a public certificate for your domain (validated via DNS or email) and attach it directly to the listener. ACM-issued certificates used with an integrated service like ELB are free and auto-renew, removing the classic "certificate silently expired" outage. A third-party certificate can also be imported into ACM (or IAM, for Regions without ACM) if you already own one.
SNI = Server Name Indication (frequently tested!): lets one HTTPS/TLS listener serve multiple certificates for multiple domains on a single load balancer. The client states which hostname it's connecting to during the TLS handshake itself, before any data is decrypted, so the listener can pick and serve the matching certificate, so no more "one load balancer per domain/certificate" just to keep certs straight. Both ALB and NLB (TLS listener) support SNI.
Security policies: a predefined set of TLS protocol versions and cipher suites the listener will accept; you choose one per HTTPS/TLS listener. A stricter policy (disabling older TLS 1.0/1.1, weak ciphers) is the answer whenever a question mentions a compliance requirement (e.g. PCI-DSS) or "must not allow outdated/weak encryption"; a looser policy trades that off for compatibility with older clients that can't negotiate modern TLS.
ALB also supports mutual TLS (mTLS): the listener can require and verify a client certificate too, not just present its own, for scenarios needing client-identity verification at the load balancer (e.g. B2B/IoT APIs) rather than purely server-side TLS.
Exam pattern: "offload TLS/SSL processing from the application instances" β†’ terminate at the load balancer with a secure listener. "Host several domains, each with its own certificate, behind one load balancer" β†’ SNI. "Enforce modern TLS only for compliance" β†’ pick a stricter security policy, not a certificate change.
🎯 Target Groups & Health Checks
A Target Group is the set of destinations a load balancer routes to: EC2 instances, IP addresses, or even Lambda functions.
The load balancer continuously health-checks every registered target (a ping to a configurable path/port) and stops routing to any that fail; traffic only ever goes to healthy targets.
One load balancer can route to multiple target groups (e.g. ALB path-based rules splitting traffic across several groups). This is how one ALB serves several backend services.
Deregistration delay (connection draining): default 300s. Once a target is deregistered (ASG scale-in, a deploy), the load balancer stops sending it new requests but lets in-flight ones finish before removing it, instead of dropping them mid-request.
Slow Start mode (ALB target group feature, frequently tested!): when enabled, a newly-registered healthy target gets a gradually ramping share of traffic over a configured warm-up window (30s–15min) instead of jumping straight to a full, even share. Protects an app that needs a moment to actually perform well once traffic starts (JVM JIT warm-up, local cache fill, connection pool priming) from being swamped the instant it passes its health check. ALB only: NLB has no equivalent feature, since NLB has no application-layer visibility into "warmed up" at all; an exam option offering "configure a Network Load Balancer with slow start" is a distractor for exactly this reason.
πŸ“ˆ Auto Scaling Group Basics
What EC2 Auto Scaling is: a service that automatically launches and terminates EC2 instances in an Auto Scaling Group to keep capacity matched to actual demand (or a known schedule), instead of a fixed fleet size someone has to resize by hand. What it works with: a Launch Template (the blueprint for every instance it creates: AMI, instance type, key pair, security groups), CloudWatch alarms (the trigger: a metric breach fires a scaling policy), an Elastic Load Balancer (optional but typical; the ASG auto-registers new instances into the target group and deregisters ones it terminates, and can use the ELB's own health checks alongside its own), and plain EC2 instances as the thing actually being scaled. How it works: continuously compares current capacity against the Min/Desired/Max bounds and each active scaling policy's target, launches instances from the Launch Template when it needs to scale out, health-checks every instance and terminates+replaces any that fail, and terminates instances (respecting deregistration delay via the ELB, if attached) when it needs to scale back in.
Auto Scaling Group (ASG): the actual resource you create: a logical, named collection of EC2 instances that AWS manages as one unit. You don't launch or track individual instances yourself; you point EC2 Auto Scaling (the service, above) at the group, and it adds/removes instances within that group to satisfy the group's own settings. An ASG is defined by a Launch Template (which AMI, instance type, etc.) plus three numbers: Min, Desired, and Max capacity.
Spans multiple AZs by design: the ASG tries to balance instances evenly across the AZs you configure, for the same resilience reasons as any multi-AZ design.
An unhealthy instance (per its health check) is automatically terminated and replaced. This is EC2-fleet-level self-healing, distinct from Auto Recovery (which recovers the same instance after a hardware failure; see the EC2 & Compute topic). See "ASG Health Checks: EC2 vs. ELB" below for what counts as unhealthy and the grace period that protects a slow-booting instance from this.
Scaling up vs. scaling out (frequently tested!): Scaling up (vertical) = swap an instance for a bigger one: more vCPU/RAM on the same single box. Scaling out (horizontal) = add more instances of the same size to share the load. ASG only ever scales out: it changes the instance count, never the instance size, which is exactly why the fleet needs to be stateless (see below): any of those interchangeable new instances has to be able to pick up traffic with zero special setup. Example: Amazon's checkout service on Black Friday scales out: the ASG launches more identical web/app instances behind the ALB to absorb the surge, rather than scaling up, which would mean resizing each existing server to a bigger type (a slower, disruptive change with a hard ceiling once you hit the largest instance size).
Typical scale-up use caseA traditional single-writer relational database (RDS), e.g. resizing to a bigger DB instance class when queries are CPU/memory-bound, since you generally can't just add more independent writer nodes the way you can with stateless web servers
Typical scale-out use caseA stateless website/web or app tier behind an ALB, e.g. an ASG adding more identical EC2 instances as request volume grows, exactly the checkout-service pattern above
Scaling out is the generally preferred option, and the exam's default-favored answer, when the architecture allows it. Reasons: no downtime (new instances join alongside the existing ones, so nothing needs a stop/resize/restart cycle); no hard ceiling (keep adding instances, versus scaling up which eventually hits the biggest instance type AWS offers); better fault tolerance (load spread across many instances/AZs, instead of everything riding on one bigger single point of failure); and it matches the pay-for-what-you-use elasticity ASG is built for: capacity can shrink back out just as easily once demand drops, which a bigger single instance can't do without another disruptive resize. It's only off the table when the component itself can't be horizontally distributed, the single-writer relational database case above being the classic example.
πŸ“ Configuring a Launch Template
A Launch Template is the reusable blueprint an ASG (or a one-off "Launch instance from template" action) uses to create every instance. What you configure in it:
AMIWhich machine image to boot from: OS, pre-installed software, any baked-in config (see AMIs in EC2 & Compute)
Instance typee.g. t3.micro, m5.large; vCPU/RAM/network profile for every instance created, or a list of compatible types if using a Mixed Instances Policy (below)
Key pairFor SSH access, often skipped in favor of Session Manager on fleets managed entirely through automation (see Connecting to an Instance in EC2 & Compute)
Security group(s)Which inbound/outbound firewall rules every launched instance gets
StorageRoot EBS volume size/type, plus any additional volumes to attach on launch
IAM instance profileThe Role every instance assumes, so it can call other AWS APIs without an embedded Access Key (see Roles + EC2 in IAM)
User dataA bootstrap script that runs once on first boot (installs software, pulls app code, joins a cluster, etc.), the mechanism that lets a freshly launched instance become "ready" with zero manual setup
Network settingsAuto-assign public IP, etc. For an ASG specifically, the actual subnets/AZs it launches into are set on the ASG itself, not the template
Versioning: editing a Launch Template creates a new numbered version instead of overwriting the old one. Point an ASG at a specific version, or at $Latest/a chosen $Default so new instances automatically pick up template changes (or don't, if pinned to a fixed version). This versioning is the main reason Launch Templates replaced the older Launch Configurations, which are immutable once created: change anything and you must create and swap in an entirely new one.
Mixed Instances Policy: an ASG feature layered on top of a Launch Template: give it a list of compatible instance types (and optionally a mix of On-Demand + Spot) instead of one fixed type, and let the ASG pick whichever is available/cheapest at launch time. Not available with a legacy Launch Configuration.
Exam pattern: "Launch Configuration" still appears as a distractor answer on the exam. AWS no longer recommends creating new ones (no versioning, no Mixed Instances Policy, no Spot/On-Demand mixing). For a new ASG, the correct answer is always Launch Template.
🩺 ASG Health Checks: EC2 vs. ELB (frequently tested!)
An ASG's health check type setting decides which signal it trusts to decide an instance is unhealthy and due for replacement, a different, group-level setting from the System/Instance Status Checks covered in EC2 & Compute (those drive Auto Recovery for a single instance; this drives the ASG replacing instances across the whole group).
EC2 health check (default)ELB health check
What it checksJust the instance's own EC2 status checks: is it running at all (see Status Checks & Auto Recovery in EC2 & Compute)The load balancer's configured health check against the target: a real request to a path/port, so it also catches an app that's up but broken (e.g. returning 500s, or hung)
CatchesA crashed/unreachable instanceEverything EC2 health checks catch, plus an instance that's healthy at the OS level but whose application has failed
Enabled by default?Yes (always on)No, opt-in. Must be explicitly turned on for the ASG even when a load balancer/target group is already attached
Exam pattern: "instances all pass their status checks but the application behind them is actually erroring, and the ASG isn't replacing them" β†’ the ASG is relying on EC2 health checks only; enable ELB health checks so it also trusts the load balancer's application-level result.
Health check grace period (frequently tested!): new instances get a configurable grace period (default 300s) before a failed health check (of either type) is allowed to terminate them. Without it, a slow-booting app (still running its user-data script, still warming up) gets marked unhealthy and killed/relaunched in a loop before it ever finishes starting. Set it to comfortably cover the instance's real boot + app-startup time.
πŸ“Š Auto Scaling & CloudWatch Monitoring
Three different monitoring granularities can feed the CloudWatch metrics a scaling policy reacts to. Don't confuse "how often" with "at what level":
ScopeIntervalCost
Group metricsASG-level (aggregated across the whole group, e.g. group-average CPU)1 minuteFree, but must be explicitly enabled, not on by default
Basic monitoringPer-instance5 minutesFree (the default for every EC2 instance)
Detailed monitoringPer-instance1 minuteOpt-in, and unlike the other two, charges apply
Exam pattern: a scaling policy needs to react quickly (sub-5-minute) to real load; the cheaper fix is usually enabling Group metrics on the ASG itself, not paying for Detailed monitoring on every individual instance, unless the question specifically needs per-instance-level 1-minute data.
πŸ› οΈ ASG Advanced Settings: Cooldowns, Termination & Lifecycle Hooks (frequently tested!)
CooldownDefault 300 seconds (5 minutes). Pairs specifically with Simple Scaling: after a scaling activity, the ASG waits out the cooldown before launching or terminating again, giving the previous change time to actually show up in the metric before reacting further (see Simple vs. Step Scaling above for why Step Scaling isn't blocked by this the same way).
Termination PolicyControls which instance(s) get picked first when a scale-in event needs to remove capacity. The default policy balances across AZs first (so scale-in doesn't leave one AZ empty), then picks the instance using the oldest Launch Template/Configuration, then whichever is closest to its next billing hour. Can be customized (e.g. OldestInstance, NewestInstance) if a question calls for different behavior.
Instance Protection (ASG-scoped Termination Protection)A per-instance flag that excludes that one instance from being picked during scale-in, even if the Termination Policy would otherwise choose it. Distinct from EC2's own account/instance-level Termination Protection (DisableApiTermination; see EC2 & Compute), which blocks a manual/API TerminateInstances call entirely rather than just influencing which instance scale-in picks.
Standby StateManually move a running instance from InService to Standby to patch or troubleshoot it without terminating it or leaving it serving live traffic. While in Standby, the ASG stops health-checking it and excludes it from the desired-capacity count (so it may launch a replacement to make up the difference). Move it back to InService when you're done.
Lifecycle HooksPause an instance in a Pending:Wait state (after launch, before it goes InService) or a Terminating:Wait state (after it's picked for termination, before it actually shuts down) for up to 1 hour by default (extendable via a heartbeat) so a custom action can run. Use cases: download/install software or run final setup before the instance takes traffic; drain in-flight connections or flush/upload data before it's terminated. The ASG publishes the hook event via SNS/EventBridge; the instance stays paused until your code calls CompleteLifecycleAction or the timeout elapses.
Exam pattern: "need Auto Scaling to actually wait while a custom script runs before an instance goes live, or before it's torn down" β†’ Lifecycle Hooks. Launch Template user data alone runs a bootstrap script, but doesn't pause the ASG's own state transition the way a hook does.
πŸ—ΊοΈ Traffic Flow: Internet β†’ ALB β†’ Target Group β†’ ASG
🌍 Internet Application Load Balancer Target Group Auto Scaling Group (min 2 / desired 4 / max 6) AZ-a EC2 EC2 AZ-b EC2 EC2 The ALB only ever routes to healthy instances in the target group; the ASG keeps the group at (or moving toward) the desired count across both AZs.
🎚️ Types of Auto Scaling & Scaling Policies
Four types of Auto Scaling, from least to most automated: Manual: you directly change Min/Desired/Max yourself via console/CLI/API, no automation involved. Dynamic: a CloudWatch alarm reacts to a live metric breach (Target Tracking, Step, or Simple Scaling: the three policy types below). Scheduled: a capacity change fires at a known date/time. Predictive: ML forecasts future load from historical patterns and scales ahead of it automatically, with no alarm or schedule to configure.
ManualChange Min/Desired/Max by hand: no CloudWatch alarm, no schedule, no forecast. Fine for a one-off/rare change; not something you'd rely on for routine demand handling.
Target Tracking"Keep average CPU at 50%": you pick a target metric value, AWS calculates the scaling math for you. The default, easiest, most-recommended approach. Not limited to CPU: any CloudWatch metric works, including a custom metric like an SQS queue's ApproximateNumberOfMessagesVisible (worker-fleet ASGs scaling on backlog size rather than their own CPU), or a built-in ALB metric like ALBRequestCountPerTarget; scale on requests-completed-per-instance instead of CPU when request handling time (not CPU) is what actually strains each instance. The exam uses this to test whether you assume Target Tracking = CPU only.
Step ScalingDifferent scaling amounts depending on how far a CloudWatch alarm's breach goes (e.g. +1 instance if CPU >70%, +3 if CPU >90%); more control, more setup
Simple ScalingOne alarm, one fixed action, then waits out a cooldown before evaluating again. The older, simpler predecessor to Step Scaling
Scheduled ScalingChange capacity at a known future time, e.g. scale up before a Friday-night traffic pattern you already know about
Predictive ScalingUses ML on historical load patterns to scale ahead of forecasted demand, instead of reacting after the metric breaches
Exam pattern: "simplest way to maintain a target CPU/request-count" β†’ Target Tracking, essentially always the expected default answer.
Simple vs. Step Scaling: the classic exam differentiator is cooldown behavior, not just "more granular":
Simple ScalingStep Scaling
Alarm β†’ actionOne alarm, one fixed action (e.g. average CPU >70% β†’ +1 instance)Multiple actions sized to how far the breach goes off one alarm (e.g. average CPU 70–90% β†’ +1, >90% β†’ +3)
Cooldown behaviorBlind during cooldown: must wait the full cooldown period out after firing before it will evaluate the alarm again, even if CPU keeps climbingCan keep adjusting capacity while still in cooldown if the alarm remains in ALARM state, reacting to the current step instead of waiting it out
Exam framingOlder, simpler predecessor; rarely the "best" answer once Step Scaling is also an optionPreferred whenever the question wants a faster or finer-grained response to how severe the breach is, not just that it happened
Default Instance Warmup (frequently tested; easily confused with the health check grace period and ALB Slow Start above, but a distinct third thing): a configurable period during which a newly-launching instance is excluded from the ASG's own aggregated CloudWatch metrics (e.g. group-average CPU). Without it, a still-booting instance's artificially low/spiky CPU can skew the metric a Target Tracking/Step policy is reacting to, causing the ASG to misjudge real demand and launch more instances than actually needed (overscaling) while everything is still warming up. This is scaling-math protection, different from the health check grace period (which only stops a slow-booting instance from being wrongly killed) and from ALB Slow Start (which only throttles how much live traffic a new target gets); a question can legitimately need all three configured together on the same ASG.
πŸ§ͺ Worked Example: Three Things That Can Change an ASG's Instance Count (frequently tested!)
Same ASG (min 2 / desired 4 / max 10): three completely independent triggers can each change what's running. The exam likes to test whether you can tell these apart, especially #1 vs. #2 (only one of them is actually "scaling"):
TriggerWhat happensWorked example
1. Metric-based (dynamic) scalingA CloudWatch alarm watches a metric (usually average CPU across the whole group) and fires a scaling policy once it breaches a threshold for the alarm's evaluation periodAverage CPU across the ASG hits and holds 80% β†’ the CloudWatch alarm goes into ALARM state β†’ a Target Tracking or Step Scaling policy adds instances (e.g. +2) from the Launch Template β†’ CPU falls back toward target as load spreads across more instances β†’ a separate low-CPU alarm can later trigger scale-in
2. Health-check-driven replacementNot a scaling policy at all: no CloudWatch alarm, no policy evaluation. The ASG continuously health-checks every instance and swaps out any that fails, purely to hold the group at its existing desired capacityOne instance fails its EC2 or ELB health check ("a unit goes down") β†’ ASG terminates it and launches a replacement from the same Launch Template β†’ desired capacity stays at 4 throughout; it's a 1-for-1 swap, never a scale-out
3. Scheduled scalingA scaling action set for a specific date/time (once) or a recurring cron-style schedule, for load you already know is coming; no metric involvedA retail site's traffic reliably jumps every Monday morning when staff return and place orders β†’ a scheduled action raises min/desired capacity ahead of that, e.g. 07:00 every Monday β†’ desired 8, then a second scheduled action scales back down Friday evening β†’ nothing has to breach a threshold first, because the pattern is already known rather than detected
Exam pattern: "predictable, recurring load (payroll runs, Monday-morning traffic, Black Friday)" β†’ Scheduled Scaling: you already know it's coming, no need to wait for a metric to breach. "React to unpredictable load" β†’ CloudWatch-alarm-driven Target Tracking/Step Scaling. "An instance/AZ fails" β†’ this is not a scaling-policy trigger at all. It's the ASG's own health check replacing a broken instance to hold the group at its already-existing desired count, which is why it can happen even when the group is sitting well under Max.
πŸ”„ Stateful vs. Stateless (frequently tested!)
Stateless: any instance can answer any request because it holds no session/user data locally: it reads whatever context it needs from a shared store instead. This is the design ASG assumes: instances get terminated and replaced constantly (scale-in, unhealthy replacement, AZ failure), so anything kept only in an instance's memory or local disk is lost the moment that instance goes. Example: checking the weather on a public site: you type a city, get a forecast, and nothing about that visit is remembered; the very next request could land on a completely different instance with no loss of function.
Stateful: an instance holds session/user data locally (in-memory or on local disk), so a request has to keep landing on that same instance to work correctly. Fights against elastic scaling: kill that one instance and its state is gone. Example: shopping on Amazon: your cart contents, browsing/purchase history, and login session all have to persist and follow you across every subsequent request, not just the one instance that first handled you. (In practice AWS-scale apps like this externalize that state (see the table below) rather than pinning you to one instance, which is exactly the "stateless design, statefully-feeling app" pattern the exam is testing for.)
Three-tier breakdown of the Amazon example: Web tier (renders the storefront pages): stateless, any instance can serve any request. Application tier (checkout/order logic), the layer that handles your purchase: it validates the order and writes it to a database. Data tier (RDS/DynamoDB): where the order actually becomes durable, permanently tied to your account regardless of which web/app instance you hit next time.
Exam nuance: "writes to a database" is not the same as "is stateful" in the ASG sense. An application-tier instance that records your order straight to the database and keeps nothing about you in its own memory afterward is still stateless and fully replaceable: ASG can terminate it a second later with zero data loss. It's the database, not the compute instance, that's genuinely stateful here (it's the one thing in this whole flow that must not be casually replaced/wiped). This is exactly why the exam's recommended pattern is: keep every compute tier (web and app) stateless, and push all real state down into RDS/DynamoDB/ElastiCache.
StatelessStateful (with sticky sessions)
Where session data livesExternalized: ElastiCache (Redis/Memcached), DynamoDB, RDSOn the instance itself (memory or local disk)
ALB routingAny healthy target in the group (normal round-robin/least-outstanding-requests)Sticky sessions (session affinity) pin a client to the same target via a cookie
Effect of scale-in / instance replacementNo data loss: any other instance already has access to the same shared stateThat client's session data is lost when its instance is terminated
Exam-favored designYes: the "best practice" answer whenever ASG/ELB is in playOnly when the question explicitly requires sticky sessions (legacy app, no session-store option)
Sticky sessions (ALB feature): a load-balancer-generated cookie pins a client to one target for the life of the session. This makes a stateful app work despite multiple targets, but it doesn't make the app stateless, and it can unbalance load onto whichever instances got "sticky" early.
Exam pattern: "app stores session data on the instance and users get logged out/lose their cart after scaling events" β†’ the fix is to externalize session state (ElastiCache or DynamoDB), not just to enable sticky sessions: sticky sessions patch the symptom, externalizing state removes the dependency on any one instance entirely.
πŸ›οΈ Architecture Patterns: Auto Scaling and ELB (frequently tested!)
A set of specific, exam-style requirement β†’ solution pairings: the kind of one-line scenario the exam drops into a question stem, expecting you to recognize the pattern instantly.
RequirementSolution
High availability and elastic scalability for web serversEC2 Auto Scaling + an Application Load Balancer, spread across multiple AZs
Low-latency connections over UDP to a pool of instances running a gaming applicationNetwork Load Balancer with a UDP listener (NLB is the only ELB type that handles UDP at all)
Clients need to whitelist static IP addresses for a highly available load-balanced application in a RegionNLB: it exposes one static IP per AZ automatically, unlike ALB/CLB which only ever have DNS names
Application on EC2 in an ASG requires disaster recovery across RegionsCreate a matching ASG in a second Region with capacity set to 0; take AMI/EBS snapshots and copy them across Regions on a schedule (Lambda or Data Lifecycle Manager). The standby ASG scales up from those snapshots only if the primary Region fails
Application on EC2 must scale in larger increments for a big traffic increase than for a small oneStep Scaling with a larger capacity-increase step configured for the bigger breach (see Simple vs. Step Scaling above)
Need to scale EC2 instances behind an ALB based on the number of requests each instance has completedTarget Tracking policy on the built-in ALBRequestCountPerTarget metric (see the Types of Auto Scaling section above)
Application runs on EC2 behind an ALB; once authenticated, a user shouldn't have to reauthenticate if their instance failsExternalize session state to DynamoDB or ElastiCache. See Stateful vs. Stateless above; don't rely on sticky sessions alone
Company is deploying an IDS/IPS system using virtual appliances and needs it to scale horizontallyGateway Load Balancer in front of the virtual-appliance fleet (see Load Balancer Types above)
🧭 Putting It Together: Which One Do I Reach For? (frequently tested!)
The exam loves to describe a scenario in plain business language and make you pick the right ELB/ASG/session-state tool. This table collects the "if the question says X, the answer is Y" patterns from across this whole topic in one place.
ScenarioReach forWhy
[Scaling] Traffic reliably spikes at a known time (payroll runs, Monday mornings, Black Friday)Scheduled ScalingYou already know it's coming, no need to wait for a metric to breach first
[Scaling] Traffic is variable/unpredictable, want to hold average CPU (or another metric) at a target automaticallyTarget Tracking (Dynamic)Simplest, most-recommended default; you state the goal, AWS does the scaling math
[Scaling] Need different-sized reactions depending on how severe the breach is, and to keep reacting through a fast-moving spikeStep ScalingMultiple actions off one alarm, and, unlike Simple Scaling, keeps adjusting during cooldown
[Scaling] One-off, rare capacity change (a planned migration, a short test)Manual scalingNot worth building automation for something that happens once
[Scaling] A big known future event with historical seasonal data to learn from (e.g. last year's Black Friday numbers)Predictive ScalingML forecasts demand and pre-scales ahead of it, instead of reacting after a threshold breach
[Scaling] A relational, single-writer database (RDS) is CPU/memory-boundScale up (bigger instance class)You generally can't add more independent writer nodes the way you can with stateless web servers
[Scaling] A stateless web/app tier is running out of headroom as request volume growsScale out (ASG adds more instances)Any instance can answer any request, so more identical instances behind the LB is the natural fix
[Load Balancer] Public web app needs path-based or host-based routing (/api/* vs /images/*)ALBLayer 7, the only type that reads request content to route on
[Load Balancer] Need extreme throughput/ultra-low latency, or a static IP for a client allow-listNLBLayer 4, built for raw speed and static IP support, not content-aware routing
[Load Balancer] Need to insert third-party firewall/IDS/IPS appliances transparently in front of trafficGWLBLayer 3 transparent pass-through built specifically for security-appliance fleets
[Load Balancer] A backend/microservice tier should only ever be reachable from inside the VPC, never the internetInternal load balancerDNS resolves to private IPs only, deployed in private subnets; simpler and more correct than trying to lock it down purely with security groups
[Load Balancer] Uneven traffic across AZs because instance counts per AZ don't matchEnable Cross-Zone Load Balancing (already on for ALB; opt-in for NLB)Spreads every LB node's traffic across all targets in all AZs, not just the targets in its own AZ
[Session State] Users get logged out or lose their cart after a scale-in/replacement eventExternalize session state (ElastiCache or DynamoDB)Removes the dependency on any one instance entirely (the actual fix, not a patch)
[Session State] A legacy app can't be quickly refactored to share session state externallySticky sessions (ALB)Makes a stateful app work across multiple targets as a stopgap, but doesn't remove the underlying risk of losing that instance
[Health] A newly-launched instance keeps getting killed as "unhealthy" while it's still bootingIncrease the Health Check Grace PeriodGives the app time to finish starting before health checks start counting against it
[Health] Instances pass their status checks but the app returns errors under real requestsEnable/use the ELB (target group) health check, not just the EC2 status checkOnly the ELB health check makes a real request to the app; EC2 status checks only see infrastructure-level failures
🧠 Quick Memory Hooks
ALB reads the letter (URL path/host): layer 7. NLB just moves the envelope fast: layer 4. GWLB is a transparent tollbooth for security appliances.
Load Balancer = where traffic goes. Auto Scaling Group = how many places it can go.
Target Tracking = tell AWS the goal, not the math.
Stateless = any instance can answer, because the state isn't on the instance. Sticky sessions only patch a stateful app: they don't remove the risk.
Cross-Zone Load Balancing = ALB shares fairly for free; NLB has to be asked, and it costs to ask.
🧭 Route 53 & DNS: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Route 53 is AWS's DNS service: it answers "what IP does this name point to," and its routing policies are really just different rules for which answer to give. Also a domain registrar: two separate jobs bundled under one name.
πŸ“‚ Hosted Zones
A Hosted Zone is a container of DNS records for one domain (e.g. example.com), where you actually create A, CNAME, MX, etc. records.
Public Hosted ZoneAnswers queries from the public internet (a normal, internet-facing domain)
Private Hosted ZoneOnly resolves inside one or more associated VPCs: internal-only DNS names, invisible to the outside world
πŸ“‡ Record Types: Alias vs. CNAME
A record β†’ domain name to IPv4. AAAA β†’ IPv6. CNAME β†’ domain name to another domain name (not an IP).
Alias record (AWS-only, frequently tested): looks like a CNAME but works at the zone apex (example.com, not just www.example.com; a plain CNAME is not allowed at the apex per the DNS spec) and points straight at an AWS resource (ALB, CloudFront, S3 website endpoint) by its AWS-internal identity, not just its current hostname. Free to query, and automatically tracks the target's IP if it changes.
Exam pattern: "point the domain's root/apex directly at an ALB/CloudFront distribution" β†’ Alias record, not CNAME (CNAME literally can't be used at the apex).
🧭 Routing Policies (frequently tested!)
PolicyWhat it doesTypical use case
SimpleOne record, one (or a random) answer, no logicA single-server site, no need for anything fancier
WeightedSplit traffic across records by assigned weight/percentageCanary/A-B testing a new version, gradual migration
Latency-basedRoute to whichever Region gives the requester the lowest latencyGlobal app, multiple Regional deployments, want the fastest response
FailoverPrimary/secondary pair; routes to secondary only when primary's health check failsActive-passive DR setup
GeolocationRoute by the requester's geographic locationLegal/licensing restrictions, localized content by country
GeoproximityRoute by geographic distance, with an adjustable "bias" to shift traffic toward/away from a RegionFine-tuned traffic shaping by location (Traffic Flow feature)
Multi-value AnswerReturns several healthy IPs in response, client picks oneSimple client-side load distribution + basic health checking, not a substitute for a real load balancer
Exam pattern: "lowest latency for global users" β†’ Latency-based. "Active-passive DR, switch on failure" β†’ Failover. "Gradually shift 10% of traffic to a new version" β†’ Weighted.
🩺 Health Checks
Route 53 can monitor an endpoint (HTTP/HTTPS/TCP) from multiple global locations and mark it healthy/unhealthy.
Unhealthy records are automatically excluded from what Route 53 returns. This is the mechanism behind Failover routing, and it also works within Weighted/Latency/Multi-value policies to skip a failed target.
Can also health-check a CloudWatch alarm instead of an endpoint directly, useful for checking something not reachable over the network (e.g. a database's internal metric).
🏷️ Domain Registration
Route 53 can also act as the registrar, the entity that reserves a domain name for you (separate from DNS hosting, though it usually sets up a matching Hosted Zone automatically).
You can register a domain elsewhere and still point it at a Route 53 Hosted Zone: registrar and DNS host don't have to be the same service.
🧠 Quick Memory Hooks
Alias = the only way to point the bare domain at an AWS resource, and it's free.
Latency = fastest route. Failover = backup plan. Weighted = percentage split. Geolocation = where the visitor is from.
Health checks are what let Route 53 skip a broken record instead of blindly handing it out.
πŸ“‘ CloudFront & Edge: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
CloudFront is AWS's CDN: it caches your content at locations physically close to viewers, so most requests never have to travel back to the origin at all.
🎯 Origins
The Origin is where CloudFront fetches content from when it doesn't already have a fresh cached copy.
S3 OriginStatic assets, website hosting (the most common origin)
Custom OriginAnything with a public HTTP(S) endpoint: an ALB, EC2 instance, or even a non-AWS server
One distribution can route to multiple origins based on the request path (Origin Groups can also fail over from a primary to a secondary origin).
πŸ” OAC (Origin Access Control)
Locks an S3 origin down so it's only reachable through CloudFront; direct requests straight to the bucket's S3 URL are rejected.
Works by having CloudFront sign every request to the origin with SigV4; the bucket policy then only trusts that specific distribution.
OAC replaces the older OAI (Origin Access Identity). It's AWS's current recommendation for any new distribution. Exam material may still mention OAI by name; know it's the predecessor being phased out, same underlying purpose.
⚑ Caching Behavior
Viewer Edge Location (cache) Origin Request Response (cached or fresh) Cache miss only: fetch from origin, then cache it A cache hit never touches the origin at all. This is the entire performance and cost benefit of a CDN.
TTL controls how long an object stays cached before CloudFront re-checks the origin. Can be set per-object (origin's Cache-Control header) or overridden at the distribution level.
Cache key can include (or ignore) specific headers, cookies, and query strings. Configuring this correctly matters: caching per-user content under one shared key would leak one viewer's response to another.
Invalidation forces CloudFront to drop cached copies of specific paths before their TTL expires, useful right after a deploy, but has a cost per invalidation path at scale (versioned filenames are the cheaper long-term pattern).
🌍 Edge Locations vs. Regional Edge Caches
Edge Locations: hundreds of small points of presence worldwide: the first stop for a viewer's request, holding the most popular/recently-requested content.
Regional Edge Caches: fewer, larger caches sitting between edge locations and the origin; catch less-popular content that fell out of an edge location's cache, reducing how often the origin itself gets hit.
🎟️ Signed URLs & Signed Cookies
Restrict access to private content served through CloudFront. A request without a valid signature is rejected.
Signed URLOne URL, one object (or a few), e.g. a single paid download link with an expiry
Signed CookieOne cookie grants access to many objects, e.g. an entire video course a logged-in user has paid for
πŸ›‘οΈ AWS WAF & AWS Shield (frequently tested!)
Two different edge-protection services, commonly confused because both attach to the same resources (CloudFront, ALB, API Gateway, AppSync); they defend against completely different attack types.
AWS WAFAWS Shield
Protects againstApplication-layer (Layer 7) exploits: SQL injection, cross-site scripting (XSS), bad bots, rate-based abuseDDoS attacks: Layer 3/4 (and, with Advanced, Layer 7) volumetric/protocol floods
How it worksYou define rules (or use AWS Managed Rule groups) that inspect each request and Allow/Block/Count it. This is where "block SQLi/XSS" configuration actually livesDetects and automatically mitigates attack traffic, nothing to author yourself
TiersPay per rule + per request; no free tier, but no separate Standard/Advanced split eitherStandard: free, automatic, protects every AWS customer by default. Advanced: paid, adds near-real-time attack visibility, cost protection against scaling charges caused by an attack, and 24/7 access to the AWS Shield Response Team (SRT)
Exam pattern: "block SQL injection / cross-site scripting requests" β†’ AWS WAF with a rule/managed rule group, never Shield (Shield has no concept of inspecting request content for an exploit pattern). "Protect against a DDoS attack" β†’ Shield. AWS Firewall Manager is a distinct third service: it centrally deploys and enforces WAF rules (and Shield Advanced protections) across many accounts/resources in an Organization at once, rather than doing the blocking itself.
🧠 Quick Memory Hooks
OAC = the lock on the origin door; only CloudFront has the key.
Edge Location = the corner shop (fast, small, popular items). Regional Edge Cache = the regional warehouse behind it.
Signed URL = one ticket, one item. Signed Cookie = a season pass to everything.
WAF reads the request content (SQLi/XSS rules). Shield just absorbs the flood (DDoS); it never looks inside a request.
⚑ Lambda & Serverless: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Lambda runs your code without a server to manage: you pay per invocation/duration, and AWS handles all the provisioning. The exam mostly tests how a given trigger invokes it, not the code itself.
🧩 Lambda Basics
Max execution time: 15 minutes per invocation. Long-running work beyond that needs Step Functions, ECS/Fargate, or a different compute model entirely.
Memory (128 MB–10,240 MB) and CPU are coupled: you only set memory, and CPU allocation scales with it automatically.
Charged per invocation + duration (rounded to the millisecond): near-zero cost when nothing is running, unlike an idle EC2 instance you're still paying for.
πŸ”€ Invocation Types (frequently tested!)
Synchronous Caller (API GW/ALB) Lambda Caller waits for response Asynchronous Event (S3, SNS) Lambda DLQ / failure dest. Fire-and-forget; retries on failure Poll-based (event source mapping) SQS / Kinesis / Streams Lambda AWS polls the source, invokes with a batch The trigger's identity determines the invocation type, not something you choose independently.
TypeExample triggersRetry behavior
SynchronousAPI Gateway, ALB, direct SDK callNone: caller sees the error immediately, handles retry itself
AsynchronousS3, SNS, EventBridgeAutomatic retries, then a Dead-Letter Queue or on-failure Destination if still failing
Poll-basedSQS, Kinesis, DynamoDB StreamsLambda service itself polls and re-tries batches; a persistently-failing record can block the batch, hence configurable retry/skip settings per source
🧡 Concurrency
Lambda scales by running many concurrent copies of your function, one per in-flight invocation, not by scaling a single instance's throughput.
Reserved ConcurrencyCaps how many concurrent executions a function can use; protects other functions from being starved of the account-wide limit, but throttles this one once the cap is hit
Provisioned ConcurrencyKeeps a set number of execution environments pre-initialized and warm; eliminates cold starts for that reserved capacity, at a standing cost even when idle
❄️ Cold Starts
A cold start happens when Lambda has to spin up a brand-new execution environment (download code, start the runtime) before running your function. This adds noticeable latency to that one request.
A warm invocation reuses an already-initialized environment from a recent invocation: fast, no setup overhead.
Mitigations: Provisioned Concurrency (guaranteed warm), smaller deployment packages, avoiding heavyweight SDK initialization in the function's global scope where it can't be reused efficiently.
πŸ”’ Lambda + VPC (frequently tested!)
By default a Lambda function runs outside any customer VPC, in an AWS-managed network: it can reach the public internet and any public AWS service endpoint, but has no route to a private resource (an RDS instance in a private subnet, an internal-only ElastiCache cluster).
To reach a private VPC resource, configure the function with VPC settings: the subnets and security group(s) it should launch into. Lambda then attaches an ENI in those subnets to route traffic, the same fundamental mechanism an EC2 instance uses.
Trade-off: a VPC-attached function loses default internet access. Once attached to private subnets, outbound internet calls (e.g. to a public third-party API) need the same NAT Gateway/Instance path any other private-subnet resource would use. Attaching to a VPC does not, by itself, add internet access, it only adds VPC access.
Historical cold-start cost (largely resolved): earlier Lambda networking created a new ENI per unique subnet/security-group combination on a cold start, which could take tens of seconds, a frequently-cited exam gotcha. AWS's Hyperplane-based networking model now pre-creates and shares ENIs across functions/accounts, so a modern VPC-attached function's cold start is close to a non-VPC function's. Older exam material may still describe VPC attachment as a heavy cold-start penalty; know both the old reputation and the current reality.
Exam pattern: "Lambda function times out trying to reach an RDS instance in a private subnet" β†’ the function isn't configured with that VPC's subnets/security group yet. "VPC-attached Lambda function can no longer reach a public API it used to call" β†’ it needs a NAT Gateway/Instance route from those subnets now, the same as any private-subnet resource.
πŸšͺ API Gateway
The front door for a serverless API: routes HTTP requests to backends (usually Lambda), handles auth, throttling, and request/response transformation.
REST APIFull feature set: request validation, API keys/usage plans, caching. More expensive, more setup.
HTTP APINewer, lighter, cheaper (~70% less); covers the common Lambda-backend case without the full REST API feature set
Exam pattern: "cheapest way to expose a Lambda function over HTTP with no advanced needs" β†’ HTTP API.
πŸͺœ Step Functions
Orchestrates a sequence of steps (a state machine): Lambda calls, waits, parallel branches, choice logic, error handling/retries, without writing that coordination logic yourself.
Standard workflowsLong-running (up to 1 year), exactly-once execution, full execution history; for auditable, durable business processes
Express workflowsHigh-volume, short-duration (up to 5 min), at-least-once execution; for fast event-processing pipelines
Exam pattern: "coordinate multiple Lambda functions with retries, branching, and a visual audit trail" β†’ Step Functions, instead of hand-rolling orchestration logic inside one giant Lambda.
🧠 Quick Memory Hooks
Sync = wait on hold. Async = leave a voicemail (queued, retried). Poll-based = AWS checks your mailbox for you.
Provisioned Concurrency = pay to keep the engine running so it's never cold.
Step Functions = the flowchart that runs itself.
βœ‰οΈ Messaging: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Messaging services exist to decouple producers from consumers, so one side can fail, slow down, or scale independently without taking the other down with it.
πŸ”“ Why Decouple?
Tightly coupled: Service A calls Service B directly and waits. If B is slow or down, A is stuck too.
Loosely coupled: Service A drops a message somewhere durable; Service B picks it up whenever it's ready. A never waits on B's health or speed.
This is also what absorbs a traffic spike: messages queue up instead of overwhelming the consumer, which processes at its own sustainable pace.
πŸ“¬ SQS: Standard vs. FIFO
StandardFIFO
OrderingBest-effort, not guaranteedStrict, first-in-first-out
DeliveryAt-least-once (a duplicate is possible)Exactly-once processing
ThroughputNearly unlimitedUp to 3,000 msg/sec with batching (per API action)
Use caseMost workloads (order/duplication don't matter)Order-sensitive: financial transactions, sequential commands
Visibility TimeoutOnce a consumer receives a message, it's hidden from other consumers for this window. If the consumer doesn't delete it in time (crashed, still processing), it reappears for someone else to pick up
Dead-Letter Queue (DLQ)After a message fails processing too many times (maxReceiveCount), it's moved here instead of retrying forever. This lets you inspect poison-pill messages without blocking the main queue
Long PollingConsumer's receive call waits up to 20s for a message to arrive instead of returning empty immediately; cheaper (fewer empty API calls) than short polling
πŸ“’ SNS & the Fan-out Pattern (frequently tested!)
SNS is pub/sub: one message published to a Topic is pushed out to every subscriber at once (email, SMS, Lambda, HTTP endpoint, or an SQS queue).
Producer SNS Topic SQS Queue A SQS Queue B SQS Queue C One publish, three independent queues: each consumer processes at its own pace, and a slow/broken consumer B doesn't affect A or C.
This SNS→multiple-SQS pattern is the standard answer whenever a question describes "one event needs to trigger several independent downstream processes."
🚌 EventBridge
An event bus: routes events matching a rule (pattern-matched on the event's content) to one or more targets, from a much wider range of sources than SNS.
SourcesAWS services natively, your own custom application events, and direct integrations with 3rd-party SaaS (e.g. Zendesk, Datadog)
SchedulingBuilt-in cron/rate-based scheduled rules, no need for a separate scheduler service
Content filteringRules can match on the event's actual field values, not just its source; more routing intelligence than an SNS topic
🧭 SQS vs. SNS vs. EventBridge: how to choose
SQSOne producer, one (logical) consumer group, need durable buffering/retry: a queue, not a broadcast
SNSOne event, multiple known subscribers, need it pushed out immediately (pub/sub)
EventBridgeEvent-driven architecture with content-based routing rules, scheduled events, or ingesting from SaaS/3rd-party sources
🧠 Quick Memory Hooks
SQS = a mailbox (pull, buffered). SNS = a megaphone (push, broadcast). EventBridge = a smart switchboard (rules-based routing).
Standard SQS = fast, maybe-duplicate. FIFO SQS = strict order, exactly-once, capped throughput.
Fan-out = SNS topic, many SQS queues, each consumer independent.
πŸ“Š Monitoring & Management: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Four different questions, four different tools: Is it healthy right now? (CloudWatch) Β· Who did what? (CloudTrail) Β· What changed, and when? (Config) Β· Am I following best practice? (Trusted Advisor).
πŸ“ˆ CloudWatch
MetricsTime-series numeric data (CPU%, request count, queue depth). Most AWS services publish these automatically; custom metrics can be pushed from your own app
AlarmsWatch a metric against a threshold and trigger an action (SNS notification, Auto Scaling policy, EC2 Auto Recovery) when breached
LogsCentralized, searchable log storage: application/Lambda/VPC Flow Logs all ship here; Logs Insights lets you query them
DashboardsCustom visual boards combining metrics/alarms from across services into one view
Resolution: varies by service, e.g. EC2's basic monitoring is 5-minute granularity by default, with Detailed Monitoring (an extra cost) dropping that to 1-minute for faster alarm reaction; many other services (e.g. Lambda) already publish at 1-minute resolution with no separate "detailed" tier. Don't assume 5-minute-by-default applies account-wide.
βš–οΈ CloudWatch vs. CloudTrail (frequently tested!)
CloudWatchCloudTrail
Answers"Is the resource healthy / how is it performing?""Who called which API, when, from where?"
DataMetrics, logs, alarms (operational/performance data)An audit log of every API call (console, CLI, SDK) in the account
Enabled by default?Basic metrics yes, most other features opt-inEvent history exists automatically with no setup (management events, rolling 90-day window), but that's not a "trail": for continuous delivery to S3/CloudWatch Logs and retention beyond 90 days, you must explicitly create a Trail
Exam pattern: "someone deleted a security group, find out who" β†’ CloudTrail. "The app's error rate spiked, alert the team" β†’ CloudWatch.
πŸ•°οΈ AWS Config
Records a configuration history of your resources: what a resource's settings looked like at any past point in time, and every change in between.
Config Rules continuously check resources against a desired configuration (e.g. "every EBS volume must be encrypted") and flag non-compliant ones. This is compliance/drift detection, distinct from CloudTrail's "who did it" audit trail.
Exam pattern: "prove a security group's rules over the last 3 months" or "detect when a resource drifts out of compliance" β†’ Config, not CloudTrail.
πŸ’‘ Trusted Advisor
Automated checks across your account against AWS best practice, grouped into 6 categories: Cost Optimization, Performance, Security, Fault Tolerance, Service Limits, Operational Excellence.
Free tier gives a limited set of core checks (mostly security); the full check list requires a Business or Enterprise Support plan.
Overlaps in spirit with Config/Compute Optimizer but is broader and higher-level: a quick account-wide health/hygiene scorecard rather than deep resource history.
πŸ› οΈ Systems Manager (SSM): brief mention
Parameter StoreFree, secure key-value config/secrets storage, a lighter alternative to Secrets Manager when rotation isn't needed
Session ManagerShell access to an instance with no open SSH port, no bastion host, no key pair; access is governed entirely by IAM
🩻 AWS X-Ray (frequently tested!)
Distributed tracing: answers a different question than everything else on this page. CloudWatch tells you a metric is bad; X-Ray tells you where inside a multi-service request the time actually went.
A single request (e.g. API Gateway β†’ Lambda β†’ DynamoDB) is stitched into one trace, made of per-service segments (and finer subsegments within a service, e.g. one for the DynamoDB call specifically), visualized as a service map showing every hop and its latency.
Requires minimal code change: the X-Ray SDK (or, for supported services, a built-in "Active tracing" toggle; Lambda and API Gateway both have one) instruments the service without a rewrite.
Exam pattern: "find which specific downstream service/call is causing latency in a chain of microservices" β†’ X-Ray, not CloudWatch (CloudWatch alerts you that latency is high overall; it doesn't show you the call chain that produced it).
🧠 Quick Memory Hooks
CloudWatch = the vital-signs monitor. CloudTrail = the security camera log. Config = the change-history book. Trusted Advisor = the health inspector's checklist. X-Ray = the delivery tracking map, hop by hop.
If the question is about who, think CloudTrail. If it's about what changed over time, think Config. If it's about right now, think CloudWatch. If it's about where in the chain the time went, think X-Ray.
πŸ›οΈ Well-Architected: One-Page Study Guide
πŸ” Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
The Well-Architected Framework is AWS's lens for judging a design: many SAA questions are really asking "which pillar does this trade-off belong to," not testing new facts.
πŸ›οΈ The Six Pillars (frequently tested!)
PillarCore question
Operational ExcellenceCan we run and monitor this reliably, and improve it over time (as code, with automation)?
SecurityIs access, data, and infrastructure protected: least privilege, encryption, traceability?
ReliabilityDoes it recover from failure and meet demand without manual intervention (Multi-AZ, Auto Scaling, backups)?
Performance EfficiencyAre we using the right resource type/size for the workload, and adapting as it evolves?
Cost OptimizationAre we avoiding unnecessary spend: right-sizing, the right purchasing option, eliminating waste?
SustainabilityAre we minimizing the environmental impact of running this workload (added as the 6th pillar in 2021)?
Exam pattern: a question describing "encrypt data, apply least privilege" is testing Security; "add Multi-AZ, automate failover" is Reliability; "switch to Spot, right-size instances" is Cost Optimization. Even if the word "pillar" never appears, that's the lens being tested.
🀝 Shared Responsibility Model
AWS: Security of the CloudCustomer: Security in the Cloud
Physical data centers, hardware, networking infrastructureData encryption, IAM configuration, security group/NACL rules
The virtualization layer (hypervisor)Guest OS patching, application-level security
Managed service internals (e.g. RDS engine patching)What you put in the database, who can access it
Rule of thumb: AWS secures the infrastructure the cloud runs on; you secure everything you configure and put in it. For managed services (RDS, Lambda), the line shifts further toward AWS than for unmanaged ones (raw EC2), but data and access control are always yours.
πŸ’° Cost Optimization Levers: a recap across this guide
Cost Optimization isn't its own set of new facts. It's the same tools covered elsewhere in this guide, applied with cost as the goal:
1Right-size: don't pay for provisioned capacity you don't use (EC2 & Compute topic)
2Match the purchasing option to the workload: Reserved/Savings Plans for steady-state, Spot for fault-tolerant/flexible (EC2 & Compute topic)
3Lifecycle old data to cheaper storage tiers automatically instead of leaving everything on Standard (S3 & Storage topic)
4Cache aggressively: CloudFront and ElastiCache both cut repeated, avoidable work (CloudFront & Edge, Databases topics)
5Go serverless where usage is spiky/idle: Lambda's near-zero idle cost beats a constantly-running server for bursty workloads (Lambda & Serverless topic)
🧰 AWS Well-Architected Tool
A free console tool that walks through a structured questionnaire per pillar against your actual workload, then flags risks and links out to the specific guidance for fixing each one.
Not something you configure once, meant to be revisited as a workload evolves, the same way the pillars are a lens you keep re-applying, not a one-time checklist.
🧠 Quick Memory Hooks
Six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, Sustainability. "Our Systems Run Pretty Cheap, Sustainably."
AWS secures the cloud. You secure what's in it.
Every cost-optimization answer on this exam is really "stop paying for unused/oversized/idle capacity" in a new outfit.
🏷️ AWS Service Names: Exam Reference
AWS Certification exams reduce reading load by using official short names for certain well-known services. Know both the short name and the full name: a question may use either.
🏷️ Official Short Names (frequently tested!)
Short NameFull NameExample use case
AWS CDKAWS Cloud Development KitDefine your infrastructure in Python/TypeScript instead of hand-writing raw CloudFormation YAML
AWS CLIAWS Command Line InterfaceScript aws s3 cp in a deploy pipeline instead of clicking through the console
AWS DMSAWS Database Migration ServiceMigrate an on-prem Oracle database into RDS with minimal downtime
Amazon DocumentDBAmazon DocumentDB (with MongoDB compatibility)Run a MongoDB-style workload without managing MongoDB servers yourself
Amazon EBSAmazon Elastic Block StoreAttach a persistent virtual hard drive to an EC2 instance
Amazon EC2Amazon Elastic Compute CloudLaunch a virtual server to host an application
Amazon ECRAmazon Elastic Container RegistryStore and version your Docker container images before deploying them
Amazon ECSAmazon Elastic Container ServiceRun and orchestrate Docker containers without managing Kubernetes yourself
Amazon EFSAmazon Elastic File SystemShare one common file system across many EC2 instances at once
Amazon EKSAmazon Elastic Kubernetes ServiceRun a managed Kubernetes cluster when you specifically need Kubernetes
IAMAWS Identity and Access ManagementAdd a new user, create a Role, or attach a permissions policy
Amazon KeyspacesAmazon Keyspaces (for Apache Cassandra)Run a Cassandra-compatible workload without managing Cassandra nodes
AWS KMSAWS Key Management ServiceCreate and manage the encryption key used to encrypt an S3 bucket or EBS volume
AWS Managed Microsoft ADAWS Directory Service for Microsoft Active DirectoryStand up a real Active Directory domain for Windows workloads, without running your own domain controllers
AWS Private CAAWS Private Certificate AuthorityIssue private TLS certificates for internal services that don't need a public CA
Amazon RDSAmazon Relational Database ServiceLaunch a managed MySQL/PostgreSQL database without patching or backing it up yourself
Amazon S3Amazon Simple Storage ServiceStore and serve static files, backups, or website assets
AWS SAMAWS Serverless Application ModelDefine and deploy a serverless Lambda application from a simplified template
AWS SCTAWS Schema Conversion ToolConvert a database schema from one engine (e.g. Oracle) to another (e.g. PostgreSQL) before migrating with DMS
Amazon SESAmazon Simple Email ServiceSend transactional or marketing emails from an application
Amazon SNSAmazon Simple Notification ServiceFan out one notification to many subscribers (email, SMS, Lambda, SQS) at once
Amazon SQSAmazon Simple Queue ServiceDecouple two application components with a durable message queue
AWS STSAWS Security Token ServiceIssue the temporary credentials handed out when a Role is assumed
Amazon VPCAmazon Virtual Private CloudCreate an isolated private network to launch your resources into
Source: AWS Certification General Information policy page, the authoritative, exam-official list. Check back there periodically since AWS can add to it.
Exam tip: if a question describes a scenario rather than naming a service, match the verb: "issue temporary credentials" β†’ STS, "decouple" β†’ SQS, "fan out" β†’ SNS, "add a user" β†’ IAM.
🧠 Memory hook: pair the ones that get confused. SCT before DMS: convert the Schema, then migrate the Data. SNS pushes out, SQS holds until pulled. EBS = one disk for one instance, EFS = one shared filesystem for many instances, S3 = a bucket, not a drive at all.
🧠 Quick Memory Hooks
This isn't a technical domain. It's exam vocabulary. A question naming "AWS STS" or "Amazon Keyspaces" cold is testing whether you recognize the service, not a new concept.
Acronyms you'll see spelled out constantly elsewhere in this guide (IAM, KMS, RDS, VPC, SNS, SQS) are exactly the ones on this list. That's not a coincidence.