One-page study guides for the AWS Certified Solutions Architect β Associate (SAA-C03) exam. Pick a topic on the left. Official exam page β
π IAM: One-Page Study Guide
Identity (Who) + Policy (What's allowed) β access to a ResourceIAM is global (not region-specific) and free.
π€Identities
Type
What it is
Credentials
User
A person or app (long-term identity)
Password (console) and/or Access Keys (CLI/API)
Group
A bucket of Users, for bulk permission assignment
None: can't log in as a Group, can't nest groups
Role
A temporary identity "assumed" by a User, AWS service, or federated identity
None stored: temporary credentials issued via STS, auto-expire
Rule of thumb: Users/apps needing standing access β User. Anything temporary or service-to-service (e.g. EC2 β S3) β Role. Role is always the more secure/best-practice choice when it fits.
Service-linked role: a distinct Role subtype. Some AWS services (e.g. Auto Scaling, RDS) create and manage their own Role automatically the first time you use a feature that needs it. You don't create it, and you usually can't edit its permissions directly. It's pre-linked to that one service's use case and AWS controls when it can be deleted (typically only once nothing is using it). A common wrong-answer distractor: it's still an IAM Role, just not one you author yourself.
Cross-account Role + External ID (frequently tested!) A Role's trust policy can name another AWS account (not just a service) as the principal allowed to assume it. This is how one account (e.g. a monitoring/audit vendor, or a security team's account) gets scoped, temporary access into another without a User ever existing in the target account. When the trusting account and the assuming account belong to different organizations (a third-party SaaS vendor, not your own AWS Organization), add an External ID (a shared secret string the third party must supply in its AssumeRole call) to the trust policy's condition. This defeats the "confused deputy" problem: without it, the third party could be tricked into assuming a role on behalf of one of its other customers into your account, because the trust policy alone can't tell which of the vendor's own customers is actually asking. An External ID is unnecessary for a Role assumed purely within your own Organization (IAM Identity Center Permission Sets, member-account access); it specifically targets the third-party-vendor scenario.
An Access Key has no permissions of its own. It's just a credential pair (Access Key ID + Secret Access Key) that proves "I am this User". The actual permissions come entirely from whatever Policies are attached to that User, directly or via a Group. Two keys belonging to the same User are equally powerful; rotating or deactivating a key changes what can authenticate, not what the User is allowed to do.
Two ways to get credentials, and they're not interchangeable. (1) Long-term: a User's Access Key ID + Secret Access Key, valid until rotated/deleted, permissions = that User's attached Policies. (2) Temporary: a Role is assumed via STS (AssumeRole), which hands back a short-lived Access Key ID + Secret Access Key + session token that auto-expires, permissions = that Role's attached Policies. Same credential shape (a key pair), completely different lifecycle and source of permissions. This is the exam distinction, not "Users have keys, Roles don't."
IAM Role (temporary)
Access Key (User, long-term)
Where credentials live
Nowhere persistent: for an AWS service (EC2 instance profile, Lambda execution role) they're fetched from instance metadata / injected at invoke time and held only in memory; never written to disk
Typically ~/.aws/credentials on disk, an environment variable, or (worst case) hardcoded in app config/source, a real, repeatedly-exploited leak surface: committed to a public repo, left on a compromised laptop, pulled via an IMDSv1 SSRF bug
Lifetime
Short-lived, auto-expires (as little as 15 min, default 1 hr, up to 12 hr for most AssumeRole calls); a leaked one is only useful for a limited window
Indefinite: stays valid until someone manually deactivates or deletes it, which is exactly why forgotten/leaked keys are such a common real-world breach vector
Rotation
None needed: a brand-new set is issued every time the Role is assumed
Manual, and easy to forget (this is AWS's official best-practice topic #4, below)
Setup cost
More upfront config: a trust policy (who/what may assume it) plus the permission policies themselves
One click to generate: simple, which is also why it's overused where a Role would be safer
Best for
AWS services acting on your behalf (EC2, Lambda, ECS), cross-account access, federated/SSO human logins, anywhere the caller can be handed a credential rather than typing one in
Genuinely long-term, non-interactive use where nothing can assume a Role for you (some legacy tools/third-party integrations with no role-assumption support)
The "no keys stored" win is strongest for AWS services, not for a human at a CLI. An EC2 instance profile or Lambda execution role genuinely never has any static credential anywhere. That's the classic exam scenario above. A person running the CLI on their own laptop still needs to authenticate the AssumeRole call itself somehow: either with a User's long-term Access Key sitting in ~/.aws/credentials as a source_profile (which only partially avoids the problem: the day-to-day CLI credentials are temporary, but a static key still exists behind the scenes) or, cleaner, via IAM Identity Center / SSO (see below), which needs no long-term key at all. Don't over-claim "Roles mean nothing is ever in ~/.aws" for the human case. It depends on how the Role is being assumed.
πPolicies
JSON documents with statements: Allow or Deny on specific Actions + Resources.
Version: the policy language version, always 2012-10-17 (the current/only version; don't overthink it). Statement: an array of one or more permission blocks, each with an Effect (Allow/Deny), one or more Action(s) (the API action(s) being granted, e.g. s3:PutObject), and one or more Resource(s) (the ARN(s) it applies to, wildcards allowed).
Same shape, just Effect: Deny. This blocks s3:DeleteObject on the bucket no matter what any other Allow statement says, on this policy or any other attached to the identity. This is the "Explicit Deny always wins" rule below in practice.
Identity-basedAttached to a User/Group/Role
Resource-basedAttached to the resource itself (e.g. S3 bucket policy)
AWS ManagedCreated/maintained by AWS, can't edit
Customer ManagedYou create it, reusable across users/groups/roles in your account
InlineEmbedded 1:1 in a single User/Group/Role; deleted when that entity is deleted
AWS recommends Customer Managed over Inline in most cases.
IAM Policy Simulator tests what a policy actually allows/denies before you attach it; IAM Access Analyzer continuously scans your account and flags unintended external/public access.
βοΈEvaluation Logic (frequently tested!)
1Default = implicit deny (nothing is allowed unless stated)
2Explicit Allow overrides the implicit deny
3Explicit Deny always wins: overrides any Allow, no matter where it comes from
Explicit Deny is checked first and short-circuits everything else. This is why "Deny always wins" regardless of how many Allow statements exist.
Advanced: A Permissions Boundary sets the maximum permissions an entity (User/Role) can have, no matter what its identity policies allow. Org SCPs work the same way but at the account/OU level. More relevant at Pro level, but good to recognize.
π₯οΈRoles + EC2 (classic exam scenario)
EC2 instance needs an Instance Profile containing exactly one Role.
App on the instance retrieves temporary credentials automatically from instance metadata, no keys to manage.
This is always preferred over storing a User's Access Keys on the instance.
No Access Keys ever touch the instance; the app fetches short-lived credentials from instance metadata and AWS handles the rest.
Terminology note: the temporary credentials STS hands back are technically the same shape as a long-term Access Key: an Access Key ID + Secret Access Key, plus a session token. "No Access Keys" above means no long-term User Access Key is created, stored, or exposed on the instance, not that the credential format is different. See the "Two ways to get credentials" note under Identities above.
The launching User and the running instance are two separate principals: a classic gotcha. Creating the EC2 instance only needs ec2:RunInstances (and related) permissions on your User/Role. Once it's running, anything the instance itself does (an app calling S3, a script calling DynamoDB, or you SSHing in and running the AWS CLI from inside it) authenticates as the instance profile's Role, not as the User who launched it. Your own permissions don't carry over onto the box at all. This is why a full-admin User can launch an instance whose app then gets AccessDenied calling S3; the Role attached to the instance simply wasn't granted s3:*, regardless of what the launching User could do.
You can bridge that gap with a User's Access Key, but it's an anti-pattern. It's technically possible to put a User's long-term Access Key ID + Secret Access Key on the instance (env vars, aws configure, a config file) so the instance authenticates as that User instead of via its Role. This is exactly the "storing a User's Access Keys on the instance" practice called out as never-preferred above: it reintroduces a static, leakable, manually-rotated credential on a box that could have had none at all. The exam-correct answer for "an EC2 instance needs AWS access" is always an Instance Profile Role; a hardcoded Access Key on the instance is the wrong-answer distractor.
πAuthentication Methods
Access via
Credential
Console
Username + Password (+ MFA recommended)
CLI / API (User)
Access Key ID + Secret Access Key (long-term; avoid where possible)
CLI / API (Role)
Temporary credentials via STS (Security Token Service): short-lived, auto-expire
Every action you take (console click, CLI command, SDK call) is an API call. The Console is just a UI wrapper around the same AWS APIs the CLI/SDK use directly. This is why IAM policies grant/deny specific Actions (e.g. s3:PutObject) rather than "console access" vs "CLI access". The permission is on the underlying API action, not the interface used to trigger it. It's also why CloudTrail can log every single one, regardless of which interface made the call.
π’IAM Identity Center (formerly AWS SSO)
Solves a different problem than everything above: one login for a workforce across many AWS accounts (and even non-AWS apps), instead of a separate IAM User in every account.
What it replacesCreating individual IAM Users in each member account of an Organization; instead, people sign in once and get a portal listing every account/role they're allowed into
Permission SetsReusable templates of permissions assigned to a user/group per-account. Under the hood these provision IAM Roles in the target account, so it's built on the same Role mechanics as above, just centrally managed
Identity sourceIts own built-in directory, or federate from an external IdP (Microsoft AD, Okta, Azure AD, etc.) via SAML
Exam signal: a question about one person needing access to multiple AWS accounts without a separate User in each one is pointing at IAM Identity Center, not at creating more IAM Users or hand-rolling cross-account Roles yourself. It's an AWS Organizations-level feature, not something you turn on inside a single account in isolation.
ποΈAWS Organizations & Service Control Policies (SCPs) (frequently tested!)
AWS Organizations groups multiple AWS accounts (e.g. one per team, environment, or department) under one management account, arranged into Organizational Units (OUs), folders you can nest and apply policy to as a group instead of account-by-account.
A Service Control Policy (SCP) is a guardrail attached to the whole Organization, an OU, or a single account. It sets the maximum available permissions for every principal in scope, including that account's own root user/admin. An SCP never grants anything by itself (it has no effect on a management account); it only restricts what IAM policies inside the account are allowed to permit.
IAM Policy
Service Control Policy (SCP)
Attached to
A User, Group, or Role
An AWS account, an OU, or the whole Organization
Can it grant access?
Yes, an Allow here actually grants permission
No, it only sets a ceiling; an Allow in an SCP does nothing on its own without a matching IAM policy Allow
Affects the account's root user?
No, Users/Roles only
Yes, even the account root/admin cannot exceed an SCP's ceiling
Typical use
Day-to-day least-privilege access for a specific identity
Org-wide guardrails: block a whole category of action everywhere it applies, regardless of any individual permissions
Final permission = the intersection of both. A principal can only actually do something if an IAM policy allows it and no applicable SCP blocks it. An SCP Allow plus an IAM policy Deny still results in Deny (same "explicit Deny always wins" rule as above, just evaluated across two separate policy layers instead of one).
Common exam SCP patterns: deny cloudtrail:StopLogging org-wide so no admin in any account can quietly disable audit logging; deny ec2:RunInstances above a certain instance size in a sandbox/dev OU to cap accidental cost; deny actions outside an approved AWS Region.
Exam pattern: "must apply even if the user/account has full admin/root privileges, with the least operational overhead across many accounts" β an SCP at the OU level, not a per-account IAM policy repeated in every account. That's the "least overhead across an Organization" signal.
β Best Practices: the ones that actually get tested
Lock away the root account: don't use it day-to-day; create an admin IAM User instead.
Enable MFA, especially for privileged/root accounts.
Least privilege: grant only what's needed.
Prefer Roles over long-term Access Keys wherever possible.
Use Groups to manage User permissions at scale, not per-user policies.
1Require human users to use federation with an identity provider for temporary credentials
2Require workloads to use temporary credentials with IAM Roles
3Require multi-factor authentication (MFA)
4Update access keys when needed, for use cases that genuinely require long-term credentials
5Follow best practices to protect your root user credentials
6Apply least-privilege permissions
7Get started with AWS managed policies, then move toward least privilege
8Use IAM Access Analyzer to generate least-privilege policies based on access activity
9Regularly review and remove unused users, roles, permissions, policies, and credentials
10Use conditions in IAM policies to further restrict access
11Verify public and cross-account access to resources with IAM Access Analyzer
12Use IAM Access Analyzer to validate your IAM policies for secure and functional permissions
13Establish permissions guardrails across multiple accounts (Org SCPs/RCPs)
14Use permissions boundaries to delegate permissions management within an account
π§ Quick Memory Hooks
User = person Β· Group = filing cabinet for people Β· Role = borrowed hat Β· Policy = rulebook Deny always beats Allow. No standing credentials = Role = best practice.
π VPC & Networking: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
A VPC is your own isolated network inside a Region; everything else on this page is about controlling what can reach what, inside it and in/out of it.
See the EC2 & Compute topic for the subnet/AZ/ENI layout diagram. This page covers the routing and filtering layer on top of it.
πVPC Basics: CIDR sizing
A VPC is defined by a CIDR block (e.g. 10.0.0.0/16), the range of private IPs available inside it. Size: /16 to /28 (biggest to smallest).
Every subnet's CIDR must be a subset of its VPC's CIDR, and subnets in the same VPC can't overlap each other.
AWS reserves the first 4 and last 1 IP addresses in every subnet (network address, VPC router, DNS, future use, broadcast). A /24 subnet (256 addresses) really only gives you 251 usable.
A VPC can span multiple AZs (via multiple subnets) but always lives in exactly one Region.
πΊοΈSubnets & Route Tables
Every subnet is associated with exactly one route table (though one route table can serve many subnets). Unassociated subnets fall back to the VPC's main route table.
Every route table always has an implicit local route (traffic within the VPC's own CIDR) that can't be removed. This is what lets resources in different subnets of the same VPC reach each other by default.
A subnet counts as "public" purely because its route table has a route to an Internet Gateway; nothing else about the subnet changes.
Route table per AZ, for HA: a subnet lives in exactly one AZ, so routing each AZ's private subnet to that AZ's own NAT Gateway means giving each AZ its own route table. One shared route table pointing every private subnet at a single AZ's NAT Gateway would defeat the point of deploying one NAT Gateway per AZ.
πͺInternet Gateway vs. NAT Gateway vs. NAT Instance
Internet Gateway
NAT Gateway
NAT Instance
Purpose
Two-way internet access for a public subnet
Outbound-only internet for a private subnet
Same as NAT Gateway, but you manage it
Managed by
AWS (attach to VPC, no config)
AWS (provisioned, patched, scaled for you)
You; it's just an EC2 instance running NAT software
Availability
N/A (not a bottleneck resource)
Lives in one AZ; deploy one per AZ for HA
Single instance; you handle HA yourself
Bandwidth
N/A
Scales automatically up to 100 Gbps
Capped by the instance type you chose
Exam pattern: "instances in a private subnet need outbound internet only, minimal management" β NAT Gateway, essentially always the right answer over a NAT Instance now.
Exam pattern (HA, frequently tested!): "design must survive an AZ failure" + a NAT Gateway in the design β one NAT Gateway per AZ, each with its own private subnet's route table pointing to it. A single shared NAT Gateway is a single point of failure for every other AZ if its own AZ goes down.
π‘οΈSecurity Groups vs. NACLs (frequently tested!)
Every packet crosses the NACL first (subnet edge), then the Security Group (instance edge). Both must allow it.
Security Group
NACL
Applies to
The instance/ENI (instance-level firewall)
The subnet (network-level firewall). Affects every instance in it
Rules
Allow only (no explicit Deny)
Allow and Deny rules, evaluated in rule-number order, lowest first
State
Stateful: return traffic is automatically allowed, regardless of outbound rules
Stateless: inbound and outbound must each be explicitly allowed, even for a reply
Evaluation
All rules evaluated together; any match wins
Rules evaluated in order until the first match; a Deny can be hit before an Allow further down
Default
New SG: all inbound denied, all outbound allowed
Default NACL: allows all in/out; a custom NACL denies all until you add rules
VPC Flow Logs: captures metadata (source/dest IP, port, protocol, ACCEPT/REJECT, not the packet contents) for traffic at the VPC, subnet, or ENI level, sent to CloudWatch Logs or S3. The go-to tool for "why is my traffic being blocked": diagnosing which layer (SG or NACL) is actually rejecting a connection.
Two different "defaults": don't mix them up. The table's "Default" row describes a new SG you create: all inbound denied, all outbound allowed. The default security group that AWS auto-creates with every VPC is different again: it also allows all outbound, but its inbound rule is a self-reference (allow all traffic from any other resource that has that same default SG attached), not a flat deny. In practice this means two freshly launched instances left on the default SG can already talk to each other on every port, which surprises people expecting "deny all inbound" to be universal.
A Security Group rule's source/destination can be another Security Group, not just a CIDR (frequently tested!) This is the standard way to wire up a multi-tier design: the app-tier SG's inbound rule allows traffic from the web-tier SG (not from a CIDR range), and the database-tier SG's inbound rule allows traffic from the app-tier SG. Each tier only ever accepts traffic from the specific tier directly in front of it. Because the reference is to the group, not to fixed IPs, it stays correct automatically as instances in that tier scale in/out under an ASG, no CIDR range to keep updating. This is the best-practice answer whenever a question describes least-privilege network access between tiers, over widening a CIDR-based rule.
πVPC Peering
A private, 1:1 network connection between two VPCs (same or different accounts/Regions). Traffic stays on the AWS backbone, never touches the public internet.
Not transitive: if VPC A peers with B, and B peers with C, A cannot reach C through B; you'd need a direct AβC peering (or a Transit Gateway, at Pro-level scale).
The two VPCs' CIDR blocks must not overlap. Peering can't be established (or route tables can't be added) if they do.
Each side must add a route to the other's CIDR in its own route table; peering existing alone doesn't route traffic.
π―VPC Endpoints
Lets resources in a private subnet reach an AWS service without going through an Internet Gateway or NAT; traffic never leaves the AWS network.
Gateway EndpointOnly for S3 and DynamoDB. Free. Works by adding a target in your route table, no ENI involved.
Interface EndpointFor most other AWS services, powered by AWS PrivateLink. Creates an ENI with a private IP in your subnet. Small hourly + data cost.
Exam pattern: "private subnet, needs to reach S3, no internet access at all" β Gateway Endpoint (free) is the expected answer over a NAT Gateway (which costs more and routes through a public subnet's IGW anyway).
Scope limitation (frequently tested!) A VPC Endpoint (Gateway or Interface) only ever connects to an AWS service, or another VPC's own resource deliberately exposed as a PrivateLink endpoint service. It cannot reach an arbitrary third-party endpoint out on the public internet (an external payment processor's API, a SaaS vendor's REST API, etc.). For that, a private subnet still needs a NAT Gateway or NAT Instance routing out through an Internet Gateway. A VPC Endpoint is not a general substitute for NAT, only for the specific AWS services it supports.
Both connect an on-premises network to a VPC. The difference is what kind of link, not what it's used for.
Site-to-Site VPN
Direct Connect
What it is
An encrypted IPsec tunnel over the public internet
A private, dedicated physical network connection from your premises into AWS, never touches the public internet
Setup time
Minutes, fully software-configured
Weeks to months; needs a physical cross-connect at an AWS Direct Connect location (colo facility or partner)
Performance
Variable, subject to public internet congestion/latency
Consistent, low-latency, high-throughput: a real dedicated line
Cost model
Lower: pay for the VPN connection + data transfer
Higher fixed cost for the dedicated line, but often cheaper data transfer at scale
Typical use
Quick to stand up, backup link, lower-volume/less latency-sensitive traffic
Consistent large data volumes, latency-sensitive workloads, a strict "must not traverse the public internet" compliance requirement
Exam pattern: "must not traverse the public internet at all" or "consistent, dedicated bandwidth" β Direct Connect. "Quick to set up, encrypted, lower/no upfront hardware" β Site-to-Site VPN. The two are also commonly paired: a VPN as an automatic failover path if the Direct Connect link goes down.
π§ Quick Memory Hooks
NACL = bouncer at the building door (subnet, stateless, checks you both ways). Security Group = bouncer at your apartment door (instance, stateful, remembers who let you in). Peering is a one-hop friendship, not a chain. Gateway Endpoint = free, S3/DynamoDB only. Interface Endpoint = everything else, small cost, uses PrivateLink.
π₯οΈ EC2 & Compute: One-Page Study Guide
Pick an instance type for the workload, a purchasing option for the billing commitment, and a storage type for the data, three independent choices.
Security Groups/NACLs live in VPC & Networking; scaling policies live in ELB & Auto Scaling, both separate topics.
π·οΈInstance Type Families
Family
Optimized for
Example use case
General Purpose (M, T, A)
Balanced compute/memory/network
Web servers, small-to-medium apps, dev/test
Compute Optimized (C)
High vCPU-to-memory ratio
Batch processing, media transcoding, gaming servers, HPC
Memory Optimized (R, X, z)
Large RAM per vCPU
In-memory caches, large databases, real-time big-data analytics
Storage Optimized (I, D, H)
High, fast local disk I/O
NoSQL databases, data warehousing, distributed file systems
Accelerated Computing (P, G, Inf, Trn)
Hardware GPU/accelerator
ML training/inference, graphics rendering, video encoding
Naming decode (m6g.large):m = family, 6 = generation, g = processor attribute (g=AWS Graviton/ARM, a=AMD, n=network optimized, no letter=Intel), large = size. Higher generation number β newer, better price/performance. The exam often expects "pick the newer generation" as the answer when everything else ties.
T-family only is burstable: earns CPU credits at baseline, spends them to burst above baseline; Unlimited mode lets it keep bursting past its credit balance for a small extra charge instead of throttling back to baseline.
m5.largem = General Purpose family Β· 5 = generation 5 Β· large = instance size (2 vCPU / 8 GiB RAM)
r5.larger = Memory Optimized family Β· 5 = generation 5 Β· large = instance size (2 vCPU / 16 GiB RAM; same vCPU count as m5.large, double the RAM)
Common mix-up: the leading letter is the tell for the family, not the number: m5.large is General Purpose (the "m" is short for a balanced, "middle-of-the-road" mix of compute/memory), not memory-optimized. If a question is testing "high RAM per vCPU," the letter to look for is r (or x/z for even more extreme ratios). See the family table above.
πΊοΈWhere EC2 Lives: VPC, Subnets & ENIs
Terminology: Region vs. Availability Zone (AZ). A Region is a large geographic area (e.g. us-east-1, eu-west-2) that's fully independent of every other Region: its own copy of most services, its own data, nothing replicates between Regions unless you set that up yourself. Each Region is made up of multiple Availability Zones, one or more physically separate data centers within that Region, far enough apart to survive an independent failure (power, fire, flooding) but close enough together for low-latency links between them. Why it's tested: "high availability" almost always means spread across multiple AZs in one Region; "disaster recovery" almost always means spread across multiple Regions.
Every EC2 instance launches into a VPC, which spans a whole Region, inside one specific Availability Zone (AZ), inside one specific subnet in that AZ. There's no such thing as an instance outside a VPC, and a subnet never crosses AZ boundaries.
Public subnetIts route table has a route to an Internet Gateway (IGW). Instances here can hold both a private IP and a public IP/Elastic IP, and reach (and be reached from) the internet directly
Private subnetNo route to an IGW; instances here only ever get a private IP, never a public one; outbound-only internet access needs a NAT Gateway sitting in a public subnet
A subnet decides what a private IP can reach; a public/Elastic IP association decides whether an ENI is also reachable from the internet.
Default VPC: AWS auto-creates one per Region on your behalf. It comes pre-wired with a public subnet in every AZ, an IGW already attached, and a route table already sending 0.0.0.0/0 β IGW. Convenient for quickly launching a test instance with internet access, no networking setup required.
Custom VPC: you build the CIDR range, subnets, and route tables yourself. Internet connectivity is not automatic; you must create and attach an IGW, then add a route to it in the relevant subnet's route table before that subnet counts as "public." This is the standard, best-practice setup for real workloads (lets you control exactly which subnets are public vs. private).
Public IP by default: depends on which VPC you're in. Every subnet has an "auto-assign public IPv4" setting that decides whether a launched instance gets a public IP with no extra step. In the default VPC, every subnet has this switched on. Launch an instance with default settings and it gets a public IP automatically, alongside the default SG's permissive self-reference rule above, which is why a brand-new default-VPC instance is reachable from the internet (if you also allow the right port) with zero networking configuration. In a custom VPC, auto-assign public IPv4 is off by default on any subnet you create; you must enable it on the subnet (or attach an Elastic IP after launch) for an instance there to get a public IP at all.
ENI = think "virtual NIC." Every instance gets a primary ENI automatically at launch, fixed to the subnet you launched into (and so to that subnet's AZ). It always carries a private IP, and in most cases can't be detached while the instance is running. You can also create and hot-attach additional ENIs to a running instance for a second network presence (e.g. a separate management interface, or fast failover: detach the ENI from a failed instance and attach it to a standby, keeping the same IP). An ENI is "public" only in the sense that it sits in a public subnet and has a public IPv4/Elastic IP associated with it. The ENI object itself doesn't have a fixed public/private type.
Full detail on subnets, route tables, IGW/NAT, and security groups vs. NACLs lives in the VPC & Networking topic. This section is just the "where does my instance actually sit" mental model for Compute.
π·οΈPublic IP vs. Private IP vs. Elastic IP
Type
Persistence
Cost
Typical use case
Private IPv4
Persists for the life of the ENI (survives stop/start)
Free
Internal-only traffic: DB tier, backend services, anything reached only via a load balancer/NAT/VPN
Public IPv4 (auto-assigned)
Not persistent: released and re-assigned to a new random address on every stop/start
Billed hourly (all public IPv4 addresses, attached or not)
Quick/throwaway dev-test instances where the IP itself doesn't matter
Elastic IP
Static: yours until you release it; remap between instances/ENIs on demand
Billed hourly, same rate as any public IPv4 (see note), but still a soft-limited resource (5 per account by default), so release ones you're not using
Production internet-facing endpoints needing a stable, memorable address: a DNS record, a partner's IP allowlist, or a fast-failover target
Pros/cons in one line each: Private IP: free and stable, but never internet-reachable on its own. Public IP (auto): zero setup, but changes every restart, so nothing should hardcode it. Elastic IP: the only one that's both stable and reachable, but it's a finite, deliberately-nudged-away-from resource (a soft limit of 5 per account by default); don't hold one you're not using.
Billing gotcha (frequently tested!): since Feb 2024, AWS bills every public IPv4 address (auto-assigned or Elastic), a small hourly rate, whether it's attached to a running resource or not. Before that change, an Elastic IP was only charged while unattached (to discourage hoarding) and an auto-assigned public IP was free. The "avoid unattached EIPs" instinct from older material is still good practice, but it's no longer the only cost driver. Minimizing how many public IPv4 addresses you provision at all is now the bigger lever.
Why the instance can't see its own public/Elastic IP: that address is never actually on the instance's ENI. The ENI only ever carries the private IP. The public IP is a 1:1 NAT mapping held at the Internet Gateway, translated on the way in/out; the instance only learns it by asking instance metadata or an external service.
πENI vs. ENA vs. EFA (don't mix these up)
ENIElastic Network Interface, the virtual NIC itself: its identity (private/public IPs, MAC address, security groups). What you attach/detach.
ENAElastic Network Adapter, the high-performance driver/hardware behind an ENI on modern instance types, enabling "enhanced networking" up to 100 Gbps. Almost every current-generation instance uses this by default.
EFAElastic Fabric Adapter, an ENA variant that adds an OS-bypass hardware interface for ultra-low-latency, tightly-coupled inter-node traffic. Built for HPC/distributed ML training (MPI-style workloads), typically paired with a Cluster Placement Group.
Mental model: ENI is the network card you can see and manage (IP/MAC/security groups). ENA is what makes that card fast. EFA is a specialized version of ENA for the narrow case of HPC nodes that need to talk to each other with minimal latency, not something you'd pick for a typical web app.
Why ENA/EFA are possible at all (the Nitro System): the underlying hardware/hypervisor platform for the next generation of EC2 instances. Virtually every current-generation instance type runs on it, offloading networking, storage, and management functions to dedicated Nitro cards instead of the host CPU. Performance win: with almost nothing left for a traditional hypervisor to do on the host CPU, virtually all of that CPU/RAM is handed to your instance instead. This is what delivers near-bare-metal performance and enables enhanced networking (ENA/EFA) up to 100 Gbps. Nitro Enclaves is the exam-relevant spin-off: an isolated, hardened compute environment carved out of an instance with no persistent storage, no interactive access, and no external networking, for processing highly sensitive data (PII, cryptographic keys) with a minimized attack surface.
Steady-state, predictable usage on a specific instance family/region
Savings Plans
1 or 3 years, $/hr spend commitment
Up to ~72%
Same as Reserved but flexible across instance family/size/OS/region (Compute Savings Plans) or just size within a family (EC2 Instance Savings Plans)
Spot Instances
None (can be reclaimed)
Up to ~90%
Fault-tolerant, flexible, interruptible workloads (batch, CI, stateless web tiers)
Dedicated Hosts
On-Demand or 1/3-yr Reservation
Varies
Compliance/licensing needs a physical server mapped to you (BYOL, per-socket/core licensing)
Dedicated Instances
None or Reservation
Varies
Physical isolation from other accounts, but no control over host placement
Capacity Reservations
None (no term); reserves capacity, not discount
None by itself (pair with RI/Savings Plan for a discount)
Guarantee capacity is available in a specific AZ, e.g. for a known future launch/DR
Reserved: Standard vs. Convertible. Standard = bigger discount, can't change instance family. Convertible = smaller discount, can swap instance family/OS/tenancy during the term.
Spot gotcha: AWS can reclaim a Spot instance with a 2-minute warning (via CloudWatch event / instance metadata) when it needs the capacity back. Never use Spot for anything that can't tolerate sudden termination or can't checkpoint its work.
πSavings Plans vs. Reserved Instances: closer look
Attribute
Reserved Instances
Savings Plans
What the commitment attaches to
A specific instance family + Region/AZ + OS + tenancy (Standard); Convertible can swap family/OS/tenancy
A $/hr spend commitment; applies automatically to any matching usage, no instance-level binding
Flexibility
Standard: none. Convertible: family/OS/tenancy, but stays EC2-only
Compute Savings Plans: any instance family/size/OS/Region, plus Fargate and Lambda. EC2 Instance Savings Plans: size/OS/tenancy flexible, but locked to one family + Region
Payment options
All/Partial/No Upfront
All/Partial/No Upfront
Capacity guarantee
Zonal RI reserves actual capacity in a specific AZ; Regional RI does not
None: never guarantees capacity, purely a billing discount
Resale if plans change
Unused RIs can be sold on the AWS Reserved Instance Marketplace to recoup cost
No resale mechanism; you're committed for the term
Best for
A known, fixed instance configuration you won't change, especially if you also need a capacity guarantee (Zonal)
Steady-state spend across a mix that may shift over time (instance types, or even EC2 β Fargate/Lambda)
Exam pattern: "steady-state usage, but the instance mix might change" or "want the discount to also cover Fargate/Lambda" β Savings Plans (Compute). "Need a capacity guarantee in a specific AZ" or "might resell the commitment later" β Reserved Instances.
π°You're Billed for What You Provision, Not What You Use (frequently tested!)
For most EC2-adjacent resources, AWS charges for the size/capacity you provisioned, regardless of how much of it your workload actually consumes. Idle capacity still costs money. This is exactly why "rightsizing" is a real, tested cost-optimization lever, not just a nice-to-have.
Ex. 1Launch an m5.xlarge (4 vCPU / 16 GiB) but your app only ever uses ~10% CPU and 2 GiB RAM β you're still billed the full m5.xlarge hourly rate, not 10% of it.
Ex. 2Provision a 500 GB gp3 EBS volume but only store 50 GB of actual data on it β you're billed for all 500 GB provisioned, every month, not the 50 GB used.
Tools that help catch this: AWS Compute Optimizer (recommends better-fitting instance types/sizes from actual utilization history) and Trusted Advisor (flags idle/over-provisioned resources). Downsizing an over-provisioned instance or shrinking an oversized volume is one of the most common "how do I reduce cost" exam answers.
πΎEBS vs. Instance Store
Terminology: "Ephemeral." AWS's own name for instance store is ephemeral storage. Ephemeral just means short-lived/temporary. It's the exam's go-to vocabulary word for "this data disappears the moment the instance stops or terminates". If a question describes something as ephemeral, it's pointing you at instance store, not EBS.
EBS
Instance Store
Persistence
Survives stop/start & instance termination (if not set to delete)
Physically attached to the host, lost on stop or termination
Performance
Network-attached (slightly higher latency)
Physically local, highest possible IOPS/throughput
Flexibility
Detach/reattach to another instance, resize, snapshot
Cache, buffer, scratch/temp data, data already replicated elsewhere
Rule of thumb: if data must survive the instance, it belongs on EBS (or S3/EFS), never instance store.
EBS lives in one AZ. A volume is created inside a single Availability Zone and is automatically, redundantly replicated within that AZ by AWS for durability. That replication is what protects against a single hardware failure. It's also why a volume can only attach to an instance in the same AZ, and why moving one to a different AZ (or Region) requires taking a snapshot first (snapshots live in S3, which is Region-wide) and restoring a new volume from it there.
What it looks like from inside the instance: once attached, an EBS volume just shows up as an ordinary local disk to the OS: a drive letter like D:\ on Windows, or a device like /dev/xvdf on Linux. The instance has no idea it's talking to network-attached storage under the hood.
πEBS Volume Types
gp3General purpose SSD: default choice; IOPS/throughput provisioned independently of size. Use for: boot volumes, dev/test, most everyday app and small-to-medium database workloads
gp2Older general purpose SSD: IOPS scales with volume size (3 IOPS/GB). Use for: legacy volumes not yet migrated to gp3; rarely the right pick for anything new
io1 / io2Provisioned IOPS SSD: highest performance, mission-critical low-latency DBs; io2 Block Express for the largest scale. Use for: large relational DBs (Oracle, SQL Server, SAP HANA) with sustained heavy IOPS and sub-millisecond latency needs
st1Throughput Optimized HDD: big sequential workloads; cannot be a boot volume. Use for: big data/Hadoop, data warehouses, log processing; anything reading large sequential chunks fast, not small random reads
sc1Cold HDD: lowest cost, infrequently accessed data; cannot be a boot volume. Use for: archive-style data, infrequently-accessed backups (cost matters more than speed)
Exam pattern: "cheapest option for infrequent access" β sc1. "Boot volume" β must be SSD (gp2/gp3/io1/io2), never st1/sc1. "Need consistent sub-millisecond latency for a DB" β io1/io2.
EBS Multi-Attach: io1/io2 only. Normally an EBS volume attaches to exactly one instance at a time. io1/io2 can opt into Multi-Attach: the same volume attached to up to 16 instances at once, but only instances in the same AZ. It doesn't give you a shared filesystem for free. Without a cluster-aware filesystem (or your app coordinating writes itself), multiple instances writing to the same blocks will corrupt data. Used for cluster-aware apps designed for shared block storage.
πΈEBS Snapshots
A snapshot is a point-in-time, incremental backup of an EBS volume, stored durably in S3 behind the scenes (you don't manage the bucket; it's not visible in your S3 console). Incremental means the first snapshot copies every used block, and every snapshot after that only stores the blocks that changed since the last one, but each snapshot still restores to a complete, standalone volume; deleting an older snapshot in the chain doesn't break the newer ones.
Unlike the volume itself, a snapshot is Region-wide, not AZ-locked. This is exactly what makes it the mechanism for moving a volume's data across AZs or Regions (see the AZ-locked note above): snapshot the volume, then restore a new volume from that snapshot in whichever AZ (same Region) you need, or copy the snapshot to another Region first if you need to cross Regions.
Backup / DRScheduled snapshots (e.g. via Amazon Data Lifecycle Manager or AWS Backup) protect against accidental deletion, corruption, or a bad deploy; restore a new volume from any snapshot in the chain
Move across AZ/RegionSnapshot in the source AZ β restore a volume from it in the target AZ (same Region), or copy the snapshot to another Region first for cross-Region moves
Building an AMIAn AMI's block device mapping is literally a set of references to EBS snapshots. Creating a custom AMI creates the underlying snapshot(s) for you
Resize/retype a volumeRestore a snapshot into a larger volume, or onto a different volume type (e.g. gp2 β gp3), instead of resizing the original in place
Frequently tested: a snapshot can be taken while the volume is in use, but for a fully consistent backup (especially a root/boot volume) best practice is to stop the instance first, or at minimum flush filesystem writes. An "in flight" write not yet flushed to disk at snapshot time can leave the restored volume in a slightly inconsistent state. A snapshot of an encrypted volume is automatically encrypted too, and copying a snapshot lets you encrypt an originally-unencrypted one along the way.
πEBS Encryption
An encrypted EBS volume protects data at rest on the volume, data in transit between the instance and the volume, and every snapshot made from it (plus any new volume restored from those snapshots), all using AES-256 under an AWS KMS key. It's fully transparent: the OS and your application see a normal disk, with encrypt/decrypt happening on the underlying host hardware, at negligible performance cost.
You can't flip encryption on an existing volume directly. The path is: snapshot the unencrypted volume β copy that snapshot with encryption enabled (this is the one step where you can turn it on) β create a new volume from the encrypted copy β attach the new volume in place of the old one. Same trick works for changing which KMS key protects a volume.
Compliance / regulated dataPCI-DSS, HIPAA, GDPR-style requirements for encryption at rest, e.g. a volume holding payment records or health data
Org-wide "encrypt by default"Enable default encryption per Region so every new volume/snapshot is encrypted automatically, without anyone remembering to tick a box
Safer snapshot sharingAn encrypted snapshot can never be made public, a built-in guardrail against accidentally exposing sensitive data when sharing/copying snapshots across accounts
Key choice: the default AWS-managed key (aws/ebs) works with zero setup; a Customer Managed Key (CMK) via KMS costs a little more but gives you rotation control, a separate key per project/environment, and the ability to revoke access (or delete the key) to instantly make the data unreadable.
π¦AMIs (Amazon Machine Images)
Terminology: AMI. An AMI is a template for launching an instance. Think of it like a "frozen disk image" or a phone's factory restore image, not a running server. Example: you configure one EC2 instance by hand (install Nginx, deploy your app code, apply OS patches), then create an AMI from it. That AMI now captures the whole machine at that moment. Launch 20 new instances from it and every one boots up already running Nginx with your app pre-installed, with no setup script needed.
An AMI is a template: OS + installed software + configuration + permissions + a block device mapping (which EBS snapshots become which volumes).
Golden AMI pattern: bake app/config/patches into a custom AMI ahead of time so new instances launch pre-configured, much faster than running a bootstrap script on every launch.
AMIs are region-specific. Copy an AMI to another region to launch instances there.
Sources: AWS-provided, AWS Marketplace, community, or your own (created from an existing instance/snapshot).
πPlacement Groups
A logical grouping that controls how AWS places instances on underlying hardware. Pick a strategy to optimize for either performance (low latency, high throughput) or fault isolation.
Attribute
Cluster
Spread
Partition
Goal
Lowest latency, highest throughput
Maximize instance isolation
Isolate groups of instances
AZ scope
Single AZ
Multi-AZ allowed
Multi-AZ allowed (within a Region)
Hardware isolation
Low (packed together)
Highest (one instance per rack)
Medium (isolated per-partition)
Instance limits
None
7 instances per AZ
Up to 7 partitions per AZ; scales to hundreds of instances overall
Risk of correlated failure
Higher
Lowest
Lower
Typical use case
HPC, big data, tightly-coupled/high-speed apps
Critical workloads that must not fail together
Large distributed systems (HDFS, Cassandra, Kafka); partition info exposed to the instance for rack-aware replication
Cluster trades fault isolation for speed; Spread and Partition trade some speed back for isolation, at increasing scale.
Can't move instances inExisting instances can't easily be added to a group after the fact; launch new instances into it; relaunching is usually simpler than migrating
Enhanced networkingCluster benefits most when every instance type in the group supports it; mixing types limits the payoff
CostPlacement groups themselves are free; no extra charge to use one
πInstance Lifecycle: Stop vs. Terminate
States: pending β running β either stopping/stopped (Stop) or shutting-down/terminated (Terminate). Reboot isn't a state change at all: it's just an OS restart; instance ID, IPs, and volumes are untouched.
Stop only exists for EBS-backed instances (frequently tested!): an instance-store-backed instance has no EBS root volume to power down onto and preserve. It can only be terminated, never stopped. If a question offers "Stop" as an option for an instance-store-backed instance, that option is wrong.
Stop
Terminate
Instance ID
Kept (same instance, can be started again)
Gone permanently
Private IP
Kept (survives stop/start)
Released (the primary ENI is deleted with the instance)
Elastic IP
Stays associated (still billed per the note above)
Released back to your account
Root EBS volume
Kept, billed as ordinary EBS storage while stopped
Deleted by default, unless DeleteOnTermination was turned off
Instance store data
Lost the moment power stops
Lost (same reason)
RAM contents
Lost, unless Hibernate is enabled, which saves RAM to the root EBS volume first (see Hibernate note below)
Lost
Compute billing
Stops (no per-hour instance charge while stopped)
Stops
A stopped instance is the state both Hibernate and Auto Recovery build on top of. See their own notes on this page. Both Stop and Terminate can be gated by a guardrail flag. See Termination & Stop Protection below.
Every instance runs two independent automated checks every minute. Telling them apart is the whole exam question: they imply completely different fixes.
Check
What it's really testing
Typical cause
Fix
System Status Check
The underlying AWS hardware/hypervisor/network the instance runs on
Host power loss, hardware failure, network connectivity issue at the host level
Stop/Start the instance. This moves it onto different underlying hardware (a reboot does not help, since it doesn't change hardware)
Reboot the instance; the hardware is fine, just the OS needs to restart
EC2 Auto Recovery: a CloudWatch alarm (on the StatusCheckFailed_System metric) that automatically stops and starts the instance for you the moment a System Status Check fails, no human needed to notice and intervene. The recovered instance keeps the same instance ID, private IP, Elastic IP, and EBS volumes; it just lands on healthy hardware. This is EC2-level reliability, distinct from Auto Scaling (which replaces an unhealthy instance with a brand-new one rather than recovering the same one).
Instance Retirement: AWS-initiated, not something you trigger. When the underlying hardware an instance sits on is degrading or scheduled for decommission, AWS schedules a retirement date and notifies you in advance (Personal Health Dashboard/email), giving you a window to Stop/Start the instance yourself (moving it to healthy hardware, same as Auto Recovery does) before AWS does it for you at the retirement date.
πTermination & Stop Protection
Termination ProtectionA per-instance flag (DisableApiTermination) that blocks a Terminate request from the console/CLI/API until it's switched off, the classic guardrail against accidentally deleting a production instance (and its instance-store data, and any EBS volumes set to delete-on-termination)
Stop ProtectionThe same idea for Stop (DisableApiStop), useful for an instance where even a stop/start (e.g. losing its public IP, or interrupting a long-running process) would be disruptive
Both are just flags you toggle on the instance: no extra cost, no separate service. Neither is a substitute for IAM permissions (a user with ec2:TerminateInstances can still disable the flag first, then terminate). Think of it as a safety catch against accidental clicks, not a security control.
Every instance can query its own metadata (instance ID, AMI ID, IAM Role credentials, user data, etc.) from inside the instance at:
http://169.254.169.254/latest/meta-data/
IMDSv2 vs. IMDSv1: IMDSv1 answers plain GET requests to that address, a classic SSRF target (a vulnerable app can be tricked into fetching it and leaking the instance's Role credentials). IMDSv2 requires a session token (via a PUT request first) before it will answer, closing that SSRF path. AWS/the exam now treats IMDSv2 as the best-practice default.
User DataScript run once, automatically, on first boot, used to bootstrap/configure a new instance (install packages, pull config, join a fleet)
Elastic IPStatic public IPv4 you own and can remap between instances; billed hourly whether attached or not (since Feb 2024); release ones you're not using, not just unattached ones
ENIElastic Network Interface, a virtual NIC (private IPs, security groups, MAC); additional ENIs can be hot-attached/detached between instances for failover. Full breakdown (incl. ENA/EFA) in "Where EC2 Lives" above
HibernateSaves in-memory RAM state to the root EBS volume on stop, restores it on start; faster warm boot than a cold start's OS+app re-init. Constraints: root volume must be encrypted, RAM capped at ~150 GiB, and it's not supported on instance-store-backed instances (nothing to durably save the RAM contents to)
πConnecting to an Instance (frequently tested!)
Four ways to get a shell on an instance. The exam usually asks "what's the most secure / only way to connect here," so the differences matter more than the mechanics.
Method
Requires
Pros
Cons
Best for
EC2 Instance Connect
SG allows inbound 22 (or 3389), a supported AMI (Amazon Linux 2/2023 have it preinstalled), IAM permission to push the key
No key pair to manage or store; pushes a short-lived, one-time SSH key via the API; works straight from the browser console
Still needs an open inbound port and a real network path (public IP, or same-VPC/VPN reachability); doesn't help a fully private instance with no such path
Quick one-off browser access to a reachable (public or in-VPC) instance without setting up a local SSH client
Session Manager (Systems Manager)
SSM Agent running (preinstalled on most current AMIs), an instance IAM role with SSM permissions, and outbound HTTPS reachability to the SSM endpoints (public internet, NAT, or SSM VPC endpoints)
Zero inbound ports open at all: no SG rule for 22/3389, no bastion host, no key pair; every session is logged (who, when, commands) via CloudTrail/S3/CloudWatch Logs for audit
Needs the Agent + IAM role provisioned ahead of time; still needs an outbound path to the SSM service (a fully air-gapped private subnet needs SSM VPC endpoints added)
The AWS-recommended default, especially for private-subnet instances; removes the need for a bastion host entirely
SSH client (traditional)
A key pair, an SG allowing inbound 22 from your source, and network reachability (public IP directly, or a bastion host / VPN / Direct Connect into a private subnet)
Familiar tooling, full terminal control, works the same regardless of AWS-specific agents or console access
You're responsible for storing/rotating the private key yourself; some inbound port must be opened somewhere; a private instance needs an extra bastion/VPN hop
Existing key-based workflows, automation/scripts, or environments not using Systems Manager
EC2 Serial Console
Enabled at the account level, plus an OS-level console user/password configured on the instance in advance
Works even when the instance has no working network path at all: locked out by a bad SG/NACL/route table change, broken sshd, corrupted network config
Text-only, no file transfer, no IAM-role convenience, and useless in the emergency it's meant for if you didn't configure OS console access beforehand
True last-resort troubleshooting of an instance that's unreachable over the network by every other method here
Exam callout: "most secure way to connect to a private instance, no bastion host" β Session Manager, every time. It's the only option here needing no inbound rule whatsoever. Serial Console is the odd one out: the only method that still works when the instance's own network stack is the thing that's broken, which is exactly why it can't rely on SSH, SSM, or Instance Connect (all of which need a working network).
π§ Quick Memory Hooks
Compute-optimized = CPU-heavy Β· RAM-optimized = memory Β· I = IOPS-heavy local disk Spot = cheapest but can vanish with 2 min warning. Reserved/Savings Plans = commit for a discount. On-Demand = pay for flexibility. If it must survive termination, it's not on Instance Store. Session Manager = no open ports, no keys, full audit trail. Serial Console = the only way in when the network itself is the problem.
πͺ£ S3 & Storage: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
S3 is object storage, not a disk: think "a bucket of files with a key," not a filesystem with folders. Almost every exam question is really asking which storage class or which access control layer.
Same "11 nines" (99.999999999%) durability across every class. What changes between classes is availability and retrieval cost/speed, not durability.
ποΈStorage Classes (frequently tested!)
Class
Access pattern
Retrieval time
Typical use case
Standard
Frequent access
Milliseconds
Actively-used data (website assets, app content)
Intelligent-Tiering
Unknown/changing access pattern
Milliseconds
Auto-moves objects between tiers by observed access, no retrieval fees for tier changes
Standard-IA
Infrequent, but need fast access when needed
Milliseconds
Backups, disaster recovery files, older but still-needed data
One Zone-IA
Infrequent, re-creatable data
Milliseconds
Same as Standard-IA but single-AZ; cheaper, less resilient (fine for easily-regenerated data)
Glacier Instant Retrieval
Archive, rarely accessed, but needs instant access
Milliseconds
Quarterly reports, medical images; archived but occasionally opened immediately
Glacier Flexible Retrieval
Archive
Minutes to hours
Backups you don't need in a hurry
Glacier Deep Archive
Long-term archive, almost never accessed
Standard: ~12h Β· Bulk: up to 48h
Compliance/regulatory retention (7-10 year records), cheapest storage class
Exam pattern: "cheapest, access once or twice a year, can wait hours" β Glacier Deep Archive. "Unpredictable access pattern, don't want to manage tiering manually" β Intelligent-Tiering.
β»οΈLifecycle Rules
A lifecycle rule automatically transitions objects to a cheaper class, and/or expires (deletes) them, after a set number of days. No manual intervention once configured.
Example rule: every threshold above is fully configurable per rule, and you can skip tiers or stop at any point.
Applies to a whole bucket, or scoped by prefix/tag, e.g. only objects under logs/.
πVersioning
Once enabled (bucket-level), every overwrite or delete keeps the old version instead of destroying it. A "delete" just adds a delete marker on top.
Protects against accidental overwrite/deletion. Restore by removing the delete marker or fetching an older version ID.
Can't be fully turned off once enabled, only suspended. MFA Delete adds a second-factor requirement to permanently delete a version.
Lifecycle rules can target noncurrent versions separately, e.g. auto-expire old versions after 90 days to control storage cost growth.
πEncryption
SSE-S3Amazon-managed keys: encryption at rest with zero setup; the default for every new object/bucket
SSE-KMSKMS-managed keys; adds an audit trail (every use logged to CloudTrail) and granular per-key IAM control; costs a bit more (KMS API calls). Comes in two flavors: an AWS managed key (created/rotated automatically, zero setup, but you can't set its key policy) or a customer managed CMK (you create it, control its key policy/rotation, can disable or schedule deletion). "Full audit + keys managed by us" in a question points at a customer managed CMK specifically, not just "SSE-KMS" in general
SSE-CYou supply your own encryption key with every request; AWS never stores it
Client-sideYou encrypt before upload; S3 only ever sees ciphertext
In transit: enforce HTTPS-only access via a bucket policy condition on aws:SecureTransport.
πAccess Control
Bucket PolicyA resource-based JSON policy on the bucket itself, the standard way to grant/restrict access, including cross-account
IAM PolicyAttached to a User/Role instead of the bucket; same underlying permission model, different attachment point
ACLLegacy, object/bucket-level access list. AWS now recommends disabling ACLs entirely and using policies instead
Block Public Access is a bucket (and account-wide) setting that overrides any policy/ACL trying to make something public. It's on by default for new buckets, and the exam expects you to know it must be deliberately turned off before a bucket can be made public, no matter what the policy says.
CORS is a browser security rule, not an S3-specific concept: a webpage's JavaScript is blocked by the browser from calling a different origin (domain, scheme, or port) than the page itself was loaded from, unless that other origin's response explicitly says it's allowed.
If a webpage hosted anywhere makes a JavaScript fetch/XHR request straight to an S3 bucket (a different origin from the page), the browser blocks it by default. Fix: add a CORS configuration on the bucket itself: a small JSON/XML rule set naming the allowed origins, HTTP methods (GET/PUT/POST/etc.), and headers. This is a setting on the bucket's CORS configuration, separate from a bucket policy or IAM policy.
Exam pattern (the classic trap): a question describes a script/webpage making authenticated requests to S3 and failing, with a bucket policy and IAM permissions already correctly in place. The fix is still enabling CORS. A bucket policy governs authorization (is this caller allowed to do this), while CORS governs a completely separate browser-enforced check that happens regardless of whether the caller is authorized. Fixing one doesn't fix the other; versioning and encryption settings are unrelated distractors here too.
βοΈS3 vs. EFS vs. EBS
S3
EFS
EBS
What it is
Object storage (a bucket, not a drive)
Managed NFS file system
A virtual block-storage disk
Attaches to
Nothing; accessed over HTTP(S) API from anywhere
Many EC2 instances at once, across multiple AZs
One instance at a time (Multi-Attach io1/io2 excepted, same AZ only)
Scope
Region-wide
Region-wide (multi-AZ)
Single AZ
Typical use
Static assets, backups, data lake, hosting
Shared content/config across a fleet of instances
Boot volumes, databases; anything needing a real filesystem for one instance
Exam pattern: "multiple EC2 instances behind a load balancer must all see the same files/documents". This is a shared-filesystem need, so EBS (one instance at a time) is disqualified regardless of how the data got out of sync in the first place; the fix is EFS, not "copy the same files to every instance's own EBS volume" (which just re-breaks the moment anything changes).
πGetting Data Into AWS: Transfer & Migration Options (frequently tested!)
Which tool is right depends entirely on data volume and available bandwidth. The exam tests whether you can match the scenario's numbers to the right option, not just recognize the service names.
Tool
What it does
Best for
S3 Transfer Acceleration
Routes an upload through CloudFront's global edge locations over AWS's private backbone instead of the public internet path to the bucket's Region
Frequent, ongoing uploads from geographically distributed users/sites where the network path itself is the bottleneck, not a one-off bulk migration
AWS DataSync
An online, automated, ongoing transfer/sync agent between on-premises storage (NFS/SMB) and S3/EFS/FSx; encrypts in transit, validates data, can run on a schedule
Repeated or ongoing transfers over an existing network link (works well alongside Direct Connect); not for a single one-time bulk migration with very limited bandwidth
AWS Snowball Edge
AWS ships you a physical device; you load it with data on-site, ship it back, AWS ingests it directly into S3; bypasses the network entirely
Large one-time (or infrequent) bulk transfers (tens of TB to PB) where available bandwidth would otherwise take days/weeks
AWS Storage Gateway (File Gateway)
An on-premises virtual appliance presenting an NFS/SMB share that's actually backed by S3: files written locally land in S3, with the most-recently-used data cached locally for low-latency access
Keeping an existing on-prem file-based workflow while transparently using S3 as the actual backing store, including S3 Lifecycle rules on the data it writes
Exam pattern: "as quickly as possible, minimal operational complexity, one-off, very large volume, limited/expensive bandwidth" β Snowball Edge. This beats Transfer Acceleration whenever the numbers imply days/weeks over the wire (a classic giveaway: a large TB/PB figure paired with a modest daily bandwidth figure that doesn't divide down to a reasonable transfer window). "Ongoing/repeated programmatic sync between on-prem and AWS storage" β DataSync. "Keep an existing file-share workflow, back it with S3, apply lifecycle rules" β Storage Gateway File Gateway.
Amazon Athena: a serverless query service that runs standard SQL directly against data already sitting in S3 (CSV, JSON, Parquet, etc.): no database to provision, no ETL pipeline, pay only per query (per data scanned). The default answer whenever a question wants ad-hoc/on-demand analysis of S3-resident logs or data with minimal setup, over standing up Redshift (a provisioned data warehouse) or a Glue+EMR pipeline (built for heavier, repeated ETL processing).
π§ Quick Memory Hooks
Durability never changes across classes: only availability and retrieval speed/cost do. S3 = a bucket (HTTP API) Β· EFS = a shared network drive for many instances Β· EBS = a private disk for one instance. Block Public Access beats everything else: a wide-open bucket policy still won't be public if this is on.
ποΈ Databases: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Pick the database by the shape of the data and the access pattern, not familiarity: relational + complex queries β RDS/Aurora; key-value at massive scale β DynamoDB; sub-millisecond lookups β ElastiCache.
π’οΈRDS Basics
Amazon RDS = a managed relational database: AWS handles patching, backups, and failover setup. Engines: MySQL, PostgreSQL, MariaDB, Oracle, SQL Server, and Aurora.
"Managed" means AWS handles the undifferentiated heavy lifting (OS patching, engine patching, backups); you still choose instance size, storage, and when maintenance windows apply.
Automated backups + manual snapshots both available; point-in-time recovery restores to any second within the backup retention window.
RDS Proxy (frequently tested!) A fully-managed connection pooler that sits between the application and the database. Solves a specific problem: a highly-concurrent, short-lived-connection caller (classically Lambda, where every concurrent invocation can open its own DB connection) can exhaust a database's max-connections limit and slow it down with constant connect/disconnect overhead. RDS Proxy pools and reuses a small number of real DB connections behind many app-side connections. It also speeds up failover: the proxy holds the connection and transparently re-points it to the new primary, so the application doesn't need to detect and reconnect itself. Exam pattern: "Lambda functions are overwhelming the database with connections" β RDS Proxy, not simply "increase the DB instance size."
πMulti-AZ vs. Read Replicas (frequently tested!)
These solve two different problems and are often combined, not alternatives to each other.
Multi-AZ = disaster insurance (same Region, automatic failover). Read Replicas = a scaling lever (same/cross-Region, manual promotion).
Automatic (AWS-managed, DNS endpoint stays the same)
Manual promotion to a standalone instance
Location
A second AZ in the same Region
Same Region or cross-Region, plus a narrow Aurora-specific case: an RDS MySQL instance can replicate into an Aurora MySQL cluster as a migration path (not general cross-engine replication, e.g. no MySQLβPostgreSQL)
πAurora
AWS's own MySQL/PostgreSQL-compatible engine, same wire protocol/drivers, but a re-architected storage layer that auto-replicates data 6 ways across 3 AZs and claims up to 5x MySQL / 3x PostgreSQL throughput.
Aurora ReplicasUp to 15, low replication lag (shared storage layer, not a full copy); can also fail over automatically, unlike a standard RDS read replica
Global DatabaseOne primary Region + up to 5 read-only secondary Regions, typically <1s lag; for globally-distributed reads or DR
Aurora ServerlessAuto-scales capacity up/down (even to zero) with demand, for spiky or unpredictable workloads instead of sizing an instance by hand
Exam pattern: when a question wants "the best relational performance/availability on AWS," Aurora is usually the intended answer over plain RDS MySQL/PostgreSQL.
β‘DynamoDB (frequently tested!)
Fully-managed key-value / NoSQL store, single-digit millisecond latency at any scale, no servers to manage at all (serverless).
Partition Key (hash key)Determines which physical partition an item lives on; the core design decision for even data distribution and query performance. A high-cardinality key (lots of distinct values) spreads load evenly; a low-cardinality one creates a "hot partition"
Sort Key (range key), optionalPaired with the partition key to form a composite primary key. Items sharing a partition key are stored together, ordered by sort key, so a query can efficiently pull a range (e.g. all of one customer's orders, sorted by date)
On-Demand capacityPay per request, auto-scales instantly; for unpredictable/spiky traffic
Provisioned capacitySet Read/Write Capacity Units (RCU/WCU) ahead of time; cheaper at steady, predictable load. Pair with DynamoDB Auto Scaling (a target-utilization setting, e.g. 70%, conceptually identical to ASG Target Tracking) so capacity adjusts within a min/max band automatically instead of a human resizing it. This is the "most cost-effective" answer far more often than switching to On-Demand or adding DAX, when the workload is steady rather than spiky. A ThrottledRequests CloudWatch alarm signals capacity is set too low for either mode.
DAXDynamoDB Accelerator, an in-memory read-through/write-through cache in front of DynamoDB, microsecond latency for read-heavy workloads. Solves a read-latency problem, not a cost/throttling problem; a common exam distractor when the real fix is just right-sizing capacity (above)
DynamoDB StreamsAn ordered, 24-hour log of every item-level change (insert/update/delete) on a table, the trigger source for reacting to data changes (e.g. invoking a Lambda function per change), not just a Lambda poll-source detail
TTL (Time To Live)Auto-deletes an item once its designated timestamp attribute passes: free (no WCU consumed) background deletion, for data with a natural expiry (session data, temporary tokens) instead of a manual cleanup job. A completely different "TTL" from CloudFront's cache TTL (same term, unrelated concept)
Global TablesMulti-Region, multi-active replication: write to any Region, reads stay fast and local everywhere
Global Secondary Index (GSI)
Local Secondary Index (LSI)
Keys
A different partition key (and optionally a different sort key) from the base table
Same partition key as the base table, a different sort key
When it can be created
Any time (added to an existing table)
Only at table creation time; cannot be added or removed later
Capacity
Its own separate RCU/WCU, provisioned independently of the base table
Shares the base table's capacity
Consistency
Eventually consistent reads only
Eventually or strongly consistent reads (your choice)
Limit
Up to 20 per table
Up to 5 per table
Exam pattern: a question needing to query by a completely different attribute than the base table's partition key (e.g. table keyed by Song ID, but the app needs to query "all songs by this Artist") β GSI, essentially always. An LSI can't change the partition key at all. Reach for LSI only when the scenario explicitly needs strong consistency on an alternate sort order for the same partition key, and the index was planned before the table was created.
π§ElastiCache
Redis
Memcached
Data structures
Rich (lists, sets, sorted sets, hashes, pub/sub)
Simple key-value only
Persistence
Optional snapshotting/AOF (can survive a restart)
None: pure in-memory, gone on restart
HA
Multi-AZ with automatic failover, read replicas
None built-in; just multiple independent nodes
Scaling
Cluster mode shards data across nodes
Simple horizontal scaling via multiple nodes
Exam pattern: "need persistence, replication, or pub/sub" β Redis. "Simplest possible object cache, no HA needed" β Memcached.
4Analytics over huge historical datasets, complex OLAP queries β Redshift (data warehouse, out of scope of this page but good to recognize by name)
π§ Quick Memory Hooks
Multi-AZ = insurance policy (sync, unreadable standby). Read Replica = extra hands (async, readable, scales reads). Aurora = RDS's faster, AWS-native cousin. DynamoDB = no servers, no limits, simple lookups. ElastiCache = everything in RAM, blink-fast. GSI = a new door with a new key. LSI = the same door, a different sort order, and only if you said so on move-in day. RDS Proxy = a bouncer for the database's front door, so Lambda's crowd of short visits doesn't overwhelm it.
βοΈ ELB & Auto Scaling: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
A Load Balancer spreads traffic across instances; Auto Scaling changes how many instances exist. Together: elastic capacity that survives both traffic spikes and instance failures.
π‘οΈHigh Availability vs. Fault Tolerance (frequently tested!)
High Availability (HA): the system keeps running with minimal downtime when a component fails: it recovers automatically, but recovery can involve a brief, real interruption while failover happens. The goal is "downtime measured in seconds/minutes, not hours," not necessarily zero.
Fault Tolerance (FT): the system keeps running with zero service interruption when a component fails: end users never notice. This needs redundant capacity that's already active and already absorbing load before the failure, not a standby that has to be switched to.
High Availability
Fault Tolerance
Downtime on failure
Brief (a failover event), small but real
None: no interruption at all
How it's achieved
A standby/replica takes over after detecting failure
Redundant components are already active in parallel, each individually sized to absorb the loss of another
Cost
Lower: standby capacity, not necessarily fully active
Higher: every "extra" unit of redundancy is running (and billed) all the time, not just on standby
AWS example
RDS Multi-AZ: synchronous standby in another AZ, but there's still a real (~1-2 min) failover event on the primary's failure
An ASG spread across 3 AZs behind an ALB, sized so any one AZ can be lost and the remaining two still handle full load with zero dropped requests; no failover event, nothing to "switch to"
Relationship: Fault Tolerance is a stronger, more expensive goal that sits on top of High Availability, not a separate, unrelated concept. Every fault-tolerant system is highly available (it has no meaningful downtime at all), but not every highly-available system is fault tolerant (Multi-AZ RDS is HA, but that failover window means it isn't strictly FT). S3 and DynamoDB are fault tolerant by design: their multi-AZ redundancy is invisible to you, with no failover event to reason about.
Exam pattern: "recovers automatically, brief/acceptable interruption is fine" β High Availability (Multi-AZ, Auto Recovery). "Must keep serving traffic with zero interruption even if a whole AZ goes down" β Fault Tolerance (over-provisioned ASG/ELB across β₯3 AZs, or a managed service that's fault tolerant by default like S3/DynamoDB).
πLoad Balancer Types (frequently tested!)
Elastic Load Balancing (ELB) is the umbrella AWS service, a fully managed load balancer that AWS scales, patches, and runs across multiple AZs for you (no servers of your own to manage). It comes in four types, split across the two networking layers the exam cares about: Application (Layer 7) and Network (Layer 4), plus Gateway (Layer 3) and a legacy fourth type.
Application (ALB)
Network (NLB)
Gateway (GWLB)
Classic (CLB), legacy
Layer
7 (HTTP/HTTPS)
4 (TCP/UDP)
3 (IP), transparent network gateway
Both 4 and 7, but shallowly; predates ALB/NLB's split
Routing
Path/host-based rules, e.g. /api/* β one target group, /images/* β another
Just forwards connections, no content awareness
Passes all traffic through to a fleet of virtual appliances
One flat routing model: no path/host rules, no per-listener target groups
Performance
Very fast, but not the fastest
Millions of requests/sec, ultra-low latency, static IP support
N/A (not about speed)
Lower throughput than ALB/NLB, no modern performance features
Typical use
Web apps, microservices, containers
Extreme performance needs, TCP-only protocols, static IP requirement
Inserting a fleet of firewalls/IDS/IPS appliances transparently in front of traffic
Only ever the right answer for an old EC2-Classic-era app that predates VPC. AWS actively steers new designs toward ALB/NLB instead
Exam pattern: "route by URL path or hostname" β ALB. "Extreme throughput, need a static IP" β NLB. "Deploy third-party security appliances inline" β GWLB. If "Classic Load Balancer" appears as an answer option at all, it's almost always the wrong one. The exam uses it to test whether you know it's the deprecated previous-generation type, not a live recommendation.
GWLB + GENEVE (frequently tested!): GWLB listens for every IP packet on every port (not specific listeners like ALB/NLB) and hands each one to a fleet of virtual appliances using the GENEVE protocol on port 6081: the appliance inspects/filters the packet and sends it back, transparently, before it continues to its real destination. Knowing "GENEVE, port 6081" by name is a recurring exact-recall question.
Target types: ALB can register EC2 instances, IP addresses (including on-prem/peered-VPC targets), and, uniquely, Lambda functions as targets (the request becomes the Lambda's event payload; no traditional health check needed since Lambda has no "instance" to check). NLB registers instances, IP addresses, or an Application Load Balancer (see below). GWLB registers virtual-appliance instances/IPs, not application targets.
Cross-Zone Load Balancing (frequently tested!): controls whether each load balancer node (one is provisioned per enabled AZ) spreads its share of traffic across every registered target in every AZ, or only the targets sitting in its own AZ. ALB has it on by default, free. NLB and GWLB have it off by default, and turning it on incurs cross-AZ data transfer charges for any traffic that has to cross an AZ boundary to reach a target. A classic "why is my traffic unevenly spread across AZs" gotcha.
Worked example: one load balancer spanning 2 AZs: AZ-A has 1 target, AZ-B has 3 targets. Cross-zone off: the LB node in AZ-A can only forward to AZ-A's 1 target, so that single instance absorbs 50% of all traffic (as much as the other 3 combined): an even split per node, not per target, so mismatched AZ instance counts create hot spots. Cross-zone on: every LB node spreads its traffic across all 4 targets regardless of AZ, so each instance gets roughly a 25% share instead.
Every ELB (ALB/NLB/GWLB) must be enabled in at least 2 AZs, a hard requirement, not just an HA best practice. Within each enabled AZ, an ELB can only use one subnet; size that subnet /27 or larger, with at least 8 free IP addresses, since AWS reserves capacity in it for the load balancer's own nodes to scale into; a subnet that's nearly full of other resources can silently block the ELB from scaling.
ποΈALB and NLB Deployments (frequently tested!)
Internet-facing vs. Internal (both ALB and NLB): chosen when you create the load balancer, not something you can quietly bolt on later.
Internet-facing
Internal
DNS resolves to
Public IP(s)
Private IP(s) only
Subnets
Public subnets (nodes need a route to an Internet Gateway)
Private subnets (no path to/from the public internet at all)
Typical use
A public web app's front door
Internal microservice-to-microservice traffic, or a private API only other VPC resources call
Exam pattern: "expose this tier only to other services inside the VPC, never to the internet" β Internal load balancer, not Internet-facing. A common wrong-answer trap is reaching for extra security groups instead of the simpler internal-scheme fix.
Listeners: a load balancer can run multiple listeners at once (e.g. one on port 80, one on port 443). Each listener checks for connection requests on its own protocol/port and decides what to do with them independently of the others.
ALB listener rules (content-aware): each listener evaluates a priority-ordered list of rules, each made of conditions (path, hostname, HTTP header, query string, source IP) and an action: forward to a target group, redirect (e.g. every HTTP request on port 80 β the equivalent HTTPS URL on port 443), or fixed-response (return a static status code/body with no backend call at all, e.g. a 503 maintenance page). A default catch-all rule runs when nothing else matches. This is what makes ALB "Layer 7" in practice: the routing decision can depend on the actual request content, not just where it came from.
NLB listeners (connection-aware, not content-aware): a listener typically just forwards every connection straight to one target group: no path/host rules, no redirects, no fixed responses, because NLB operates below the layer where that content is visible. Target types: instance, IP, or even an ALB (put an NLB in front of an ALB to get NLB's static IP / extreme-throughput properties while keeping ALB's content-based routing behind it).
Exam pattern: "need path-based or host-based routing" or "need an HTTPβHTTPS redirect rule" β ALB listener rules. NLB cannot do either.
πSecure Listeners for ELB (frequently tested!)
A secure listener is an HTTPS (ALB) or TLS (NLB) listener: the load balancer terminates the TLS connection itself using an X.509 certificate, decrypts the request, then talks to the targets (in plain HTTP, or re-encrypted via HTTPS/TLS to the target group for end-to-end encryption). Terminating TLS at the load balancer offloads the CPU cost of encryption from every backend instance onto the load balancer instead.
AWS Certificate Manager (ACM): the standard way to get the certificate a secure listener needs: request a public certificate for your domain (validated via DNS or email) and attach it directly to the listener. ACM-issued certificates used with an integrated service like ELB are free and auto-renew, removing the classic "certificate silently expired" outage. A third-party certificate can also be imported into ACM (or IAM, for Regions without ACM) if you already own one.
SNI = Server Name Indication (frequently tested!): lets one HTTPS/TLS listener serve multiple certificates for multiple domains on a single load balancer. The client states which hostname it's connecting to during the TLS handshake itself, before any data is decrypted, so the listener can pick and serve the matching certificate, so no more "one load balancer per domain/certificate" just to keep certs straight. Both ALB and NLB (TLS listener) support SNI.
Security policies: a predefined set of TLS protocol versions and cipher suites the listener will accept; you choose one per HTTPS/TLS listener. A stricter policy (disabling older TLS 1.0/1.1, weak ciphers) is the answer whenever a question mentions a compliance requirement (e.g. PCI-DSS) or "must not allow outdated/weak encryption"; a looser policy trades that off for compatibility with older clients that can't negotiate modern TLS.
ALB also supports mutual TLS (mTLS): the listener can require and verify a client certificate too, not just present its own, for scenarios needing client-identity verification at the load balancer (e.g. B2B/IoT APIs) rather than purely server-side TLS.
Exam pattern: "offload TLS/SSL processing from the application instances" β terminate at the load balancer with a secure listener. "Host several domains, each with its own certificate, behind one load balancer" β SNI. "Enforce modern TLS only for compliance" β pick a stricter security policy, not a certificate change.
π―Target Groups & Health Checks
A Target Group is the set of destinations a load balancer routes to: EC2 instances, IP addresses, or even Lambda functions.
The load balancer continuously health-checks every registered target (a ping to a configurable path/port) and stops routing to any that fail; traffic only ever goes to healthy targets.
One load balancer can route to multiple target groups (e.g. ALB path-based rules splitting traffic across several groups). This is how one ALB serves several backend services.
Deregistration delay (connection draining): default 300s. Once a target is deregistered (ASG scale-in, a deploy), the load balancer stops sending it new requests but lets in-flight ones finish before removing it, instead of dropping them mid-request.
Slow Start mode (ALB target group feature, frequently tested!): when enabled, a newly-registered healthy target gets a gradually ramping share of traffic over a configured warm-up window (30sβ15min) instead of jumping straight to a full, even share. Protects an app that needs a moment to actually perform well once traffic starts (JVM JIT warm-up, local cache fill, connection pool priming) from being swamped the instant it passes its health check. ALB only: NLB has no equivalent feature, since NLB has no application-layer visibility into "warmed up" at all; an exam option offering "configure a Network Load Balancer with slow start" is a distractor for exactly this reason.
πAuto Scaling Group Basics
What EC2 Auto Scaling is: a service that automatically launches and terminates EC2 instances in an Auto Scaling Group to keep capacity matched to actual demand (or a known schedule), instead of a fixed fleet size someone has to resize by hand. What it works with: a Launch Template (the blueprint for every instance it creates: AMI, instance type, key pair, security groups), CloudWatch alarms (the trigger: a metric breach fires a scaling policy), an Elastic Load Balancer (optional but typical; the ASG auto-registers new instances into the target group and deregisters ones it terminates, and can use the ELB's own health checks alongside its own), and plain EC2 instances as the thing actually being scaled. How it works: continuously compares current capacity against the Min/Desired/Max bounds and each active scaling policy's target, launches instances from the Launch Template when it needs to scale out, health-checks every instance and terminates+replaces any that fail, and terminates instances (respecting deregistration delay via the ELB, if attached) when it needs to scale back in.
Auto Scaling Group (ASG): the actual resource you create: a logical, named collection of EC2 instances that AWS manages as one unit. You don't launch or track individual instances yourself; you point EC2 Auto Scaling (the service, above) at the group, and it adds/removes instances within that group to satisfy the group's own settings. An ASG is defined by a Launch Template (which AMI, instance type, etc.) plus three numbers: Min, Desired, and Max capacity.
Spans multiple AZs by design: the ASG tries to balance instances evenly across the AZs you configure, for the same resilience reasons as any multi-AZ design.
An unhealthy instance (per its health check) is automatically terminated and replaced. This is EC2-fleet-level self-healing, distinct from Auto Recovery (which recovers the same instance after a hardware failure; see the EC2 & Compute topic). See "ASG Health Checks: EC2 vs. ELB" below for what counts as unhealthy and the grace period that protects a slow-booting instance from this.
Scaling up vs. scaling out (frequently tested!):Scaling up (vertical) = swap an instance for a bigger one: more vCPU/RAM on the same single box. Scaling out (horizontal) = add more instances of the same size to share the load. ASG only ever scales out: it changes the instance count, never the instance size, which is exactly why the fleet needs to be stateless (see below): any of those interchangeable new instances has to be able to pick up traffic with zero special setup. Example: Amazon's checkout service on Black Friday scales out: the ASG launches more identical web/app instances behind the ALB to absorb the surge, rather than scaling up, which would mean resizing each existing server to a bigger type (a slower, disruptive change with a hard ceiling once you hit the largest instance size).
Typical scale-up use caseA traditional single-writer relational database (RDS), e.g. resizing to a bigger DB instance class when queries are CPU/memory-bound, since you generally can't just add more independent writer nodes the way you can with stateless web servers
Typical scale-out use caseA stateless website/web or app tier behind an ALB, e.g. an ASG adding more identical EC2 instances as request volume grows, exactly the checkout-service pattern above
Scaling out is the generally preferred option, and the exam's default-favored answer, when the architecture allows it. Reasons: no downtime (new instances join alongside the existing ones, so nothing needs a stop/resize/restart cycle); no hard ceiling (keep adding instances, versus scaling up which eventually hits the biggest instance type AWS offers); better fault tolerance (load spread across many instances/AZs, instead of everything riding on one bigger single point of failure); and it matches the pay-for-what-you-use elasticity ASG is built for: capacity can shrink back out just as easily once demand drops, which a bigger single instance can't do without another disruptive resize. It's only off the table when the component itself can't be horizontally distributed, the single-writer relational database case above being the classic example.
πConfiguring a Launch Template
A Launch Template is the reusable blueprint an ASG (or a one-off "Launch instance from template" action) uses to create every instance. What you configure in it:
AMIWhich machine image to boot from: OS, pre-installed software, any baked-in config (see AMIs in EC2 & Compute)
Instance typee.g. t3.micro, m5.large; vCPU/RAM/network profile for every instance created, or a list of compatible types if using a Mixed Instances Policy (below)
Key pairFor SSH access, often skipped in favor of Session Manager on fleets managed entirely through automation (see Connecting to an Instance in EC2 & Compute)
Security group(s)Which inbound/outbound firewall rules every launched instance gets
StorageRoot EBS volume size/type, plus any additional volumes to attach on launch
IAM instance profileThe Role every instance assumes, so it can call other AWS APIs without an embedded Access Key (see Roles + EC2 in IAM)
User dataA bootstrap script that runs once on first boot (installs software, pulls app code, joins a cluster, etc.), the mechanism that lets a freshly launched instance become "ready" with zero manual setup
Network settingsAuto-assign public IP, etc. For an ASG specifically, the actual subnets/AZs it launches into are set on the ASG itself, not the template
Versioning: editing a Launch Template creates a new numbered version instead of overwriting the old one. Point an ASG at a specific version, or at $Latest/a chosen $Default so new instances automatically pick up template changes (or don't, if pinned to a fixed version). This versioning is the main reason Launch Templates replaced the older Launch Configurations, which are immutable once created: change anything and you must create and swap in an entirely new one.
Mixed Instances Policy: an ASG feature layered on top of a Launch Template: give it a list of compatible instance types (and optionally a mix of On-Demand + Spot) instead of one fixed type, and let the ASG pick whichever is available/cheapest at launch time. Not available with a legacy Launch Configuration.
Exam pattern: "Launch Configuration" still appears as a distractor answer on the exam. AWS no longer recommends creating new ones (no versioning, no Mixed Instances Policy, no Spot/On-Demand mixing). For a new ASG, the correct answer is always Launch Template.
An ASG's health check type setting decides which signal it trusts to decide an instance is unhealthy and due for replacement, a different, group-level setting from the System/Instance Status Checks covered in EC2 & Compute (those drive Auto Recovery for a single instance; this drives the ASG replacing instances across the whole group).
EC2 health check (default)
ELB health check
What it checks
Just the instance's own EC2 status checks: is it running at all (see Status Checks & Auto Recovery in EC2 & Compute)
The load balancer's configured health check against the target: a real request to a path/port, so it also catches an app that's up but broken (e.g. returning 500s, or hung)
Catches
A crashed/unreachable instance
Everything EC2 health checks catch, plus an instance that's healthy at the OS level but whose application has failed
Enabled by default?
Yes (always on)
No, opt-in. Must be explicitly turned on for the ASG even when a load balancer/target group is already attached
Exam pattern: "instances all pass their status checks but the application behind them is actually erroring, and the ASG isn't replacing them" β the ASG is relying on EC2 health checks only; enable ELB health checks so it also trusts the load balancer's application-level result.
Health check grace period (frequently tested!): new instances get a configurable grace period (default 300s) before a failed health check (of either type) is allowed to terminate them. Without it, a slow-booting app (still running its user-data script, still warming up) gets marked unhealthy and killed/relaunched in a loop before it ever finishes starting. Set it to comfortably cover the instance's real boot + app-startup time.
πAuto Scaling & CloudWatch Monitoring
Three different monitoring granularities can feed the CloudWatch metrics a scaling policy reacts to. Don't confuse "how often" with "at what level":
Scope
Interval
Cost
Group metrics
ASG-level (aggregated across the whole group, e.g. group-average CPU)
1 minute
Free, but must be explicitly enabled, not on by default
Basic monitoring
Per-instance
5 minutes
Free (the default for every EC2 instance)
Detailed monitoring
Per-instance
1 minute
Opt-in, and unlike the other two, charges apply
Exam pattern: a scaling policy needs to react quickly (sub-5-minute) to real load; the cheaper fix is usually enabling Group metrics on the ASG itself, not paying for Detailed monitoring on every individual instance, unless the question specifically needs per-instance-level 1-minute data.
CooldownDefault 300 seconds (5 minutes). Pairs specifically with Simple Scaling: after a scaling activity, the ASG waits out the cooldown before launching or terminating again, giving the previous change time to actually show up in the metric before reacting further (see Simple vs. Step Scaling above for why Step Scaling isn't blocked by this the same way).
Termination PolicyControls which instance(s) get picked first when a scale-in event needs to remove capacity. The default policy balances across AZs first (so scale-in doesn't leave one AZ empty), then picks the instance using the oldest Launch Template/Configuration, then whichever is closest to its next billing hour. Can be customized (e.g. OldestInstance, NewestInstance) if a question calls for different behavior.
Instance Protection (ASG-scoped Termination Protection)A per-instance flag that excludes that one instance from being picked during scale-in, even if the Termination Policy would otherwise choose it. Distinct from EC2's own account/instance-level Termination Protection (DisableApiTermination; see EC2 & Compute), which blocks a manual/API TerminateInstances call entirely rather than just influencing which instance scale-in picks.
Standby StateManually move a running instance from InService to Standby to patch or troubleshoot it without terminating it or leaving it serving live traffic. While in Standby, the ASG stops health-checking it and excludes it from the desired-capacity count (so it may launch a replacement to make up the difference). Move it back to InService when you're done.
Lifecycle HooksPause an instance in a Pending:Wait state (after launch, before it goes InService) or a Terminating:Wait state (after it's picked for termination, before it actually shuts down) for up to 1 hour by default (extendable via a heartbeat) so a custom action can run. Use cases: download/install software or run final setup before the instance takes traffic; drain in-flight connections or flush/upload data before it's terminated. The ASG publishes the hook event via SNS/EventBridge; the instance stays paused until your code calls CompleteLifecycleAction or the timeout elapses.
Exam pattern: "need Auto Scaling to actually wait while a custom script runs before an instance goes live, or before it's torn down" β Lifecycle Hooks. Launch Template user data alone runs a bootstrap script, but doesn't pause the ASG's own state transition the way a hook does.
πΊοΈTraffic Flow: Internet β ALB β Target Group β ASG
The ALB only ever routes to healthy instances in the target group; the ASG keeps the group at (or moving toward) the desired count across both AZs.
ποΈTypes of Auto Scaling & Scaling Policies
Four types of Auto Scaling, from least to most automated:Manual: you directly change Min/Desired/Max yourself via console/CLI/API, no automation involved. Dynamic: a CloudWatch alarm reacts to a live metric breach (Target Tracking, Step, or Simple Scaling: the three policy types below). Scheduled: a capacity change fires at a known date/time. Predictive: ML forecasts future load from historical patterns and scales ahead of it automatically, with no alarm or schedule to configure.
ManualChange Min/Desired/Max by hand: no CloudWatch alarm, no schedule, no forecast. Fine for a one-off/rare change; not something you'd rely on for routine demand handling.
Target Tracking"Keep average CPU at 50%": you pick a target metric value, AWS calculates the scaling math for you. The default, easiest, most-recommended approach. Not limited to CPU: any CloudWatch metric works, including a custom metric like an SQS queue's ApproximateNumberOfMessagesVisible (worker-fleet ASGs scaling on backlog size rather than their own CPU), or a built-in ALB metric like ALBRequestCountPerTarget; scale on requests-completed-per-instance instead of CPU when request handling time (not CPU) is what actually strains each instance. The exam uses this to test whether you assume Target Tracking = CPU only.
Step ScalingDifferent scaling amounts depending on how far a CloudWatch alarm's breach goes (e.g. +1 instance if CPU >70%, +3 if CPU >90%); more control, more setup
Simple ScalingOne alarm, one fixed action, then waits out a cooldown before evaluating again. The older, simpler predecessor to Step Scaling
Scheduled ScalingChange capacity at a known future time, e.g. scale up before a Friday-night traffic pattern you already know about
Predictive ScalingUses ML on historical load patterns to scale ahead of forecasted demand, instead of reacting after the metric breaches
Exam pattern: "simplest way to maintain a target CPU/request-count" β Target Tracking, essentially always the expected default answer.
Simple vs. Step Scaling: the classic exam differentiator is cooldown behavior, not just "more granular":
Simple Scaling
Step Scaling
Alarm β action
One alarm, one fixed action (e.g. average CPU >70% β +1 instance)
Multiple actions sized to how far the breach goes off one alarm (e.g. average CPU 70β90% β +1, >90% β +3)
Cooldown behavior
Blind during cooldown: must wait the full cooldown period out after firing before it will evaluate the alarm again, even if CPU keeps climbing
Can keep adjusting capacity while still in cooldown if the alarm remains in ALARM state, reacting to the current step instead of waiting it out
Exam framing
Older, simpler predecessor; rarely the "best" answer once Step Scaling is also an option
Preferred whenever the question wants a faster or finer-grained response to how severe the breach is, not just that it happened
Default Instance Warmup (frequently tested; easily confused with the health check grace period and ALB Slow Start above, but a distinct third thing): a configurable period during which a newly-launching instance is excluded from the ASG's own aggregated CloudWatch metrics (e.g. group-average CPU). Without it, a still-booting instance's artificially low/spiky CPU can skew the metric a Target Tracking/Step policy is reacting to, causing the ASG to misjudge real demand and launch more instances than actually needed (overscaling) while everything is still warming up. This is scaling-math protection, different from the health check grace period (which only stops a slow-booting instance from being wrongly killed) and from ALB Slow Start (which only throttles how much live traffic a new target gets); a question can legitimately need all three configured together on the same ASG.
π§ͺWorked Example: Three Things That Can Change an ASG's Instance Count (frequently tested!)
Same ASG (min 2 / desired 4 / max 10): three completely independent triggers can each change what's running. The exam likes to test whether you can tell these apart, especially #1 vs. #2 (only one of them is actually "scaling"):
Trigger
What happens
Worked example
1. Metric-based (dynamic) scaling
A CloudWatch alarm watches a metric (usually average CPU across the whole group) and fires a scaling policy once it breaches a threshold for the alarm's evaluation period
Average CPU across the ASG hits and holds 80% β the CloudWatch alarm goes into ALARM state β a Target Tracking or Step Scaling policy adds instances (e.g. +2) from the Launch Template β CPU falls back toward target as load spreads across more instances β a separate low-CPU alarm can later trigger scale-in
2. Health-check-driven replacement
Not a scaling policy at all: no CloudWatch alarm, no policy evaluation. The ASG continuously health-checks every instance and swaps out any that fails, purely to hold the group at its existing desired capacity
One instance fails its EC2 or ELB health check ("a unit goes down") β ASG terminates it and launches a replacement from the same Launch Template β desired capacity stays at 4 throughout; it's a 1-for-1 swap, never a scale-out
3. Scheduled scaling
A scaling action set for a specific date/time (once) or a recurring cron-style schedule, for load you already know is coming; no metric involved
A retail site's traffic reliably jumps every Monday morning when staff return and place orders β a scheduled action raises min/desired capacity ahead of that, e.g. 07:00 every Monday β desired 8, then a second scheduled action scales back down Friday evening β nothing has to breach a threshold first, because the pattern is already known rather than detected
Exam pattern: "predictable, recurring load (payroll runs, Monday-morning traffic, Black Friday)" β Scheduled Scaling: you already know it's coming, no need to wait for a metric to breach. "React to unpredictable load" β CloudWatch-alarm-driven Target Tracking/Step Scaling. "An instance/AZ fails" β this is not a scaling-policy trigger at all. It's the ASG's own health check replacing a broken instance to hold the group at its already-existing desired count, which is why it can happen even when the group is sitting well under Max.
πStateful vs. Stateless (frequently tested!)
Stateless: any instance can answer any request because it holds no session/user data locally: it reads whatever context it needs from a shared store instead. This is the design ASG assumes: instances get terminated and replaced constantly (scale-in, unhealthy replacement, AZ failure), so anything kept only in an instance's memory or local disk is lost the moment that instance goes. Example: checking the weather on a public site: you type a city, get a forecast, and nothing about that visit is remembered; the very next request could land on a completely different instance with no loss of function.
Stateful: an instance holds session/user data locally (in-memory or on local disk), so a request has to keep landing on that same instance to work correctly. Fights against elastic scaling: kill that one instance and its state is gone. Example: shopping on Amazon: your cart contents, browsing/purchase history, and login session all have to persist and follow you across every subsequent request, not just the one instance that first handled you. (In practice AWS-scale apps like this externalize that state (see the table below) rather than pinning you to one instance, which is exactly the "stateless design, statefully-feeling app" pattern the exam is testing for.)
Three-tier breakdown of the Amazon example:Web tier (renders the storefront pages): stateless, any instance can serve any request. Application tier (checkout/order logic), the layer that handles your purchase: it validates the order and writes it to a database. Data tier (RDS/DynamoDB): where the order actually becomes durable, permanently tied to your account regardless of which web/app instance you hit next time.
Exam nuance: "writes to a database" is not the same as "is stateful" in the ASG sense. An application-tier instance that records your order straight to the database and keeps nothing about you in its own memory afterward is still stateless and fully replaceable: ASG can terminate it a second later with zero data loss. It's the database, not the compute instance, that's genuinely stateful here (it's the one thing in this whole flow that must not be casually replaced/wiped). This is exactly why the exam's recommended pattern is: keep every compute tier (web and app) stateless, and push all real state down into RDS/DynamoDB/ElastiCache.
Any healthy target in the group (normal round-robin/least-outstanding-requests)
Sticky sessions (session affinity) pin a client to the same target via a cookie
Effect of scale-in / instance replacement
No data loss: any other instance already has access to the same shared state
That client's session data is lost when its instance is terminated
Exam-favored design
Yes: the "best practice" answer whenever ASG/ELB is in play
Only when the question explicitly requires sticky sessions (legacy app, no session-store option)
Sticky sessions (ALB feature): a load-balancer-generated cookie pins a client to one target for the life of the session. This makes a stateful app work despite multiple targets, but it doesn't make the app stateless, and it can unbalance load onto whichever instances got "sticky" early.
Exam pattern: "app stores session data on the instance and users get logged out/lose their cart after scaling events" β the fix is to externalize session state (ElastiCache or DynamoDB), not just to enable sticky sessions: sticky sessions patch the symptom, externalizing state removes the dependency on any one instance entirely.
ποΈArchitecture Patterns: Auto Scaling and ELB (frequently tested!)
A set of specific, exam-style requirement β solution pairings: the kind of one-line scenario the exam drops into a question stem, expecting you to recognize the pattern instantly.
Requirement
Solution
High availability and elastic scalability for web servers
EC2 Auto Scaling + an Application Load Balancer, spread across multiple AZs
Low-latency connections over UDP to a pool of instances running a gaming application
Network Load Balancer with a UDP listener (NLB is the only ELB type that handles UDP at all)
Clients need to whitelist static IP addresses for a highly available load-balanced application in a Region
NLB: it exposes one static IP per AZ automatically, unlike ALB/CLB which only ever have DNS names
Application on EC2 in an ASG requires disaster recovery across Regions
Create a matching ASG in a second Region with capacity set to 0; take AMI/EBS snapshots and copy them across Regions on a schedule (Lambda or Data Lifecycle Manager). The standby ASG scales up from those snapshots only if the primary Region fails
Application on EC2 must scale in larger increments for a big traffic increase than for a small one
Step Scaling with a larger capacity-increase step configured for the bigger breach (see Simple vs. Step Scaling above)
Need to scale EC2 instances behind an ALB based on the number of requests each instance has completed
Target Tracking policy on the built-in ALBRequestCountPerTarget metric (see the Types of Auto Scaling section above)
Application runs on EC2 behind an ALB; once authenticated, a user shouldn't have to reauthenticate if their instance fails
Externalize session state to DynamoDB or ElastiCache. See Stateful vs. Stateless above; don't rely on sticky sessions alone
Company is deploying an IDS/IPS system using virtual appliances and needs it to scale horizontally
Gateway Load Balancer in front of the virtual-appliance fleet (see Load Balancer Types above)
π§Putting It Together: Which One Do I Reach For? (frequently tested!)
The exam loves to describe a scenario in plain business language and make you pick the right ELB/ASG/session-state tool. This table collects the "if the question says X, the answer is Y" patterns from across this whole topic in one place.
Scenario
Reach for
Why
[Scaling] Traffic reliably spikes at a known time (payroll runs, Monday mornings, Black Friday)
Scheduled Scaling
You already know it's coming, no need to wait for a metric to breach first
[Scaling] Traffic is variable/unpredictable, want to hold average CPU (or another metric) at a target automatically
Target Tracking (Dynamic)
Simplest, most-recommended default; you state the goal, AWS does the scaling math
[Scaling] Need different-sized reactions depending on how severe the breach is, and to keep reacting through a fast-moving spike
Step Scaling
Multiple actions off one alarm, and, unlike Simple Scaling, keeps adjusting during cooldown
[Scaling] One-off, rare capacity change (a planned migration, a short test)
Manual scaling
Not worth building automation for something that happens once
[Scaling] A big known future event with historical seasonal data to learn from (e.g. last year's Black Friday numbers)
Predictive Scaling
ML forecasts demand and pre-scales ahead of it, instead of reacting after a threshold breach
[Scaling] A relational, single-writer database (RDS) is CPU/memory-bound
Scale up (bigger instance class)
You generally can't add more independent writer nodes the way you can with stateless web servers
[Scaling] A stateless web/app tier is running out of headroom as request volume grows
Scale out (ASG adds more instances)
Any instance can answer any request, so more identical instances behind the LB is the natural fix
[Load Balancer] Public web app needs path-based or host-based routing (/api/* vs /images/*)
ALB
Layer 7, the only type that reads request content to route on
[Load Balancer] Need extreme throughput/ultra-low latency, or a static IP for a client allow-list
NLB
Layer 4, built for raw speed and static IP support, not content-aware routing
[Load Balancer] Need to insert third-party firewall/IDS/IPS appliances transparently in front of traffic
GWLB
Layer 3 transparent pass-through built specifically for security-appliance fleets
[Load Balancer] A backend/microservice tier should only ever be reachable from inside the VPC, never the internet
Internal load balancer
DNS resolves to private IPs only, deployed in private subnets; simpler and more correct than trying to lock it down purely with security groups
[Load Balancer] Uneven traffic across AZs because instance counts per AZ don't match
Enable Cross-Zone Load Balancing (already on for ALB; opt-in for NLB)
Spreads every LB node's traffic across all targets in all AZs, not just the targets in its own AZ
[Session State] Users get logged out or lose their cart after a scale-in/replacement event
Externalize session state (ElastiCache or DynamoDB)
Removes the dependency on any one instance entirely (the actual fix, not a patch)
[Session State] A legacy app can't be quickly refactored to share session state externally
Sticky sessions (ALB)
Makes a stateful app work across multiple targets as a stopgap, but doesn't remove the underlying risk of losing that instance
[Health] A newly-launched instance keeps getting killed as "unhealthy" while it's still booting
Increase the Health Check Grace Period
Gives the app time to finish starting before health checks start counting against it
[Health] Instances pass their status checks but the app returns errors under real requests
Enable/use the ELB (target group) health check, not just the EC2 status check
Only the ELB health check makes a real request to the app; EC2 status checks only see infrastructure-level failures
π§ Quick Memory Hooks
ALB reads the letter (URL path/host): layer 7. NLB just moves the envelope fast: layer 4. GWLB is a transparent tollbooth for security appliances. Load Balancer = where traffic goes. Auto Scaling Group = how many places it can go. Target Tracking = tell AWS the goal, not the math. Stateless = any instance can answer, because the state isn't on the instance. Sticky sessions only patch a stateful app: they don't remove the risk. Cross-Zone Load Balancing = ALB shares fairly for free; NLB has to be asked, and it costs to ask.
π§ Route 53 & DNS: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Route 53 is AWS's DNS service: it answers "what IP does this name point to," and its routing policies are really just different rules for which answer to give.
Also a domain registrar: two separate jobs bundled under one name.
πHosted Zones
A Hosted Zone is a container of DNS records for one domain (e.g. example.com), where you actually create A, CNAME, MX, etc. records.
Public Hosted ZoneAnswers queries from the public internet (a normal, internet-facing domain)
Private Hosted ZoneOnly resolves inside one or more associated VPCs: internal-only DNS names, invisible to the outside world
πRecord Types: Alias vs. CNAME
A record β domain name to IPv4. AAAA β IPv6. CNAME β domain name to another domain name (not an IP).
Alias record (AWS-only, frequently tested): looks like a CNAME but works at the zone apex (example.com, not just www.example.com; a plain CNAME is not allowed at the apex per the DNS spec) and points straight at an AWS resource (ALB, CloudFront, S3 website endpoint) by its AWS-internal identity, not just its current hostname. Free to query, and automatically tracks the target's IP if it changes.
Exam pattern: "point the domain's root/apex directly at an ALB/CloudFront distribution" β Alias record, not CNAME (CNAME literally can't be used at the apex).
π§Routing Policies (frequently tested!)
Policy
What it does
Typical use case
Simple
One record, one (or a random) answer, no logic
A single-server site, no need for anything fancier
Weighted
Split traffic across records by assigned weight/percentage
Canary/A-B testing a new version, gradual migration
Latency-based
Route to whichever Region gives the requester the lowest latency
Global app, multiple Regional deployments, want the fastest response
Failover
Primary/secondary pair; routes to secondary only when primary's health check fails
Active-passive DR setup
Geolocation
Route by the requester's geographic location
Legal/licensing restrictions, localized content by country
Geoproximity
Route by geographic distance, with an adjustable "bias" to shift traffic toward/away from a Region
Fine-tuned traffic shaping by location (Traffic Flow feature)
Multi-value Answer
Returns several healthy IPs in response, client picks one
Simple client-side load distribution + basic health checking, not a substitute for a real load balancer
Exam pattern: "lowest latency for global users" β Latency-based. "Active-passive DR, switch on failure" β Failover. "Gradually shift 10% of traffic to a new version" β Weighted.
Route 53 can monitor an endpoint (HTTP/HTTPS/TCP) from multiple global locations and mark it healthy/unhealthy.
Unhealthy records are automatically excluded from what Route 53 returns. This is the mechanism behind Failover routing, and it also works within Weighted/Latency/Multi-value policies to skip a failed target.
Can also health-check a CloudWatch alarm instead of an endpoint directly, useful for checking something not reachable over the network (e.g. a database's internal metric).
π·οΈDomain Registration
Route 53 can also act as the registrar, the entity that reserves a domain name for you (separate from DNS hosting, though it usually sets up a matching Hosted Zone automatically).
You can register a domain elsewhere and still point it at a Route 53 Hosted Zone: registrar and DNS host don't have to be the same service.
π§ Quick Memory Hooks
Alias = the only way to point the bare domain at an AWS resource, and it's free. Latency = fastest route. Failover = backup plan. Weighted = percentage split. Geolocation = where the visitor is from.
Health checks are what let Route 53 skip a broken record instead of blindly handing it out.
π‘ CloudFront & Edge: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
CloudFront is AWS's CDN: it caches your content at locations physically close to viewers, so most requests never have to travel back to the origin at all.
π―Origins
The Origin is where CloudFront fetches content from when it doesn't already have a fresh cached copy.
S3 OriginStatic assets, website hosting (the most common origin)
Custom OriginAnything with a public HTTP(S) endpoint: an ALB, EC2 instance, or even a non-AWS server
One distribution can route to multiple origins based on the request path (Origin Groups can also fail over from a primary to a secondary origin).
πOAC (Origin Access Control)
Locks an S3 origin down so it's only reachable through CloudFront; direct requests straight to the bucket's S3 URL are rejected.
Works by having CloudFront sign every request to the origin with SigV4; the bucket policy then only trusts that specific distribution.
OAC replaces the older OAI (Origin Access Identity). It's AWS's current recommendation for any new distribution. Exam material may still mention OAI by name; know it's the predecessor being phased out, same underlying purpose.
β‘Caching Behavior
A cache hit never touches the origin at all. This is the entire performance and cost benefit of a CDN.
TTL controls how long an object stays cached before CloudFront re-checks the origin. Can be set per-object (origin's Cache-Control header) or overridden at the distribution level.
Cache key can include (or ignore) specific headers, cookies, and query strings. Configuring this correctly matters: caching per-user content under one shared key would leak one viewer's response to another.
Invalidation forces CloudFront to drop cached copies of specific paths before their TTL expires, useful right after a deploy, but has a cost per invalidation path at scale (versioned filenames are the cheaper long-term pattern).
πEdge Locations vs. Regional Edge Caches
Edge Locations: hundreds of small points of presence worldwide: the first stop for a viewer's request, holding the most popular/recently-requested content.
Regional Edge Caches: fewer, larger caches sitting between edge locations and the origin; catch less-popular content that fell out of an edge location's cache, reducing how often the origin itself gets hit.
ποΈSigned URLs & Signed Cookies
Restrict access to private content served through CloudFront. A request without a valid signature is rejected.
Signed URLOne URL, one object (or a few), e.g. a single paid download link with an expiry
Signed CookieOne cookie grants access to many objects, e.g. an entire video course a logged-in user has paid for
π‘οΈAWS WAF & AWS Shield (frequently tested!)
Two different edge-protection services, commonly confused because both attach to the same resources (CloudFront, ALB, API Gateway, AppSync); they defend against completely different attack types.
You define rules (or use AWS Managed Rule groups) that inspect each request and Allow/Block/Count it. This is where "block SQLi/XSS" configuration actually lives
Detects and automatically mitigates attack traffic, nothing to author yourself
Tiers
Pay per rule + per request; no free tier, but no separate Standard/Advanced split either
Standard: free, automatic, protects every AWS customer by default. Advanced: paid, adds near-real-time attack visibility, cost protection against scaling charges caused by an attack, and 24/7 access to the AWS Shield Response Team (SRT)
Exam pattern: "block SQL injection / cross-site scripting requests" β AWS WAF with a rule/managed rule group, never Shield (Shield has no concept of inspecting request content for an exploit pattern). "Protect against a DDoS attack" β Shield. AWS Firewall Manager is a distinct third service: it centrally deploys and enforces WAF rules (and Shield Advanced protections) across many accounts/resources in an Organization at once, rather than doing the blocking itself.
π§ Quick Memory Hooks
OAC = the lock on the origin door; only CloudFront has the key. Edge Location = the corner shop (fast, small, popular items). Regional Edge Cache = the regional warehouse behind it. Signed URL = one ticket, one item. Signed Cookie = a season pass to everything. WAF reads the request content (SQLi/XSS rules). Shield just absorbs the flood (DDoS); it never looks inside a request.
β‘ Lambda & Serverless: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Lambda runs your code without a server to manage: you pay per invocation/duration, and AWS handles all the provisioning. The exam mostly tests how a given trigger invokes it, not the code itself.
Max execution time: 15 minutes per invocation. Long-running work beyond that needs Step Functions, ECS/Fargate, or a different compute model entirely.
Memory (128 MBβ10,240 MB) and CPU are coupled: you only set memory, and CPU allocation scales with it automatically.
Charged per invocation + duration (rounded to the millisecond): near-zero cost when nothing is running, unlike an idle EC2 instance you're still paying for.
πInvocation Types (frequently tested!)
The trigger's identity determines the invocation type, not something you choose independently.
Type
Example triggers
Retry behavior
Synchronous
API Gateway, ALB, direct SDK call
None: caller sees the error immediately, handles retry itself
Asynchronous
S3, SNS, EventBridge
Automatic retries, then a Dead-Letter Queue or on-failure Destination if still failing
Poll-based
SQS, Kinesis, DynamoDB Streams
Lambda service itself polls and re-tries batches; a persistently-failing record can block the batch, hence configurable retry/skip settings per source
π§΅Concurrency
Lambda scales by running many concurrent copies of your function, one per in-flight invocation, not by scaling a single instance's throughput.
Reserved ConcurrencyCaps how many concurrent executions a function can use; protects other functions from being starved of the account-wide limit, but throttles this one once the cap is hit
Provisioned ConcurrencyKeeps a set number of execution environments pre-initialized and warm; eliminates cold starts for that reserved capacity, at a standing cost even when idle
βοΈCold Starts
A cold start happens when Lambda has to spin up a brand-new execution environment (download code, start the runtime) before running your function. This adds noticeable latency to that one request.
A warm invocation reuses an already-initialized environment from a recent invocation: fast, no setup overhead.
Mitigations: Provisioned Concurrency (guaranteed warm), smaller deployment packages, avoiding heavyweight SDK initialization in the function's global scope where it can't be reused efficiently.
πLambda + VPC (frequently tested!)
By default a Lambda function runs outside any customer VPC, in an AWS-managed network: it can reach the public internet and any public AWS service endpoint, but has no route to a private resource (an RDS instance in a private subnet, an internal-only ElastiCache cluster).
To reach a private VPC resource, configure the function with VPC settings: the subnets and security group(s) it should launch into. Lambda then attaches an ENI in those subnets to route traffic, the same fundamental mechanism an EC2 instance uses.
Trade-off: a VPC-attached function loses default internet access. Once attached to private subnets, outbound internet calls (e.g. to a public third-party API) need the same NAT Gateway/Instance path any other private-subnet resource would use. Attaching to a VPC does not, by itself, add internet access, it only adds VPC access.
Historical cold-start cost (largely resolved): earlier Lambda networking created a new ENI per unique subnet/security-group combination on a cold start, which could take tens of seconds, a frequently-cited exam gotcha. AWS's Hyperplane-based networking model now pre-creates and shares ENIs across functions/accounts, so a modern VPC-attached function's cold start is close to a non-VPC function's. Older exam material may still describe VPC attachment as a heavy cold-start penalty; know both the old reputation and the current reality.
Exam pattern: "Lambda function times out trying to reach an RDS instance in a private subnet" β the function isn't configured with that VPC's subnets/security group yet. "VPC-attached Lambda function can no longer reach a public API it used to call" β it needs a NAT Gateway/Instance route from those subnets now, the same as any private-subnet resource.
πͺAPI Gateway
The front door for a serverless API: routes HTTP requests to backends (usually Lambda), handles auth, throttling, and request/response transformation.
REST APIFull feature set: request validation, API keys/usage plans, caching. More expensive, more setup.
HTTP APINewer, lighter, cheaper (~70% less); covers the common Lambda-backend case without the full REST API feature set
Exam pattern: "cheapest way to expose a Lambda function over HTTP with no advanced needs" β HTTP API.
πͺStep Functions
Orchestrates a sequence of steps (a state machine): Lambda calls, waits, parallel branches, choice logic, error handling/retries, without writing that coordination logic yourself.
Standard workflowsLong-running (up to 1 year), exactly-once execution, full execution history; for auditable, durable business processes
Express workflowsHigh-volume, short-duration (up to 5 min), at-least-once execution; for fast event-processing pipelines
Exam pattern: "coordinate multiple Lambda functions with retries, branching, and a visual audit trail" β Step Functions, instead of hand-rolling orchestration logic inside one giant Lambda.
π§ Quick Memory Hooks
Sync = wait on hold. Async = leave a voicemail (queued, retried). Poll-based = AWS checks your mailbox for you. Provisioned Concurrency = pay to keep the engine running so it's never cold. Step Functions = the flowchart that runs itself.
βοΈ Messaging: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Messaging services exist to decouple producers from consumers, so one side can fail, slow down, or scale independently without taking the other down with it.
πWhy Decouple?
Tightly coupled: Service A calls Service B directly and waits. If B is slow or down, A is stuck too.
Loosely coupled: Service A drops a message somewhere durable; Service B picks it up whenever it's ready. A never waits on B's health or speed.
This is also what absorbs a traffic spike: messages queue up instead of overwhelming the consumer, which processes at its own sustainable pace.
π¬SQS: Standard vs. FIFO
Standard
FIFO
Ordering
Best-effort, not guaranteed
Strict, first-in-first-out
Delivery
At-least-once (a duplicate is possible)
Exactly-once processing
Throughput
Nearly unlimited
Up to 3,000 msg/sec with batching (per API action)
Visibility TimeoutOnce a consumer receives a message, it's hidden from other consumers for this window. If the consumer doesn't delete it in time (crashed, still processing), it reappears for someone else to pick up
Dead-Letter Queue (DLQ)After a message fails processing too many times (maxReceiveCount), it's moved here instead of retrying forever. This lets you inspect poison-pill messages without blocking the main queue
Long PollingConsumer's receive call waits up to 20s for a message to arrive instead of returning empty immediately; cheaper (fewer empty API calls) than short polling
π’SNS & the Fan-out Pattern (frequently tested!)
SNS is pub/sub: one message published to a Topic is pushed out to every subscriber at once (email, SMS, Lambda, HTTP endpoint, or an SQS queue).
One publish, three independent queues: each consumer processes at its own pace, and a slow/broken consumer B doesn't affect A or C.
This SNSβmultiple-SQS pattern is the standard answer whenever a question describes "one event needs to trigger several independent downstream processes."
πEventBridge
An event bus: routes events matching a rule (pattern-matched on the event's content) to one or more targets, from a much wider range of sources than SNS.
SourcesAWS services natively, your own custom application events, and direct integrations with 3rd-party SaaS (e.g. Zendesk, Datadog)
SchedulingBuilt-in cron/rate-based scheduled rules, no need for a separate scheduler service
Content filteringRules can match on the event's actual field values, not just its source; more routing intelligence than an SNS topic
π§SQS vs. SNS vs. EventBridge: how to choose
SQSOne producer, one (logical) consumer group, need durable buffering/retry: a queue, not a broadcast
SNSOne event, multiple known subscribers, need it pushed out immediately (pub/sub)
EventBridgeEvent-driven architecture with content-based routing rules, scheduled events, or ingesting from SaaS/3rd-party sources
π§ Quick Memory Hooks
SQS = a mailbox (pull, buffered). SNS = a megaphone (push, broadcast). EventBridge = a smart switchboard (rules-based routing). Standard SQS = fast, maybe-duplicate. FIFO SQS = strict order, exactly-once, capped throughput.
Fan-out = SNS topic, many SQS queues, each consumer independent.
π Monitoring & Management: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
Four different questions, four different tools: Is it healthy right now? (CloudWatch) Β· Who did what? (CloudTrail) Β· What changed, and when? (Config) Β· Am I following best practice? (Trusted Advisor).
πCloudWatch
MetricsTime-series numeric data (CPU%, request count, queue depth). Most AWS services publish these automatically; custom metrics can be pushed from your own app
AlarmsWatch a metric against a threshold and trigger an action (SNS notification, Auto Scaling policy, EC2 Auto Recovery) when breached
LogsCentralized, searchable log storage: application/Lambda/VPC Flow Logs all ship here; Logs Insights lets you query them
DashboardsCustom visual boards combining metrics/alarms from across services into one view
Resolution: varies by service, e.g. EC2's basic monitoring is 5-minute granularity by default, with Detailed Monitoring (an extra cost) dropping that to 1-minute for faster alarm reaction; many other services (e.g. Lambda) already publish at 1-minute resolution with no separate "detailed" tier. Don't assume 5-minute-by-default applies account-wide.
βοΈCloudWatch vs. CloudTrail (frequently tested!)
An audit log of every API call (console, CLI, SDK) in the account
Enabled by default?
Basic metrics yes, most other features opt-in
Event history exists automatically with no setup (management events, rolling 90-day window), but that's not a "trail": for continuous delivery to S3/CloudWatch Logs and retention beyond 90 days, you must explicitly create a Trail
Exam pattern: "someone deleted a security group, find out who" β CloudTrail. "The app's error rate spiked, alert the team" β CloudWatch.
π°οΈAWS Config
Records a configuration history of your resources: what a resource's settings looked like at any past point in time, and every change in between.
Config Rules continuously check resources against a desired configuration (e.g. "every EBS volume must be encrypted") and flag non-compliant ones. This is compliance/drift detection, distinct from CloudTrail's "who did it" audit trail.
Exam pattern: "prove a security group's rules over the last 3 months" or "detect when a resource drifts out of compliance" β Config, not CloudTrail.
π‘Trusted Advisor
Automated checks across your account against AWS best practice, grouped into 6 categories: Cost Optimization, Performance, Security, Fault Tolerance, Service Limits, Operational Excellence.
Free tier gives a limited set of core checks (mostly security); the full check list requires a Business or Enterprise Support plan.
Overlaps in spirit with Config/Compute Optimizer but is broader and higher-level: a quick account-wide health/hygiene scorecard rather than deep resource history.
π οΈSystems Manager (SSM): brief mention
Parameter StoreFree, secure key-value config/secrets storage, a lighter alternative to Secrets Manager when rotation isn't needed
Session ManagerShell access to an instance with no open SSH port, no bastion host, no key pair; access is governed entirely by IAM
Distributed tracing: answers a different question than everything else on this page. CloudWatch tells you a metric is bad; X-Ray tells you where inside a multi-service request the time actually went.
A single request (e.g. API Gateway β Lambda β DynamoDB) is stitched into one trace, made of per-service segments (and finer subsegments within a service, e.g. one for the DynamoDB call specifically), visualized as a service map showing every hop and its latency.
Requires minimal code change: the X-Ray SDK (or, for supported services, a built-in "Active tracing" toggle; Lambda and API Gateway both have one) instruments the service without a rewrite.
Exam pattern: "find which specific downstream service/call is causing latency in a chain of microservices" β X-Ray, not CloudWatch (CloudWatch alerts you that latency is high overall; it doesn't show you the call chain that produced it).
π§ Quick Memory Hooks
CloudWatch = the vital-signs monitor. CloudTrail = the security camera log. Config = the change-history book. Trusted Advisor = the health inspector's checklist. X-Ray = the delivery tracking map, hop by hop.
If the question is about who, think CloudTrail. If it's about what changed over time, think Config. If it's about right now, think CloudWatch. If it's about where in the chain the time went, think X-Ray.
ποΈ Well-Architected: One-Page Study Guide
π Under Human Review. Newly drafted, not yet read through by the site owner. Treat as a study aid, not a finished/verified reference yet.
The Well-Architected Framework is AWS's lens for judging a design: many SAA questions are really asking "which pillar does this trade-off belong to," not testing new facts.
ποΈThe Six Pillars (frequently tested!)
Pillar
Core question
Operational Excellence
Can we run and monitor this reliably, and improve it over time (as code, with automation)?
Security
Is access, data, and infrastructure protected: least privilege, encryption, traceability?
Reliability
Does it recover from failure and meet demand without manual intervention (Multi-AZ, Auto Scaling, backups)?
Performance Efficiency
Are we using the right resource type/size for the workload, and adapting as it evolves?
Cost Optimization
Are we avoiding unnecessary spend: right-sizing, the right purchasing option, eliminating waste?
Sustainability
Are we minimizing the environmental impact of running this workload (added as the 6th pillar in 2021)?
Exam pattern: a question describing "encrypt data, apply least privilege" is testing Security; "add Multi-AZ, automate failover" is Reliability; "switch to Spot, right-size instances" is Cost Optimization. Even if the word "pillar" never appears, that's the lens being tested.
π€Shared Responsibility Model
AWS: Security of the Cloud
Customer: Security in the Cloud
Physical data centers, hardware, networking infrastructure
Data encryption, IAM configuration, security group/NACL rules
The virtualization layer (hypervisor)
Guest OS patching, application-level security
Managed service internals (e.g. RDS engine patching)
What you put in the database, who can access it
Rule of thumb: AWS secures the infrastructure the cloud runs on; you secure everything you configure and put in it. For managed services (RDS, Lambda), the line shifts further toward AWS than for unmanaged ones (raw EC2), but data and access control are always yours.
π°Cost Optimization Levers: a recap across this guide
Cost Optimization isn't its own set of new facts. It's the same tools covered elsewhere in this guide, applied with cost as the goal:
1Right-size: don't pay for provisioned capacity you don't use (EC2 & Compute topic)
2Match the purchasing option to the workload: Reserved/Savings Plans for steady-state, Spot for fault-tolerant/flexible (EC2 & Compute topic)
3Lifecycle old data to cheaper storage tiers automatically instead of leaving everything on Standard (S3 & Storage topic)
4Cache aggressively: CloudFront and ElastiCache both cut repeated, avoidable work (CloudFront & Edge, Databases topics)
5Go serverless where usage is spiky/idle: Lambda's near-zero idle cost beats a constantly-running server for bursty workloads (Lambda & Serverless topic)
π§°AWS Well-Architected Tool
A free console tool that walks through a structured questionnaire per pillar against your actual workload, then flags risks and links out to the specific guidance for fixing each one.
Not something you configure once, meant to be revisited as a workload evolves, the same way the pillars are a lens you keep re-applying, not a one-time checklist.
π§ Quick Memory Hooks
Six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, Sustainability. "Our Systems Run Pretty Cheap, Sustainably." AWS secures the cloud. You secure what's in it.
Every cost-optimization answer on this exam is really "stop paying for unused/oversized/idle capacity" in a new outfit.
π·οΈ AWS Service Names: Exam Reference
AWS Certification exams reduce reading load by using official short names for certain well-known services.
Know both the short name and the full name: a question may use either.
π·οΈOfficial Short Names (frequently tested!)
Short Name
Full Name
Example use case
AWS CDK
AWS Cloud Development Kit
Define your infrastructure in Python/TypeScript instead of hand-writing raw CloudFormation YAML
AWS CLI
AWS Command Line Interface
Script aws s3 cp in a deploy pipeline instead of clicking through the console
AWS DMS
AWS Database Migration Service
Migrate an on-prem Oracle database into RDS with minimal downtime
Amazon DocumentDB
Amazon DocumentDB (with MongoDB compatibility)
Run a MongoDB-style workload without managing MongoDB servers yourself
Amazon EBS
Amazon Elastic Block Store
Attach a persistent virtual hard drive to an EC2 instance
Amazon EC2
Amazon Elastic Compute Cloud
Launch a virtual server to host an application
Amazon ECR
Amazon Elastic Container Registry
Store and version your Docker container images before deploying them
Amazon ECS
Amazon Elastic Container Service
Run and orchestrate Docker containers without managing Kubernetes yourself
Amazon EFS
Amazon Elastic File System
Share one common file system across many EC2 instances at once
Amazon EKS
Amazon Elastic Kubernetes Service
Run a managed Kubernetes cluster when you specifically need Kubernetes
IAM
AWS Identity and Access Management
Add a new user, create a Role, or attach a permissions policy
Amazon Keyspaces
Amazon Keyspaces (for Apache Cassandra)
Run a Cassandra-compatible workload without managing Cassandra nodes
AWS KMS
AWS Key Management Service
Create and manage the encryption key used to encrypt an S3 bucket or EBS volume
AWS Managed Microsoft AD
AWS Directory Service for Microsoft Active Directory
Stand up a real Active Directory domain for Windows workloads, without running your own domain controllers
AWS Private CA
AWS Private Certificate Authority
Issue private TLS certificates for internal services that don't need a public CA
Amazon RDS
Amazon Relational Database Service
Launch a managed MySQL/PostgreSQL database without patching or backing it up yourself
Amazon S3
Amazon Simple Storage Service
Store and serve static files, backups, or website assets
AWS SAM
AWS Serverless Application Model
Define and deploy a serverless Lambda application from a simplified template
AWS SCT
AWS Schema Conversion Tool
Convert a database schema from one engine (e.g. Oracle) to another (e.g. PostgreSQL) before migrating with DMS
Amazon SES
Amazon Simple Email Service
Send transactional or marketing emails from an application
Amazon SNS
Amazon Simple Notification Service
Fan out one notification to many subscribers (email, SMS, Lambda, SQS) at once
Amazon SQS
Amazon Simple Queue Service
Decouple two application components with a durable message queue
AWS STS
AWS Security Token Service
Issue the temporary credentials handed out when a Role is assumed
Amazon VPC
Amazon Virtual Private Cloud
Create an isolated private network to launch your resources into
Source: AWS Certification General Information policy page, the authoritative, exam-official list. Check back there periodically since AWS can add to it.
Exam tip: if a question describes a scenario rather than naming a service, match the verb: "issue temporary credentials" β STS, "decouple" β SQS, "fan out" β SNS, "add a user" β IAM.
π§ Memory hook: pair the ones that get confused. SCT before DMS: convert the Schema, then migrate the Data. SNS pushes out, SQS holds until pulled.EBS = one disk for one instance, EFS = one shared filesystem for many instances, S3 = a bucket, not a drive at all.
π§ Quick Memory Hooks
This isn't a technical domain. It's exam vocabulary. A question naming "AWS STS" or "Amazon Keyspaces" cold is testing whether you recognize the service, not a new concept.
Acronyms you'll see spelled out constantly elsewhere in this guide (IAM, KMS, RDS, VPC, SNS, SQS) are exactly the ones on this list. That's not a coincidence.