AWS
Resolve AWS CLI v2 credentials and profiles, filter output, and run and debug common EC2, S3, IAM, logs, RDS, Lambda and EKS operations.
On this page
Cheatsheet#
| Task | Command |
|---|---|
| Which identity and account | aws sts get-caller-identity |
| Which profile, region and credential source | aws configure list |
| Log in with IAM Identity Center | aws sso login --profile prod |
| Log in with console credentials (CLI 2.32+) | aws login --profile dev |
| Assume a role | aws sts assume-role --role-arn <arn> --role-session-name alice |
| Filter output client-side | --query 'Reservations[].Instances[].InstanceId' --output text |
| Shell on an instance, no SSH or bastion | aws ssm start-session --target i-0abc123 |
| Tail a log group | aws logs tail /aws/lambda/my-fn --follow |
| Recent management API calls | aws cloudtrail lookup-events --max-results 20 |
| Preview a sync that deletes | aws s3 sync ./dist s3://my-bucket/site --delete --dryrun |
| Temporary download URL | aws s3 presign s3://my-bucket/key --expires-in 3600 |
| Decode an encoded authorisation failure | aws sts decode-authorization-message --encoded-message "$MSG" |
| Test a principal’s permissions | aws iam simulate-principal-policy --policy-source-arn <arn> --action-names s3:GetObject |
| Service quota | aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A |
| Block until a state is reached | aws ec2 wait instance-running --instance-ids i-0abc123 |
Everything here is AWS CLI v2. v1 lacks aws logs tail, SSO sessions and aws login, and treats binary blobs differently. Check with aws --version. See the AWS CLI v2 reference.
Credentials and identity#
For each setting the CLI takes the first source that provides it, in this order: command-line options (--profile, --region), environment variables (AWS_ACCESS_KEY_ID, AWS_PROFILE, AWS_REGION), then the profile’s configuration: assume role, assume role with web identity, IAM Identity Center (SSO), the credentials file, credential_process, the config file, then container credentials (ECS task role, EKS Pod Identity) and finally EC2 instance profile credentials from instance metadata. See configuration and credential precedence.
A stale AWS_ACCESS_KEY_ID exported in a shell therefore beats a correct --profile in the config file, and an instance role is used silently when nothing else is configured.
aws sts get-caller-identity # the reliable answer to "which account and role is this"
aws configure list # each setting with its source (env, config-file, iam-role)
env | grep '^AWS_' | cut -d= -f1 # which AWS variables are set, without printing values
AWS_PROFILE=prod aws s3 ls{
"UserId": "AROAEXAMPLEID:alice",
"Account": "123456789012",
"Arn": "arn:aws:sts::123456789012:assumed-role/PlatformEngineer/alice"
}# ~/.aws/config
[sso-session corp]
sso_start_url = https://example.awsapps.com/start
sso_region = ap-southeast-2
sso_registration_scopes = sso:account:access
[profile prod]
sso_session = corp
sso_account_id = 123456789012
sso_role_name = PlatformEngineer
region = ap-southeast-2
output = json
[profile prod-admin]
source_profile = prod
role_arn = arn:aws:iam::123456789012:role/Admin
duration_seconds = 3600aws sso login caches a token under ~/.aws/sso/cache; aws login (CLI 2.32.0+) does the same for console sign-in (root, IAM user or federated) under ~/.aws/login/cache, refreshing for up to 12 hours. aws configure export-credentials --profile prod --format env prints temporary credentials for tools that cannot read profiles; the output is secret.
Prefer short-lived credentials
Use IAM Identity Center or aws login for people and roles for workloads (instance profiles, ECS task roles, EKS Pod Identity or IRSA). A static access key must be rotated and kept out of repositories, images and CI logs.
Output, queries and pagination#
--filters (and parameters such as --prefix) are applied by the service before it responds. --query is a JMESPath expression applied by the CLI after every page has been downloaded. On a large account, filter server-side first and use --query to shape the result.
aws ec2 describe-instances --filters 'Name=instance-state-name,Values=running' \
--query 'Reservations[].Instances[].[InstanceId,InstanceType,PrivateIpAddress]' --output text
aws ec2 describe-instances \
--query 'Reservations[].Instances[].[InstanceId,Tags[?Key==`Name`].Value|[0]]' --output table
aws s3api list-objects-v2 --bucket my-bucket --prefix logs/ \
--query 'sort_by(Contents, &LastModified)[-5:].[Key,Size]' --output text| Option | Effect |
|---|---|
--output json|yaml|text|table | text is tab-separated for awk and cut; json with jq for nested data |
--no-paginate | Return only the first page |
--page-size n | Smaller API pages (avoids timeouts); still returns everything |
--max-items n | Stop after n items and print a NextToken for --starting-token |
--no-cli-pager | Disable the pager (AWS_PAGER="" does the same for a session) |
--cli-read-timeout, --cli-connect-timeout | Seconds before a request is abandoned; set in scripts |
In text output, a --query that selects nothing prints None. Test for it explicitly in scripts.
JMESPath patterns#
The same handful of constructs cover nearly every --query. Literals inside a filter use backticks; a string compared with a literal must be inside backticks or single quotes, and the whole expression is single-quoted for the shell.
# Projection: a list of fields per item
--query 'Reservations[].Instances[].[InstanceId,State.Name,PrivateIpAddress]'
# Filter with a comparison, then project
--query 'Volumes[?Size > `100`].[VolumeId,Size]'
--query 'Reservations[].Instances[?State.Name==`running`].InstanceId[]' # trailing [] flattens nested lists
# A tag value: filter the Tags list, take the first match, default when missing
--query 'Reservations[].Instances[].[InstanceId, Tags[?Key==`Name`].Value | [0] || `untagged`]'
# Multi-select hash: name the output keys, then --output table gives labelled columns
--query 'DBInstances[].{id:DBInstanceIdentifier,class:DBInstanceClass,status:DBInstanceStatus,az:AvailabilityZone}'
# Functions: sort, length, contains, starts_with, to_string, join
--query 'sort_by(Functions, &LastModified)[-3:].FunctionName'
--query 'length(Reservations[].Instances[])'
--query 'Buckets[?starts_with(Name, `prod-`)].Name'
--query 'Roles[?contains(RoleName, `Deploy`)].Arn'
--query 'join(`,`, Subnets[].SubnetId)'
# Boolean OR of conditions and negation
--query 'SecurityGroups[?GroupName!=`default` && length(IpPermissions)==`0`].GroupId'
# Pipe to re-shape the result of the left side
--query 'Reservations[].Instances[] | [?Platform!=`windows`] | length(@)'--query cannot compare dates or do arithmetic beyond comparisons of numbers; do that in jq with --output json. Test an expression against saved output with aws ec2 describe-instances --output json > ec2.json and the jp CLI, or iterate quickly with --no-cli-pager --output table.
EC2 and Systems Manager#
aws ec2 describe-instances --filters 'Name=tag:Environment,Values=prod' \
--query 'Reservations[].Instances[].[InstanceId,PrivateIpAddress,State.Name]' --output text
aws ec2 start-instances --instance-ids i-0abc123
aws ec2 stop-instances --instance-ids i-0abc123 # instance-store data is lost
aws ec2 describe-instance-status --instance-ids i-0abc123 # system and instance status checks
aws ec2 get-console-output --instance-id i-0abc123 --latest --output text | tail -50
aws ec2 create-image --instance-id i-0abc123 --name "backup-$(date +%F)" --no-reboot
aws ssm start-session --target i-0abc123 # needs the SSM agent and session-manager-plugin
aws ssm start-session --target i-0abc123 \
--document-name AWS-StartPortForwardingSession --parameters 'portNumber=5432,localPortNumber=15432'--no-reboot images a running filesystem, so the image may be inconsistent for databases.
Security groups are stateful: an inbound allow implies the reply traffic. Network ACLs are stateless and need rules in both directions, including the ephemeral port range for replies. That asymmetry explains most “the security group looks correct” cases.
# Launch from the latest Amazon Linux 2023 AMI via the public SSM parameter, with tags at creation
AMI=$(aws ssm get-parameter --name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 --query Parameter.Value --output text)
aws ec2 run-instances --image-id "$AMI" --instance-type t3.small --subnet-id subnet-0abc123 \
--security-group-ids sg-0abc123 --iam-instance-profile Name=ssm-managed \
--metadata-options HttpTokens=required \
--tag-specifications 'ResourceType=instance,Tags=[{Key=Name,Value=web-01},{Key=Environment,Value=prod}]' \
--user-data file://cloud-init.yml --query 'Instances[0].InstanceId' --output text
aws ec2 authorize-security-group-ingress --group-id sg-0abc123 --ip-permissions \
'IpProtocol=tcp,FromPort=443,ToPort=443,IpRanges=[{CidrIp=192.0.2.0/24,Description="office"}]'
aws ec2 describe-security-group-rules --filters Name=group-id,Values=sg-0abc123 --query 'SecurityGroupRules[].[SecurityGroupRuleId,IsEgress,IpProtocol,FromPort,CidrIpv4]' --output table
aws ec2 modify-instance-attribute --instance-id i-0abc123 --instance-type t3.large # stopped instance only
aws ec2 modify-instance-metadata-options --instance-id i-0abc123 --http-tokens required # enforce IMDSv2
aws ec2 create-tags --resources i-0abc123 vol-0abc123 --tags Key=Owner,Value=platform
aws ec2 terminate-instances --instance-ids i-0abc123 # irreversible; check DisableApiTermination first
aws ec2 describe-instance-attribute --instance-id i-0abc123 --attribute disableApiTerminationHttpTokens=required forces IMDSv2, which stops SSRF-style credential theft through the metadata service; make it the default on every launch template. --user-data is run once by cloud-init on first boot; changes to it on a stopped instance do not re-run unless the instance is rebuilt.
Systems Manager beyond sessions#
SSM also runs commands across a fleet selected by tag, stores configuration in Parameter Store, and reports inventory and patch state, all without opening inbound ports.
aws ssm describe-instance-information --query 'InstanceInformationList[].[InstanceId,PingStatus,PlatformName,AgentVersion]' --output table
aws ssm send-command --document-name AWS-RunShellScript --targets 'Key=tag:Environment,Values=prod' \
--parameters 'commands=["dnf -y check-update || true","systemctl is-active nginx"]' \
--comment "health check" --query Command.CommandId --output text
aws ssm list-command-invocations --command-id "$CMD_ID" --details \
--query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].Output]' --output text
aws ssm get-parameter --name /my-app/prod/db_url --with-decryption --query Parameter.Value --output text
aws ssm get-parameters-by-path --path /my-app/prod/ --recursive --with-decryption --query 'Parameters[].[Name,Version]' --output table
aws ssm put-parameter --name /my-app/prod/db_url --type SecureString --value "$DB_URL" --overwrite # value lands in shell history unless read from a variable
aws ssm describe-instance-patch-states --instance-ids i-0abc123 --query 'InstancePatchStates[].[InstanceId,MissingCount,FailedCount,OperationEndTime]' --output text
aws ssm start-session --target i-0abc123 --document-name AWS-StartInteractiveCommand --parameters 'command=["journalctl -u nginx -n 50"]'send-command returns immediately; poll list-command-invocations or pass --output-s3-bucket-name for output longer than the 2,500-character inline limit. Session Manager sessions are logged to CloudWatch or S3 when the SSM-SessionManagerRunShell document preferences say so, which is the audit trail SSH never had.
S3#
aws s3 ls s3://my-bucket/prefix/ --human-readable --summarize
aws s3 sync ./dist s3://my-bucket/site --delete --dryrun # always preview --delete
aws s3 cp file.tar s3://my-bucket/key --storage-class INTELLIGENT_TIERING
aws s3 presign s3://my-bucket/key --expires-in 3600
aws s3api head-object --bucket my-bucket --key key # size, encryption, metadata
aws s3api get-bucket-policy --bucket my-bucket --query Policy --output text | jq
aws s3api list-object-versions --bucket my-bucket --prefix key \
--query 'Versions[].[Key,VersionId,IsLatest]' --output text
aws s3api get-public-access-block --bucket my-bucketaws s3 is the high-level interface (sync, recursive copy, automatic multipart). aws s3api maps one-to-one to API operations for policies, versions and metadata.
sync --delete and rm --recursive are immediate
Without versioning there is no undo. Run with --dryrun first and check the account with aws sts get-caller-identity.
Defaults that changed: since January 2023 every new object is encrypted with SSE-S3, so a bucket without an explicit encryption configuration is still encrypted. Since April 2023 new buckets have Block Public Access on and ACLs disabled (Object Ownership “bucket owner enforced”). A presigned URL stops working when the credentials that signed it expire, so a URL signed with SSO or role credentials can die before its --expires-in (maximum 7 days).
aws s3api create-bucket --bucket my-bucket --region ap-southeast-2 --create-bucket-configuration LocationConstraint=ap-southeast-2 # omit the constraint only in us-east-1
aws s3api put-bucket-versioning --bucket my-bucket --versioning-configuration Status=Enabled
aws s3api put-bucket-encryption --bucket my-bucket --server-side-encryption-configuration \
'{"Rules":[{"ApplyServerSideEncryptionByDefault":{"SSEAlgorithm":"aws:kms","KMSMasterKeyID":"alias/my-app"},"BucketKeyEnabled":true}]}'
aws s3api put-bucket-lifecycle-configuration --bucket my-bucket --lifecycle-configuration file://lifecycle.json
aws s3api get-bucket-lifecycle-configuration --bucket my-bucket
aws s3api put-bucket-policy --bucket my-bucket --policy file://policy.json
aws s3api get-bucket-location --bucket my-bucket
aws s3 rm s3://my-bucket/tmp/ --recursive --exclude '*' --include '*.log' # include/exclude are evaluated in order
aws s3api list-multipart-uploads --bucket my-bucket --query 'Uploads[].[Key,UploadId,Initiated]' --output text # abandoned uploads still bill
aws s3api abort-multipart-upload --bucket my-bucket --key key --upload-id "$UPLOAD_ID"
aws s3api restore-object --bucket my-bucket --key archive.tar --restore-request 'Days=7,GlacierJobParameters={Tier=Bulk}'{
"Rules": [
{ "ID": "logs", "Filter": { "Prefix": "logs/" }, "Status": "Enabled",
"Transitions": [{ "Days": 30, "StorageClass": "STANDARD_IA" }, { "Days": 90, "StorageClass": "GLACIER_IR" }],
"Expiration": { "Days": 365 },
"NoncurrentVersionExpiration": { "NoncurrentDays": 30 },
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } }
]
}Every versioned bucket needs NoncurrentVersionExpiration and every bucket needs AbortIncompleteMultipartUpload, or storage grows invisibly: aws s3 ls shows neither old versions nor incomplete parts. A bucket policy that denies s3:* without aws:SecureTransport enforces TLS; one that denies s3:PutObject unless s3:x-amz-server-side-encryption equals aws:kms enforces the key. Bucket policies are the resource side of IAM evaluation and can grant cross-account access on their own, so review them with get-bucket-policy during any access audit.
IAM policy evaluation#
An explicit Deny in any applicable policy wins. Otherwise an Allow must exist in an identity-based or resource-based policy, and every guardrail in play must also allow the action: service control policies (SCPs) and resource control policies (RCPs) from AWS Organizations, permission boundaries and session policies. “Denied with a policy that clearly allows it” is nearly always a guardrail. See policy evaluation logic.
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::123456789012:role/Deploy \
--action-names s3:PutObject --resource-arns 'arn:aws:s3:::my-bucket/key'
aws sts decode-authorization-message --encoded-message "$MSG" --query DecodedMessage --output text | jq
aws iam list-attached-role-policies --role-name Deploy
aws iam list-role-policies --role-name Deploy # inline policies
aws iam get-account-authorization-details > iam-dump.json # everything, for offline reviewsimulate-principal-policy evaluates identity policies, permission boundaries and SCPs. It ignores RCPs, and evaluates a resource-based policy only when passed with --resource-policy, which works for IAM users but not roles. decode-authorization-message needs sts:DecodeAuthorizationMessage.
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:PutObject"],
"Resource": "arn:aws:s3:::my-bucket/prod/*",
"Condition": { "StringEquals": { "aws:PrincipalTag/Team": "platform" } }
}]
}Policy patterns#
A role has two policies: the trust policy (who may assume it) and permission policies (what it may do). Most cross-account and CI failures are on the trust side.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "GitHubActionsOIDC",
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" },
"StringLike": { "token.actions.githubusercontent.com:sub": "repo:my-org/my-app:ref:refs/heads/main" }
}
},
{
"Sid": "CrossAccountWithExternalId",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::210987654321:role/Deployer" },
"Action": "sts:AssumeRole",
"Condition": { "StringEquals": { "sts:ExternalId": "my-app-prod" } }
}
]
}Permission patterns worth copying rather than inventing:
{
"Version": "2012-10-17",
"Statement": [
{ "Sid": "ListOnlyMyPrefix", "Effect": "Allow", "Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::my-bucket",
"Condition": { "StringLike": { "s3:prefix": ["team/${aws:PrincipalTag/Team}/*"] } } },
{ "Sid": "ObjectsInMyPrefix", "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
"Resource": "arn:aws:s3:::my-bucket/team/${aws:PrincipalTag/Team}/*" },
{ "Sid": "TagBasedEC2Control", "Effect": "Allow", "Action": ["ec2:StartInstances", "ec2:StopInstances"],
"Resource": "arn:aws:ec2:*:123456789012:instance/*",
"Condition": { "StringEquals": { "aws:ResourceTag/Owner": "${aws:PrincipalTag/Team}" } } },
{ "Sid": "RequireMFAForDeletes", "Effect": "Deny", "Action": ["s3:DeleteBucket", "rds:DeleteDBInstance"],
"Resource": "*", "Condition": { "BoolIfExists": { "aws:MultiFactorAuthPresent": "false" } } },
{ "Sid": "RegionLock", "Effect": "Deny", "NotAction": ["iam:*", "sts:*", "organizations:*", "support:*", "cloudfront:*", "route53:*"],
"Resource": "*", "Condition": { "StringNotEquals": { "aws:RequestedRegion": ["ap-southeast-2", "us-east-1"] } } },
{ "Sid": "OnlyFromVPC", "Effect": "Deny", "Action": "s3:*", "Resource": "arn:aws:s3:::my-bucket/*",
"Condition": { "StringNotEquals": { "aws:SourceVpce": "vpce-0abc123" } } }
]
}s3:ListBucket targets the bucket ARN and object actions target bucket/*; putting both actions on both resources is the classic mistake that makes a policy either fail or over-grant. NotAction with Deny is how region locks and SCPs exclude global services. Policy variables such as ${aws:PrincipalTag/Team} make one policy serve every team (attribute-based access control) instead of one policy per team. Use aws iam create-policy-version --set-as-default to update a managed policy, and aws iam get-policy-version to read the current document, since get-policy returns only metadata. IAM Access Analyzer’s aws accessanalyzer validate-policy --policy-document file://p.json --policy-type IDENTITY_POLICY catches syntax errors and over-broad grants before you attach anything. Permission boundaries cap what a role may grant to roles it creates, which is how a CI role can create IAM roles for applications without being able to grant itself administrator access.
CloudWatch Logs and CloudTrail#
aws logs tail /aws/lambda/my-fn --follow --since 10m --format short
aws logs tail /aws/eks/prod/cluster --filter-pattern 'ERROR'
aws logs start-query --log-group-name /aws/lambda/my-fn \
--start-time "$(date -d '1 hour ago' +%s)" --end-time "$(date +%s)" \
--query-string 'fields @timestamp, @message | filter @message like /ERROR/ | sort @timestamp desc | limit 20'
aws logs get-query-results --query-id <query-id> # repeat until status is Complete
aws cloudtrail lookup-events \
--lookup-attributes AttributeKey=EventName,AttributeValue=TerminateInstances \
--query 'Events[].[EventTime,Username,Resources[0].ResourceName]' --output textLogs Insights queries are billed by data scanned; narrow the time range. lookup-events searches the 90-day event history of management events in the current region only. Data events, such as S3 object reads, are recorded only by a trail or event data store configured for them.
Logs Insights recipes#
Insights queries are pipelines. fields, filter, parse, stats, sort and limit cover nearly everything, and @message, @timestamp, @logStream and @log exist in every group. JSON log lines have their fields auto-discovered, so filter level = "error" works without parsing.
-- Error count per 5 minutes
filter @message like /(?i)error/
| stats count() as errors by bin(5m)
-- Top 10 status codes and paths from JSON-structured access logs
filter ispresent(status)
| stats count() as n by status, path
| sort n desc
| limit 10
-- Parse an unstructured line into fields, then aggregate
parse @message /duration=(?<ms>\d+)ms path=(?<path>\S+)/
| stats avg(ms) as avg_ms, pct(ms, 99) as p99_ms, max(ms) as max_ms by path
| sort p99_ms desc
-- Lambda: cold starts and memory headroom from REPORT lines
filter @type = "REPORT"
| stats count() as invocations, sum(strcontains(@message, "Init Duration")) as cold_starts,
max(@maxMemoryUsed / 1000000) as max_mb, avg(@duration) as avg_ms by bin(1h)
-- EKS control plane audit: who deleted what
fields @timestamp, user.username, verb, objectRef.namespace, objectRef.resource, objectRef.name
| filter verb = "delete" and objectRef.resource not in ["events", "leases"]
| sort @timestamp desc
-- VPC Flow Logs: rejected connections by destination port
filter action = "REJECT"
| stats count() as rejects by dstPort, dstAddr
| sort rejects desc
| limit 20
-- Which log streams are noisiest (find the chatty pod)
stats count() as lines, sum(strlen(@message)) as bytes by @logStream
| sort bytes desc
| limit 10# Run a query and wait for it, one command
QID=$(aws logs start-query --log-group-names /aws/eks/prod/cluster /aws/lambda/my-fn \
--start-time "$(date -d '2 hours ago' +%s)" --end-time "$(date +%s)" \
--query-string 'filter @message like /ERROR/ | stats count() by bin(10m)' --query queryId --output text)
until [ "$(aws logs get-query-results --query-id "$QID" --query status --output text)" = Complete ]; do sleep 2; done
aws logs get-query-results --query-id "$QID" --query 'results[].[ [0].value, [1].value ]' --output text
# Cheaper than Insights for a known pattern in a short window
aws logs filter-log-events --log-group-name /aws/lambda/my-fn --start-time "$(( $(date -d '30 min ago' +%s) * 1000 ))" \
--filter-pattern '{ $.level = "error" }' --query 'events[].message' --output text
# Retention and size per log group: unset retention is the usual cost leak
aws logs describe-log-groups --query 'logGroups[].[logGroupName,retentionInDays,storedBytes]' --output text | sort -k3 -rn | head
aws logs put-retention-policy --log-group-name /aws/lambda/my-fn --retention-in-days 30--start-time and --end-time are seconds for start-query and milliseconds for filter-log-events; getting that wrong returns nothing rather than an error. Saved queries live in aws logs describe-query-definitions, and stats ... by bin() output is the same shape as a Prometheus range query if you need to compare with Prometheus.
RDS, Lambda and EKS#
aws rds describe-db-instances \
--query 'DBInstances[].[DBInstanceIdentifier,DBInstanceStatus,Endpoint.Address]' --output table
aws rds create-db-snapshot --db-instance-identifier prod --db-snapshot-identifier "pre-change-$(date +%F)"
aws rds describe-events --source-identifier prod --source-type db-instance --duration 1440
aws lambda invoke --function-name my-fn --payload '{"k":"v"}' --cli-binary-format raw-in-base64-out out.json
aws lambda get-function-configuration --function-name my-fn --query '[Timeout,MemorySize,LastUpdateStatus]'
aws eks update-kubeconfig --name prod --region ap-southeast-2 # writes a context to ~/.kube/config
aws eks describe-cluster --name prod --query 'cluster.[version,status,endpoint]'
aws eks list-nodegroups --cluster-name prod
aws eks list-access-entries --cluster-name prod # IAM principals mapped into the clusterTake a manual snapshot before any risky change to a database you cannot rebuild; automated snapshots are deleted with the instance unless retained. Without --cli-binary-format raw-in-base64-out, CLI v2 expects --payload to be base64 and rejects plain JSON. For cluster work after update-kubeconfig, see Kubernetes.
EKS operations from the CLI#
Cluster access is IAM first: update-kubeconfig writes an exec entry that runs aws eks get-token, and the resulting identity must appear in the cluster’s access entries (the replacement for the aws-auth ConfigMap, authentication mode API or API_AND_CONFIG_MAP). Workload identity comes from Pod Identity associations or the older IRSA (an OIDC provider plus a role trust policy).
aws eks describe-cluster --name prod --query 'cluster.accessConfig.authenticationMode' --output text
aws eks create-access-entry --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/PlatformEngineer --type STANDARD
aws eks associate-access-policy --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/PlatformEngineer \
--policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy --access-scope type=cluster
aws eks associate-access-policy --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/AppTeam \
--policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSEditPolicy --access-scope type=namespace,namespaces=my-namespace
aws eks list-associated-access-policies --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/AppTeam
aws eks create-pod-identity-association --cluster-name prod --namespace my-namespace --service-account my-app \
--role-arn arn:aws:iam::123456789012:role/my-app # role trusts pods.eks.amazonaws.com
aws eks list-pod-identity-associations --cluster-name prod --namespace my-namespace
aws eks describe-nodegroup --cluster-name prod --nodegroup-name workers --query 'nodegroup.[status,scalingConfig,releaseVersion,health.issues]'
aws eks update-nodegroup-version --cluster-name prod --nodegroup-name workers # rolling AMI update to the cluster version
aws eks update-nodegroup-config --cluster-name prod --nodegroup-name workers --scaling-config minSize=3,maxSize=12,desiredSize=6
aws eks list-addons --cluster-name prod
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.34 --query 'addons[0].addonVersions[0].addonVersion' --output text
aws eks update-addon --cluster-name prod --addon-name vpc-cni --addon-version v1.20.0-eksbuild.1 --resolve-conflicts PRESERVE
aws eks describe-update --name prod --update-id "$UPDATE_ID" # progress of a cluster or nodegroup update
aws eks list-insights --cluster-name prod --query 'insights[?insightStatus.status!=`PASSING`].[name,insightStatus.status]' --output text # upgrade blockersNode problems that look like Kubernetes problems: a nodegroup health.issues entry of Ec2SubnetInvalidConfiguration or InsufficientFreeAddresses means the subnet is out of IPs, and pods stuck ContainerCreating with failed to assign an IP address is the same problem at the pod level (VPC CNI takes one address per pod unless prefix delegation is on). Check aws ec2 describe-subnets --subnet-ids <id> --query 'Subnets[].AvailableIpAddressCount'. A kubectl that fails with the server has asked for the client to provide credentials means get-token returned an identity with no access entry; aws sts get-caller-identity shows which one.
Cost checks#
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-09-01 \
--granularity MONTHLY --metrics UnblendedCost --group-by Type=DIMENSION,Key=SERVICE \
--query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn | head
aws ec2 describe-volumes --filters Name=status,Values=available --query 'Volumes[].[VolumeId,Size]' --output text
aws ec2 describe-addresses --query 'Addresses[?AssociationId==null].PublicIp' --output textEach Cost Explorer API request is charged (USD 0.01 at review time). Unattached EBS volumes and unassociated Elastic IPs bill continuously; public IPv4 addresses are charged hourly whether or not they are attached.
# Daily spend for the last two weeks, to spot the day something changed
aws ce get-cost-and-usage --time-period Start="$(date -d '14 days ago' +%F)",End="$(date +%F)" --granularity DAILY --metrics UnblendedCost \
--query 'ResultsByTime[].[TimePeriod.Start,Metrics.UnblendedCost.Amount]' --output text
# This month by a cost allocation tag (tag must be activated in Billing first)
aws ce get-cost-and-usage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --granularity MONTHLY --metrics UnblendedCost \
--group-by Type=TAG,Key=Team --query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn
# Usage types inside one service: which EC2 line items are growing
aws ce get-cost-and-usage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --granularity MONTHLY --metrics UnblendedCost \
--filter '{"Dimensions":{"Key":"SERVICE","Values":["Amazon Elastic Compute Cloud - Compute"]}}' \
--group-by Type=DIMENSION,Key=USAGE_TYPE --query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn | head
# Forecast to month end
aws ce get-cost-forecast --time-period Start="$(date +%F)",End="$(date -d "$(date +%Y-%m-01) +1 month" +%F)" --metric UNBLENDED_COST --granularity MONTHLY --query Total.Amount --output text
# Rightsizing and idle recommendations
aws ce get-rightsizing-recommendation --service AmazonEC2 --query 'RightsizingRecommendations[].[CurrentInstance.ResourceId,RightsizingType,CurrentInstance.MonthlyCost]' --output text
aws compute-optimizer get-ec2-instance-recommendations --query 'instanceRecommendations[?finding==`Overprovisioned`].[instanceArn,recommendationOptions[0].instanceType]' --output text
# Savings Plans and RI coverage this month
aws ce get-savings-plans-utilization --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --query 'Total.Utilization.UtilizationPercentage' --output text
aws ce get-reservation-coverage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --query 'Total.CoverageHours.CoverageHoursPercentage' --output text
# Budgets and their current state
aws budgets describe-budgets --account-id "$(aws sts get-caller-identity --query Account --output text)" --query 'Budgets[].[BudgetName,BudgetLimit.Amount,CalculatedSpend.ActualSpend.Amount]' --output tableCost Explorer data lags by up to 24 hours and End is exclusive. Snapshots (aws ec2 describe-snapshots --owner-ids self) and old AMIs are the other quiet bill; Data Lifecycle Manager or a tag-driven cleanup script keeps them bounded.
Troubleshooting#
| Symptom | Cause | Check |
|---|---|---|
Unable to locate credentials | No source in the chain answered | aws configure list; aws sso login |
The SSO session associated with this profile has expired | Cached SSO token expired | aws sso login --profile <name> |
ExpiredToken | Exported temporary credentials outlived their session | Unset AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN |
| Right command, wrong account | Environment variables override the profile | aws sts get-caller-identity; env | grep -c '^AWS_' |
AccessDenied with an encoded message | Missing allow, or a guardrail denies | aws sts decode-authorization-message |
AccessDenied although the policy allows | SCP, RCP, permission boundary, resource policy or KMS key policy | Simulate; check the key policy for encrypted resources |
| Empty result, no error | Wrong region; most resources are regional | aws configure get region, --region |
Could not connect to the endpoint URL | Invalid region name, proxy or no network path | --debug 2>&1 | grep -i endpoint |
ThrottlingException or Rate exceeded | API rate limit | AWS_RETRY_MODE=adaptive, AWS_MAX_ATTEMPTS=10 |
SignatureDoesNotMatch or RequestTimeTooSkewed | Local clock drift | timedatectl; sync NTP |
ssm start-session fails with TargetNotConnected | SSM agent offline, no instance role, or no route to SSM endpoints | aws ssm describe-instance-information |
SessionManagerPlugin is not found | Plugin not installed on the workstation | Install session-manager-plugin; session-manager-plugin --version |
kubectl says the server has asked for the client to provide credentials | Identity from aws eks get-token has no access entry | aws sts get-caller-identity; aws eks list-access-entries --cluster-name <name> |
Pods ContainerCreating with failed to assign an IP address | Subnet out of addresses (VPC CNI) | aws ec2 describe-subnets --query 'Subnets[].[SubnetId,AvailableIpAddressCount]' |
Not authorized to perform sts:AssumeRoleWithWebIdentity | Trust policy sub/aud condition does not match the token | Decode the token claims; compare with the trust policy Condition |
AccessDenied on s3:ListBucket while GetObject works | Policy grants object actions on bucket/* but not ListBucket on the bucket ARN | aws iam simulate-principal-policy --action-names s3:ListBucket --resource-arns arn:aws:s3:::my-bucket |
KMS.AccessDeniedException reading an encrypted object or parameter | Key policy does not grant the principal kms:Decrypt | aws kms get-key-policy --key-id <id> --policy-name default |
| Logs Insights returns no rows | Time units wrong (seconds vs milliseconds), or the field is not auto-discovered | --start-time in seconds for start-query; use parse for unstructured lines |
InvalidParameterValue on --tag-specifications or --ip-permissions | Shorthand syntax mismatch | Use --generate-cli-skeleton and pass --cli-input-json file:// |
| Cost Explorer shows zero for a tag | Tag not activated as a cost allocation tag, or activated after the spend | Billing console tag activation; wait 24 h |
An error occurred (ValidationException) ... nodegroup update stuck | Pod disruption budgets block node drain | kubectl get pdb -A; aws eks describe-update shows the error |
--debug prints every request, the endpoint and the credential provider that answered. It can include signed headers, so do not paste it into tickets unedited.
Oneliners#
# Confirm account and region before anything destructive
aws sts get-caller-identity --query '[Account,Arn]' --output text; aws configure get region
# Running instances in every enabled region
for r in $(aws ec2 describe-regions --query 'Regions[].RegionName' --output text); do aws ec2 describe-instances --region "$r" --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,InstanceType]' --output text | sed "s/^/$r /"; done
# Instance ID from a Name tag
aws ec2 describe-instances --filters 'Name=tag:Name,Values=web-01' --query 'Reservations[0].Instances[0].InstanceId' --output text
# Security groups with ingress open to the internet (IPv4 or IPv6)
aws ec2 describe-security-groups --query 'SecurityGroups[?IpPermissions[?IpRanges[?CidrIp==`0.0.0.0/0`] || Ipv6Ranges[?CidrIpv6==`::/0`]]].[GroupId,GroupName]' --output text
# Buckets without a full public access block
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do aws s3api get-public-access-block --bucket "$b" --query 'PublicAccessBlockConfiguration.[BlockPublicAcls,IgnorePublicAcls,BlockPublicPolicy,RestrictPublicBuckets]' --output text 2>/dev/null | grep -q False && echo "$b"; done
# Largest objects in a bucket
aws s3api list-objects-v2 --bucket my-bucket --query 'sort_by(Contents,&Size)[-10:].[Size,Key]' --output text
# Access keys older than 90 days
aws iam list-users --query 'Users[].UserName' --output text | tr '\t' '\n' | while read -r u; do aws iam list-access-keys --user-name "$u" --query "AccessKeyMetadata[?CreateDate<='$(date -d '90 days ago' +%Y-%m-%d)'].[UserName,AccessKeyId,CreateDate]" --output text; done
# Who acted on an instance
aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,AttributeValue=i-0abc123 --query 'Events[].[EventTime,EventName,Username]' --output text
# Copy between buckets without downloading
aws s3 sync s3://src-bucket/prefix s3://dst-bucket/prefix --source-region us-east-1 --region ap-southeast-2
# Count of each value of a tag
aws ec2 describe-tags --filters Name=key,Values=Environment --query 'Tags[].Value' --output text | tr '\t' '\n' | sort | uniq -c
# Instances without an Owner tag
aws ec2 describe-instances --query 'Reservations[].Instances[?!not_null(Tags[?Key==`Owner`].Value|[0])].[InstanceId,LaunchTime]' --output text
# Instances still allowing IMDSv1
aws ec2 describe-instances --filters Name=metadata-options.http-tokens,Values=optional --query 'Reservations[].Instances[].InstanceId' --output text
# Private IP to instance name map for the whole account
aws ec2 describe-instances --query 'Reservations[].Instances[].[PrivateIpAddress, Tags[?Key==`Name`].Value|[0]]' --output text | sort -t. -k1,1n -k2,2n -k3,3n -k4,4n
# Unencrypted EBS volumes
aws ec2 describe-volumes --filters Name=encrypted,Values=false --query 'Volumes[].[VolumeId,Size,Attachments[0].InstanceId]' --output text
# Snapshots older than a year, with size
aws ec2 describe-snapshots --owner-ids self --query "Snapshots[?StartTime<='$(date -d '1 year ago' +%Y-%m-%d)'].[SnapshotId,VolumeSize,StartTime,Description]" --output text
# AMIs you own that no instance uses
comm -23 <(aws ec2 describe-images --owners self --query 'Images[].ImageId' --output text | tr '\t' '\n' | sort) <(aws ec2 describe-instances --query 'Reservations[].Instances[].ImageId' --output text | tr '\t' '\n' | sort -u)
# Security groups attached to nothing
comm -23 <(aws ec2 describe-security-groups --query 'SecurityGroups[?GroupName!=`default`].GroupId' --output text | tr '\t' '\n' | sort) <(aws ec2 describe-network-interfaces --query 'NetworkInterfaces[].Groups[].GroupId' --output text | tr '\t' '\n' | sort -u)
# Subnets running low on addresses
aws ec2 describe-subnets --query 'Subnets[?AvailableIpAddressCount<`20`].[SubnetId,CidrBlock,AvailableIpAddressCount,Tags[?Key==`Name`].Value|[0]]' --output text
# Total size of a bucket prefix without listing every object
aws s3 ls s3://my-bucket/logs/ --recursive --summarize | tail -2
# Bucket sizes from CloudWatch (free; updated daily), in GiB
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do printf '%s\t' "$b"; aws cloudwatch get-metric-statistics --namespace AWS/S3 --metric-name BucketSizeBytes --dimensions Name=BucketName,Value="$b" Name=StorageType,Value=StandardStorage --start-time "$(date -d '2 days ago' -u +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" --period 86400 --statistics Average --query 'Datapoints[-1].Average' --output text | awk '{printf "%.1f\n", $1/1073741824}'; done
# Buckets without versioning
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do [ "$(aws s3api get-bucket-versioning --bucket "$b" --query Status --output text)" = Enabled ] || echo "$b"; done
# Buckets without a lifecycle configuration
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do aws s3api get-bucket-lifecycle-configuration --bucket "$b" >/dev/null 2>&1 || echo "$b"; done
# Delete every version and delete marker under a prefix, 1000 at a time (irreversible; run the list-object-versions half alone first to review)
aws s3api list-object-versions --bucket my-bucket --prefix tmp/ --query '{Objects: [Versions[].{Key:Key,VersionId:VersionId}, DeleteMarkers[].{Key:Key,VersionId:VersionId}][] | [0:1000], Quiet: `true`}' --output json | aws s3api delete-objects --bucket my-bucket --delete file:///dev/stdin
# Roles nobody has used in 90 days
aws iam list-roles --query "Roles[?RoleLastUsed.LastUsedDate<='$(date -d '90 days ago' +%Y-%m-%d)' || !RoleLastUsed.LastUsedDate].[RoleName,RoleLastUsed.LastUsedDate]" --output text
# Users with console passwords but no MFA
aws iam generate-credential-report >/dev/null; sleep 5; aws iam get-credential-report --query Content --output text | base64 -d | awk -F, 'NR>1 && $4=="true" && $8=="false" {print $1}'
# Policies that grant Action "*" on Resource "*"
for arn in $(aws iam list-policies --scope Local --query 'Policies[].Arn' --output text); do v=$(aws iam get-policy --policy-arn "$arn" --query Policy.DefaultVersionId --output text); aws iam get-policy-version --policy-arn "$arn" --version-id "$v" --query 'PolicyVersion.Document.Statement[?Effect==`Allow` && (Action==`*` || contains(Action, `*`)) && Resource==`*`]' --output text | grep -q . && echo "$arn"; done
# Who is allowed to assume a role
aws iam get-role --role-name Deploy --query 'Role.AssumeRolePolicyDocument.Statement[].Principal' --output json
# Console sign-ins in the last day, by user
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin --start-time "$(date -d '1 day ago' -u +%FT%TZ)" --query 'Events[].Username' --output text | tr '\t' '\n' | sort | uniq -c
# Every API call by one principal today
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=alice --start-time "$(date -u +%FT00:00:00Z)" --query 'Events[].[EventTime,EventSource,EventName]' --output text
# Log groups without retention, with stored size in GiB
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].[logGroupName,storedBytes]' --output text | awk '{printf "%s\t%.2f\n", $1, $2/1073741824}' | sort -k2 -rn
# Set 30-day retention on every log group that has none
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].logGroupName' --output text | tr '\t' '\n' | xargs -r -I{} aws logs put-retention-policy --log-group-name {} --retention-in-days 30
# Lambda functions on a deprecated runtime
aws lambda list-functions --query 'Functions[?starts_with(Runtime, `python3.8`) || starts_with(Runtime, `nodejs16`)].[FunctionName,Runtime]' --output text
# Lambda error rate in the last hour for one function
aws cloudwatch get-metric-statistics --namespace AWS/Lambda --metric-name Errors --dimensions Name=FunctionName,Value=my-fn --start-time "$(date -d '1 hour ago' -u +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" --period 3600 --statistics Sum --query 'Datapoints[0].Sum' --output text
# RDS instances that are publicly accessible or unencrypted
aws rds describe-db-instances --query 'DBInstances[?PubliclyAccessible || !StorageEncrypted].[DBInstanceIdentifier,PubliclyAccessible,StorageEncrypted]' --output text
# Wait for an RDS snapshot then print its ARN
aws rds wait db-snapshot-completed --db-snapshot-identifier pre-change && aws rds describe-db-snapshots --db-snapshot-identifier pre-change --query 'DBSnapshots[0].DBSnapshotArn' --output text
# EKS: cluster version and every nodegroup's release version (mismatch means a pending upgrade)
aws eks describe-cluster --name prod --query cluster.version --output text; for ng in $(aws eks list-nodegroups --cluster-name prod --query 'nodegroups[]' --output text); do aws eks describe-nodegroup --cluster-name prod --nodegroup-name "$ng" --query 'nodegroup.[nodegroupName,version,releaseVersion,status]' --output text; done
# EKS: access entries and their policies
for p in $(aws eks list-access-entries --cluster-name prod --query 'accessEntries[]' --output text); do echo "$p"; aws eks list-associated-access-policies --cluster-name prod --principal-arn "$p" --query 'associatedAccessPolicies[].[policyArn,accessScope.type]' --output text | sed 's/^/ /'; done
# SSM: managed instances whose agent has not checked in for a day
aws ssm describe-instance-information --query "InstanceInformationList[?LastPingDateTime<='$(date -d '1 day ago' -u +%FT%TZ)'].[InstanceId,PingStatus,LastPingDateTime]" --output text
# SSM: run one command on every instance with a tag and print output per host
CMD=$(aws ssm send-command --document-name AWS-RunShellScript --targets Key=tag:Environment,Values=prod --parameters 'commands=["uptime"]' --query Command.CommandId --output text); sleep 10; aws ssm list-command-invocations --command-id "$CMD" --details --query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].Output]' --output text
# Parameter Store: every parameter under a path with its last modification
aws ssm get-parameters-by-path --path /my-app/ --recursive --query 'Parameters[].[Name,Type,Version,LastModifiedDate]' --output table
# Service quotas you are close to (EC2 vCPU example)
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A --query 'Quota.Value' --output text; aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].CpuOptions.[CoreCount,ThreadsPerCore]' --output text | awk '{s+=$1*$2} END {print s " vCPUs in use"}'
# Regions where you have any EC2 resources at all
for r in $(aws ec2 describe-regions --query 'Regions[].RegionName' --output text); do n=$(aws ec2 describe-instances --region "$r" --query 'length(Reservations[].Instances[])' --output text); [ "$n" = 0 ] || echo "$r $n"; done
# Retry-hardened settings for a scripted session
export AWS_RETRY_MODE=adaptive AWS_MAX_ATTEMPTS=10 AWS_PAGER=""Scripts#
Produce a one-page security posture report for an account: public buckets, open security groups, IMDSv1 instances, stale keys and roles, and log groups without retention.
#!/usr/bin/env bash
# account-audit.sh [profile]: read-only posture checks, one section per finding class
set -euo pipefail
export AWS_PAGER="" AWS_RETRY_MODE=adaptive AWS_MAX_ATTEMPTS=10
[ $# -eq 0 ] || export AWS_PROFILE=$1
acct=$(aws sts get-caller-identity --query '[Account,Arn]' --output text)
printf 'Account/identity: %s\nRegion: %s\n\n' "$acct" "$(aws configure get region || echo unset)"
section() { printf '== %s ==\n' "$1"; }
section "Buckets without full public access block"
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do
aws s3api get-public-access-block --bucket "$b" --query 'PublicAccessBlockConfiguration.[BlockPublicAcls,IgnorePublicAcls,BlockPublicPolicy,RestrictPublicBuckets]' --output text 2>/dev/null | grep -q False && echo "$b"
done || true
section "Security groups open to the world on ports other than 80/443"
aws ec2 describe-security-groups --query 'SecurityGroups[?IpPermissions[?(IpRanges[?CidrIp==`0.0.0.0/0`] || Ipv6Ranges[?CidrIpv6==`::/0`]) && !(FromPort==`80` || FromPort==`443`)]].[GroupId,GroupName]' --output text
section "Instances allowing IMDSv1"
aws ec2 describe-instances --filters Name=metadata-options.http-tokens,Values=optional Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,Tags[?Key==`Name`].Value|[0]]' --output text
section "Unencrypted volumes"
aws ec2 describe-volumes --filters Name=encrypted,Values=false --query 'Volumes[].[VolumeId,Size]' --output text
section "Access keys older than 90 days"
cutoff=$(date -d '90 days ago' +%Y-%m-%d)
for u in $(aws iam list-users --query 'Users[].UserName' --output text); do
aws iam list-access-keys --user-name "$u" --query "AccessKeyMetadata[?CreateDate<='$cutoff' && Status=='Active'].[UserName,AccessKeyId,CreateDate]" --output text
done
section "Roles unused for 90 days"
aws iam list-roles --query "Roles[?!starts_with(Path, '/aws-service-role/') && (RoleLastUsed.LastUsedDate<='$cutoff' || !RoleLastUsed.LastUsedDate)].RoleName" --output text | tr '\t' '\n'
section "Log groups without retention (GiB stored)"
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].[logGroupName,storedBytes]' --output text | awk '{printf "%s\t%.2f\n", $1, $2/1073741824}' | sort -k2 -rn | head -20
section "CloudTrail"
aws cloudtrail describe-trails --query 'trailList[].[Name,IsMultiRegionTrail,LogFileValidationEnabled]' --output textTag-driven cleanup of EBS snapshots: delete snapshots older than a retention period unless tagged Retain=true or still referenced by an AMI, with a dry run by default.
#!/usr/bin/env bash
# snapshot-prune.sh DAYS [--apply]: delete own EBS snapshots older than DAYS unless retained or used by an AMI
set -euo pipefail
days=$1; apply=${2:-}
export AWS_PAGER=""
cutoff=$(date -d "$days days ago" +%Y-%m-%dT%H:%M:%SZ)
in_use=$(aws ec2 describe-images --owners self --query 'Images[].BlockDeviceMappings[].Ebs.SnapshotId' --output text | tr '\t' '\n' | sort -u)
count=0; bytes=0
while read -r id size start retain; do
[ -n "$id" ] || continue
[ "$retain" = true ] && continue
grep -qx "$id" <<<"$in_use" && continue
count=$((count + 1)); bytes=$((bytes + size))
if [ "$apply" = --apply ]; then
aws ec2 delete-snapshot --snapshot-id "$id" && echo "deleted $id ($size GiB, $start)"
else
echo "would delete $id ($size GiB, $start)"
fi
done < <(aws ec2 describe-snapshots --owner-ids self --query "Snapshots[?StartTime<='$cutoff'].[SnapshotId,VolumeSize,StartTime,Tags[?Key=='Retain'].Value|[0]]" --output text)
printf '%d snapshots, %d GiB %s\n' "$count" "$bytes" "$([ "$apply" = --apply ] && echo deleted || echo 'to delete (pass --apply)')"Poll a Logs Insights query to completion and print the result as a table, for use in runbooks and cron jobs.
#!/usr/bin/env python3
"""Run a CloudWatch Logs Insights query and print a table.
Usage: insights.py LOG_GROUP MINUTES 'QUERY' [more log groups...]
Example: insights.py /aws/lambda/my-fn 60 'filter @message like /ERROR/ | stats count() by bin(5m)'
"""
import sys
import time
import boto3
group, minutes, query, *more = sys.argv[1:]
logs = boto3.client("logs")
now = int(time.time())
start = logs.start_query(logGroupNames=[group, *more], startTime=now - int(minutes) * 60, endTime=now, queryString=query)
qid = start["queryId"]
while True:
res = logs.get_query_results(queryId=qid)
if res["status"] in ("Complete", "Failed", "Cancelled", "Timeout"):
break
time.sleep(1)
if res["status"] != "Complete":
sys.exit(f"query {res['status']}")
rows = [{c["field"]: c["value"] for c in r if c["field"] != "@ptr"} for r in res["results"]]
if not rows:
sys.exit("no results")
cols = list(rows[0])
width = {c: max(len(c), *(len(r.get(c, "")) for r in rows)) for c in cols}
print(" ".join(c.ljust(width[c]) for c in cols))
for r in rows:
print(" ".join(r.get(c, "").ljust(width[c]) for c in cols))
stats = res["statistics"]
print(f"\n{stats['recordsMatched']:.0f} matched, {stats['bytesScanned'] / 1e9:.2f} GB scanned", file=sys.stderr)