# AWS

> Resolve AWS CLI v2 credentials and profiles, filter output, and run and debug common EC2, S3, IAM, logs, RDS, Lambda and EKS operations.

Canonical: https://www.wiki.jodisand.me/aws/
Reviewed: 2026-09-24
Related: [Terraform](https://www.wiki.jodisand.me/terraform/index.md), [Kubernetes](https://www.wiki.jodisand.me/kubernetes/index.md), [jq](https://www.wiki.jodisand.me/jq/index.md), [SSH](https://www.wiki.jodisand.me/ssh/index.md)


## Cheatsheet

| Task | Command |
| --- | --- |
| Which identity and account | `aws sts get-caller-identity` |
| Which profile, region and credential source | `aws configure list` |
| Log in with IAM Identity Center | `aws sso login --profile prod` |
| Log in with console credentials (CLI 2.32+) | `aws login --profile dev` |
| Assume a role | `aws sts assume-role --role-arn <arn> --role-session-name alice` |
| Filter output client-side | `--query 'Reservations[].Instances[].InstanceId' --output text` |
| Shell on an instance, no SSH or bastion | `aws ssm start-session --target i-0abc123` |
| Tail a log group | `aws logs tail /aws/lambda/my-fn --follow` |
| Recent management API calls | `aws cloudtrail lookup-events --max-results 20` |
| Preview a sync that deletes | `aws s3 sync ./dist s3://my-bucket/site --delete --dryrun` |
| Temporary download URL | `aws s3 presign s3://my-bucket/key --expires-in 3600` |
| Decode an encoded authorisation failure | `aws sts decode-authorization-message --encoded-message "$MSG"` |
| Test a principal's permissions | `aws iam simulate-principal-policy --policy-source-arn <arn> --action-names s3:GetObject` |
| Service quota | `aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A` |
| Block until a state is reached | `aws ec2 wait instance-running --instance-ids i-0abc123` |

Everything here is AWS CLI v2. v1 lacks `aws logs tail`, SSO sessions and `aws login`, and treats binary blobs differently. Check with `aws --version`. See the [AWS CLI v2 reference](https://docs.aws.amazon.com/cli/latest/reference/).

## Credentials and identity

For each setting the CLI takes the first source that provides it, in this order: command-line options (`--profile`, `--region`), environment variables (`AWS_ACCESS_KEY_ID`, `AWS_PROFILE`, `AWS_REGION`), then the profile's configuration: assume role, assume role with web identity, IAM Identity Center (SSO), the `credentials` file, `credential_process`, the `config` file, then container credentials (ECS task role, EKS Pod Identity) and finally EC2 instance profile credentials from instance metadata. See [configuration and credential precedence](https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-authentication.html#cli-chap-authentication-precedence).

A stale `AWS_ACCESS_KEY_ID` exported in a shell therefore beats a correct `--profile` in the config file, and an instance role is used silently when nothing else is configured.

```sh
aws sts get-caller-identity          # the reliable answer to "which account and role is this"
aws configure list                   # each setting with its source (env, config-file, iam-role)
env | grep '^AWS_' | cut -d= -f1     # which AWS variables are set, without printing values
AWS_PROFILE=prod aws s3 ls
```

```text
{
    "UserId": "AROAEXAMPLEID:alice",
    "Account": "123456789012",
    "Arn": "arn:aws:sts::123456789012:assumed-role/PlatformEngineer/alice"
}
```

```ini
# ~/.aws/config
[sso-session corp]
sso_start_url = https://example.awsapps.com/start
sso_region    = ap-southeast-2
sso_registration_scopes = sso:account:access

[profile prod]
sso_session    = corp
sso_account_id = 123456789012
sso_role_name  = PlatformEngineer
region         = ap-southeast-2
output         = json

[profile prod-admin]
source_profile   = prod
role_arn         = arn:aws:iam::123456789012:role/Admin
duration_seconds = 3600
```

`aws sso login` caches a token under `~/.aws/sso/cache`; `aws login` (CLI 2.32.0+) does the same for console sign-in (root, IAM user or federated) under `~/.aws/login/cache`, refreshing for up to 12 hours. `aws configure export-credentials --profile prod --format env` prints temporary credentials for tools that cannot read profiles; the output is secret.

> [!WARNING] Prefer short-lived credentials
> Use IAM Identity Center or `aws login` for people and roles for workloads (instance profiles, ECS task roles, EKS Pod Identity or IRSA). A static access key must be rotated and kept out of repositories, images and CI logs.

## Output, queries and pagination

`--filters` (and parameters such as `--prefix`) are applied by the service before it responds. `--query` is a [JMESPath](https://jmespath.org/) expression applied by the CLI after every page has been downloaded. On a large account, filter server-side first and use `--query` to shape the result.

```sh
aws ec2 describe-instances --filters 'Name=instance-state-name,Values=running' \
  --query 'Reservations[].Instances[].[InstanceId,InstanceType,PrivateIpAddress]' --output text

aws ec2 describe-instances \
  --query 'Reservations[].Instances[].[InstanceId,Tags[?Key==`Name`].Value|[0]]' --output table

aws s3api list-objects-v2 --bucket my-bucket --prefix logs/ \
  --query 'sort_by(Contents, &LastModified)[-5:].[Key,Size]' --output text
```

| Option | Effect |
| --- | --- |
| `--output json\|yaml\|text\|table` | `text` is tab-separated for `awk` and `cut`; `json` with [jq](https://www.wiki.jodisand.me/jq/) for nested data |
| `--no-paginate` | Return only the first page |
| `--page-size n` | Smaller API pages (avoids timeouts); still returns everything |
| `--max-items n` | Stop after n items and print a `NextToken` for `--starting-token` |
| `--no-cli-pager` | Disable the pager (`AWS_PAGER=""` does the same for a session) |
| `--cli-read-timeout`, `--cli-connect-timeout` | Seconds before a request is abandoned; set in scripts |

In `text` output, a `--query` that selects nothing prints `None`. Test for it explicitly in scripts.

### JMESPath patterns

The same handful of constructs cover nearly every `--query`. Literals inside a filter use backticks; a string compared with a literal must be inside backticks or single quotes, and the whole expression is single-quoted for the shell.

```sh
# Projection: a list of fields per item
--query 'Reservations[].Instances[].[InstanceId,State.Name,PrivateIpAddress]'

# Filter with a comparison, then project
--query 'Volumes[?Size > `100`].[VolumeId,Size]'
--query 'Reservations[].Instances[?State.Name==`running`].InstanceId[]'     # trailing [] flattens nested lists

# A tag value: filter the Tags list, take the first match, default when missing
--query 'Reservations[].Instances[].[InstanceId, Tags[?Key==`Name`].Value | [0] || `untagged`]'

# Multi-select hash: name the output keys, then --output table gives labelled columns
--query 'DBInstances[].{id:DBInstanceIdentifier,class:DBInstanceClass,status:DBInstanceStatus,az:AvailabilityZone}'

# Functions: sort, length, contains, starts_with, to_string, join
--query 'sort_by(Functions, &LastModified)[-3:].FunctionName'
--query 'length(Reservations[].Instances[])'
--query 'Buckets[?starts_with(Name, `prod-`)].Name'
--query 'Roles[?contains(RoleName, `Deploy`)].Arn'
--query 'join(`,`, Subnets[].SubnetId)'

# Boolean OR of conditions and negation
--query 'SecurityGroups[?GroupName!=`default` && length(IpPermissions)==`0`].GroupId'

# Pipe to re-shape the result of the left side
--query 'Reservations[].Instances[] | [?Platform!=`windows`] | length(@)'
```

`--query` cannot compare dates or do arithmetic beyond comparisons of numbers; do that in [jq](https://www.wiki.jodisand.me/jq/) with `--output json`. Test an expression against saved output with `aws ec2 describe-instances --output json > ec2.json` and the `jp` CLI, or iterate quickly with `--no-cli-pager --output table`.

## EC2 and Systems Manager

```sh
aws ec2 describe-instances --filters 'Name=tag:Environment,Values=prod' \
  --query 'Reservations[].Instances[].[InstanceId,PrivateIpAddress,State.Name]' --output text
aws ec2 start-instances --instance-ids i-0abc123
aws ec2 stop-instances --instance-ids i-0abc123                     # instance-store data is lost
aws ec2 describe-instance-status --instance-ids i-0abc123            # system and instance status checks
aws ec2 get-console-output --instance-id i-0abc123 --latest --output text | tail -50
aws ec2 create-image --instance-id i-0abc123 --name "backup-$(date +%F)" --no-reboot
aws ssm start-session --target i-0abc123                            # needs the SSM agent and session-manager-plugin
aws ssm start-session --target i-0abc123 \
  --document-name AWS-StartPortForwardingSession --parameters 'portNumber=5432,localPortNumber=15432'
```

`--no-reboot` images a running filesystem, so the image may be inconsistent for databases.

Security groups are stateful: an inbound allow implies the reply traffic. Network ACLs are stateless and need rules in both directions, including the ephemeral port range for replies. That asymmetry explains most "the security group looks correct" cases.

```sh
# Launch from the latest Amazon Linux 2023 AMI via the public SSM parameter, with tags at creation
AMI=$(aws ssm get-parameter --name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 --query Parameter.Value --output text)
aws ec2 run-instances --image-id "$AMI" --instance-type t3.small --subnet-id subnet-0abc123 \
  --security-group-ids sg-0abc123 --iam-instance-profile Name=ssm-managed \
  --metadata-options HttpTokens=required \
  --tag-specifications 'ResourceType=instance,Tags=[{Key=Name,Value=web-01},{Key=Environment,Value=prod}]' \
  --user-data file://cloud-init.yml --query 'Instances[0].InstanceId' --output text

aws ec2 authorize-security-group-ingress --group-id sg-0abc123 --ip-permissions \
  'IpProtocol=tcp,FromPort=443,ToPort=443,IpRanges=[{CidrIp=192.0.2.0/24,Description="office"}]'
aws ec2 describe-security-group-rules --filters Name=group-id,Values=sg-0abc123 --query 'SecurityGroupRules[].[SecurityGroupRuleId,IsEgress,IpProtocol,FromPort,CidrIpv4]' --output table
aws ec2 modify-instance-attribute --instance-id i-0abc123 --instance-type t3.large       # stopped instance only
aws ec2 modify-instance-metadata-options --instance-id i-0abc123 --http-tokens required   # enforce IMDSv2
aws ec2 create-tags --resources i-0abc123 vol-0abc123 --tags Key=Owner,Value=platform
aws ec2 terminate-instances --instance-ids i-0abc123                                     # irreversible; check DisableApiTermination first
aws ec2 describe-instance-attribute --instance-id i-0abc123 --attribute disableApiTermination
```

`HttpTokens=required` forces IMDSv2, which stops SSRF-style credential theft through the metadata service; make it the default on every launch template. `--user-data` is run once by cloud-init on first boot; changes to it on a stopped instance do not re-run unless the instance is rebuilt.

### Systems Manager beyond sessions

SSM also runs commands across a fleet selected by tag, stores configuration in Parameter Store, and reports inventory and patch state, all without opening inbound ports.

```sh
aws ssm describe-instance-information --query 'InstanceInformationList[].[InstanceId,PingStatus,PlatformName,AgentVersion]' --output table
aws ssm send-command --document-name AWS-RunShellScript --targets 'Key=tag:Environment,Values=prod' \
  --parameters 'commands=["dnf -y check-update || true","systemctl is-active nginx"]' \
  --comment "health check" --query Command.CommandId --output text
aws ssm list-command-invocations --command-id "$CMD_ID" --details \
  --query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].Output]' --output text
aws ssm get-parameter --name /my-app/prod/db_url --with-decryption --query Parameter.Value --output text
aws ssm get-parameters-by-path --path /my-app/prod/ --recursive --with-decryption --query 'Parameters[].[Name,Version]' --output table
aws ssm put-parameter --name /my-app/prod/db_url --type SecureString --value "$DB_URL" --overwrite   # value lands in shell history unless read from a variable
aws ssm describe-instance-patch-states --instance-ids i-0abc123 --query 'InstancePatchStates[].[InstanceId,MissingCount,FailedCount,OperationEndTime]' --output text
aws ssm start-session --target i-0abc123 --document-name AWS-StartInteractiveCommand --parameters 'command=["journalctl -u nginx -n 50"]'
```

`send-command` returns immediately; poll `list-command-invocations` or pass `--output-s3-bucket-name` for output longer than the 2,500-character inline limit. Session Manager sessions are logged to CloudWatch or S3 when the `SSM-SessionManagerRunShell` document preferences say so, which is the audit trail SSH never had.

## S3

```sh
aws s3 ls s3://my-bucket/prefix/ --human-readable --summarize
aws s3 sync ./dist s3://my-bucket/site --delete --dryrun            # always preview --delete
aws s3 cp file.tar s3://my-bucket/key --storage-class INTELLIGENT_TIERING
aws s3 presign s3://my-bucket/key --expires-in 3600
aws s3api head-object --bucket my-bucket --key key                  # size, encryption, metadata
aws s3api get-bucket-policy --bucket my-bucket --query Policy --output text | jq
aws s3api list-object-versions --bucket my-bucket --prefix key \
  --query 'Versions[].[Key,VersionId,IsLatest]' --output text
aws s3api get-public-access-block --bucket my-bucket
```

`aws s3` is the high-level interface (sync, recursive copy, automatic multipart). `aws s3api` maps one-to-one to API operations for policies, versions and metadata.

> [!WARNING] `sync --delete` and `rm --recursive` are immediate
> Without versioning there is no undo. Run with `--dryrun` first and check the account with `aws sts get-caller-identity`.

Defaults that changed: since January 2023 every new object is encrypted with SSE-S3, so a bucket without an explicit encryption configuration is still encrypted. Since April 2023 new buckets have Block Public Access on and ACLs disabled (Object Ownership "bucket owner enforced"). A presigned URL stops working when the credentials that signed it expire, so a URL signed with SSO or role credentials can die before its `--expires-in` (maximum 7 days).

```sh
aws s3api create-bucket --bucket my-bucket --region ap-southeast-2 --create-bucket-configuration LocationConstraint=ap-southeast-2   # omit the constraint only in us-east-1
aws s3api put-bucket-versioning --bucket my-bucket --versioning-configuration Status=Enabled
aws s3api put-bucket-encryption --bucket my-bucket --server-side-encryption-configuration \
  '{"Rules":[{"ApplyServerSideEncryptionByDefault":{"SSEAlgorithm":"aws:kms","KMSMasterKeyID":"alias/my-app"},"BucketKeyEnabled":true}]}'
aws s3api put-bucket-lifecycle-configuration --bucket my-bucket --lifecycle-configuration file://lifecycle.json
aws s3api get-bucket-lifecycle-configuration --bucket my-bucket
aws s3api put-bucket-policy --bucket my-bucket --policy file://policy.json
aws s3api get-bucket-location --bucket my-bucket
aws s3 rm s3://my-bucket/tmp/ --recursive --exclude '*' --include '*.log'   # include/exclude are evaluated in order
aws s3api list-multipart-uploads --bucket my-bucket --query 'Uploads[].[Key,UploadId,Initiated]' --output text   # abandoned uploads still bill
aws s3api abort-multipart-upload --bucket my-bucket --key key --upload-id "$UPLOAD_ID"
aws s3api restore-object --bucket my-bucket --key archive.tar --restore-request 'Days=7,GlacierJobParameters={Tier=Bulk}'
```

```json
{
  "Rules": [
    { "ID": "logs", "Filter": { "Prefix": "logs/" }, "Status": "Enabled",
      "Transitions": [{ "Days": 30, "StorageClass": "STANDARD_IA" }, { "Days": 90, "StorageClass": "GLACIER_IR" }],
      "Expiration": { "Days": 365 },
      "NoncurrentVersionExpiration": { "NoncurrentDays": 30 },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } }
  ]
}
```

Every versioned bucket needs `NoncurrentVersionExpiration` and every bucket needs `AbortIncompleteMultipartUpload`, or storage grows invisibly: `aws s3 ls` shows neither old versions nor incomplete parts. A bucket policy that denies `s3:*` without `aws:SecureTransport` enforces TLS; one that denies `s3:PutObject` unless `s3:x-amz-server-side-encryption` equals `aws:kms` enforces the key. Bucket policies are the resource side of [IAM evaluation](#iam-policy-evaluation) and can grant cross-account access on their own, so review them with `get-bucket-policy` during any access audit.

## IAM policy evaluation

An explicit `Deny` in any applicable policy wins. Otherwise an `Allow` must exist in an identity-based or resource-based policy, and every guardrail in play must also allow the action: service control policies (SCPs) and resource control policies (RCPs) from AWS Organizations, permission boundaries and session policies. "Denied with a policy that clearly allows it" is nearly always a guardrail. See [policy evaluation logic](https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_evaluation-logic.html).

```sh
aws iam simulate-principal-policy \
  --policy-source-arn arn:aws:iam::123456789012:role/Deploy \
  --action-names s3:PutObject --resource-arns 'arn:aws:s3:::my-bucket/key'
aws sts decode-authorization-message --encoded-message "$MSG" --query DecodedMessage --output text | jq
aws iam list-attached-role-policies --role-name Deploy
aws iam list-role-policies --role-name Deploy                       # inline policies
aws iam get-account-authorization-details > iam-dump.json           # everything, for offline review
```

`simulate-principal-policy` evaluates identity policies, permission boundaries and SCPs. It ignores RCPs, and evaluates a resource-based policy only when passed with `--resource-policy`, which works for IAM users but not roles. `decode-authorization-message` needs `sts:DecodeAuthorizationMessage`.

```json
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Action": ["s3:GetObject", "s3:PutObject"],
    "Resource": "arn:aws:s3:::my-bucket/prod/*",
    "Condition": { "StringEquals": { "aws:PrincipalTag/Team": "platform" } }
  }]
}
```

### Policy patterns

A role has two policies: the trust policy (who may assume it) and permission policies (what it may do). Most cross-account and CI failures are on the trust side.

```json
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "GitHubActionsOIDC",
      "Effect": "Allow",
      "Principal": { "Federated": "arn:aws:iam::123456789012:oidc-provider/token.actions.githubusercontent.com" },
      "Action": "sts:AssumeRoleWithWebIdentity",
      "Condition": {
        "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" },
        "StringLike": { "token.actions.githubusercontent.com:sub": "repo:my-org/my-app:ref:refs/heads/main" }
      }
    },
    {
      "Sid": "CrossAccountWithExternalId",
      "Effect": "Allow",
      "Principal": { "AWS": "arn:aws:iam::210987654321:role/Deployer" },
      "Action": "sts:AssumeRole",
      "Condition": { "StringEquals": { "sts:ExternalId": "my-app-prod" } }
    }
  ]
}
```

Permission patterns worth copying rather than inventing:

```json
{
  "Version": "2012-10-17",
  "Statement": [
    { "Sid": "ListOnlyMyPrefix", "Effect": "Allow", "Action": "s3:ListBucket",
      "Resource": "arn:aws:s3:::my-bucket",
      "Condition": { "StringLike": { "s3:prefix": ["team/${aws:PrincipalTag/Team}/*"] } } },
    { "Sid": "ObjectsInMyPrefix", "Effect": "Allow", "Action": ["s3:GetObject", "s3:PutObject", "s3:DeleteObject"],
      "Resource": "arn:aws:s3:::my-bucket/team/${aws:PrincipalTag/Team}/*" },
    { "Sid": "TagBasedEC2Control", "Effect": "Allow", "Action": ["ec2:StartInstances", "ec2:StopInstances"],
      "Resource": "arn:aws:ec2:*:123456789012:instance/*",
      "Condition": { "StringEquals": { "aws:ResourceTag/Owner": "${aws:PrincipalTag/Team}" } } },
    { "Sid": "RequireMFAForDeletes", "Effect": "Deny", "Action": ["s3:DeleteBucket", "rds:DeleteDBInstance"],
      "Resource": "*", "Condition": { "BoolIfExists": { "aws:MultiFactorAuthPresent": "false" } } },
    { "Sid": "RegionLock", "Effect": "Deny", "NotAction": ["iam:*", "sts:*", "organizations:*", "support:*", "cloudfront:*", "route53:*"],
      "Resource": "*", "Condition": { "StringNotEquals": { "aws:RequestedRegion": ["ap-southeast-2", "us-east-1"] } } },
    { "Sid": "OnlyFromVPC", "Effect": "Deny", "Action": "s3:*", "Resource": "arn:aws:s3:::my-bucket/*",
      "Condition": { "StringNotEquals": { "aws:SourceVpce": "vpce-0abc123" } } }
  ]
}
```

`s3:ListBucket` targets the bucket ARN and object actions target `bucket/*`; putting both actions on both resources is the classic mistake that makes a policy either fail or over-grant. `NotAction` with `Deny` is how region locks and SCPs exclude global services. Policy variables such as `${aws:PrincipalTag/Team}` make one policy serve every team (attribute-based access control) instead of one policy per team. Use `aws iam create-policy-version --set-as-default` to update a managed policy, and `aws iam get-policy-version` to read the current document, since `get-policy` returns only metadata. IAM Access Analyzer's `aws accessanalyzer validate-policy --policy-document file://p.json --policy-type IDENTITY_POLICY` catches syntax errors and over-broad grants before you attach anything. Permission boundaries cap what a role may grant to roles it creates, which is how a CI role can create IAM roles for applications without being able to grant itself administrator access.

## CloudWatch Logs and CloudTrail

```sh
aws logs tail /aws/lambda/my-fn --follow --since 10m --format short
aws logs tail /aws/eks/prod/cluster --filter-pattern 'ERROR'
aws logs start-query --log-group-name /aws/lambda/my-fn \
  --start-time "$(date -d '1 hour ago' +%s)" --end-time "$(date +%s)" \
  --query-string 'fields @timestamp, @message | filter @message like /ERROR/ | sort @timestamp desc | limit 20'
aws logs get-query-results --query-id <query-id>                    # repeat until status is Complete

aws cloudtrail lookup-events \
  --lookup-attributes AttributeKey=EventName,AttributeValue=TerminateInstances \
  --query 'Events[].[EventTime,Username,Resources[0].ResourceName]' --output text
```

Logs Insights queries are billed by data scanned; narrow the time range. `lookup-events` searches the 90-day event history of management events in the current region only. Data events, such as S3 object reads, are recorded only by a trail or event data store configured for them.

### Logs Insights recipes

Insights queries are pipelines. `fields`, `filter`, `parse`, `stats`, `sort` and `limit` cover nearly everything, and `@message`, `@timestamp`, `@logStream` and `@log` exist in every group. JSON log lines have their fields auto-discovered, so `filter level = "error"` works without parsing.

```sql
-- Error count per 5 minutes
filter @message like /(?i)error/
| stats count() as errors by bin(5m)

-- Top 10 status codes and paths from JSON-structured access logs
filter ispresent(status)
| stats count() as n by status, path
| sort n desc
| limit 10

-- Parse an unstructured line into fields, then aggregate
parse @message /duration=(?<ms>\d+)ms path=(?<path>\S+)/
| stats avg(ms) as avg_ms, pct(ms, 99) as p99_ms, max(ms) as max_ms by path
| sort p99_ms desc

-- Lambda: cold starts and memory headroom from REPORT lines
filter @type = "REPORT"
| stats count() as invocations, sum(strcontains(@message, "Init Duration")) as cold_starts,
        max(@maxMemoryUsed / 1000000) as max_mb, avg(@duration) as avg_ms by bin(1h)

-- EKS control plane audit: who deleted what
fields @timestamp, user.username, verb, objectRef.namespace, objectRef.resource, objectRef.name
| filter verb = "delete" and objectRef.resource not in ["events", "leases"]
| sort @timestamp desc

-- VPC Flow Logs: rejected connections by destination port
filter action = "REJECT"
| stats count() as rejects by dstPort, dstAddr
| sort rejects desc
| limit 20

-- Which log streams are noisiest (find the chatty pod)
stats count() as lines, sum(strlen(@message)) as bytes by @logStream
| sort bytes desc
| limit 10
```

```sh
# Run a query and wait for it, one command
QID=$(aws logs start-query --log-group-names /aws/eks/prod/cluster /aws/lambda/my-fn \
  --start-time "$(date -d '2 hours ago' +%s)" --end-time "$(date +%s)" \
  --query-string 'filter @message like /ERROR/ | stats count() by bin(10m)' --query queryId --output text)
until [ "$(aws logs get-query-results --query-id "$QID" --query status --output text)" = Complete ]; do sleep 2; done
aws logs get-query-results --query-id "$QID" --query 'results[].[ [0].value, [1].value ]' --output text

# Cheaper than Insights for a known pattern in a short window
aws logs filter-log-events --log-group-name /aws/lambda/my-fn --start-time "$(( $(date -d '30 min ago' +%s) * 1000 ))" \
  --filter-pattern '{ $.level = "error" }' --query 'events[].message' --output text

# Retention and size per log group: unset retention is the usual cost leak
aws logs describe-log-groups --query 'logGroups[].[logGroupName,retentionInDays,storedBytes]' --output text | sort -k3 -rn | head
aws logs put-retention-policy --log-group-name /aws/lambda/my-fn --retention-in-days 30
```

`--start-time` and `--end-time` are seconds for `start-query` and milliseconds for `filter-log-events`; getting that wrong returns nothing rather than an error. Saved queries live in `aws logs describe-query-definitions`, and `stats ... by bin()` output is the same shape as a Prometheus range query if you need to compare with [Prometheus](https://www.wiki.jodisand.me/prometheus/).

## RDS, Lambda and EKS

```sh
aws rds describe-db-instances \
  --query 'DBInstances[].[DBInstanceIdentifier,DBInstanceStatus,Endpoint.Address]' --output table
aws rds create-db-snapshot --db-instance-identifier prod --db-snapshot-identifier "pre-change-$(date +%F)"
aws rds describe-events --source-identifier prod --source-type db-instance --duration 1440

aws lambda invoke --function-name my-fn --payload '{"k":"v"}' --cli-binary-format raw-in-base64-out out.json
aws lambda get-function-configuration --function-name my-fn --query '[Timeout,MemorySize,LastUpdateStatus]'

aws eks update-kubeconfig --name prod --region ap-southeast-2       # writes a context to ~/.kube/config
aws eks describe-cluster --name prod --query 'cluster.[version,status,endpoint]'
aws eks list-nodegroups --cluster-name prod
aws eks list-access-entries --cluster-name prod                     # IAM principals mapped into the cluster
```

Take a manual snapshot before any risky change to a database you cannot rebuild; automated snapshots are deleted with the instance unless retained. Without `--cli-binary-format raw-in-base64-out`, CLI v2 expects `--payload` to be base64 and rejects plain JSON. For cluster work after `update-kubeconfig`, see [Kubernetes](https://www.wiki.jodisand.me/kubernetes/#start-with-a-failing-workload).

### EKS operations from the CLI

Cluster access is IAM first: `update-kubeconfig` writes an exec entry that runs `aws eks get-token`, and the resulting identity must appear in the cluster's access entries (the replacement for the `aws-auth` ConfigMap, authentication mode `API` or `API_AND_CONFIG_MAP`). Workload identity comes from Pod Identity associations or the older IRSA (an OIDC provider plus a role trust policy).

```sh
aws eks describe-cluster --name prod --query 'cluster.accessConfig.authenticationMode' --output text
aws eks create-access-entry --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/PlatformEngineer --type STANDARD
aws eks associate-access-policy --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/PlatformEngineer \
  --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy --access-scope type=cluster
aws eks associate-access-policy --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/AppTeam \
  --policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSEditPolicy --access-scope type=namespace,namespaces=my-namespace
aws eks list-associated-access-policies --cluster-name prod --principal-arn arn:aws:iam::123456789012:role/AppTeam

aws eks create-pod-identity-association --cluster-name prod --namespace my-namespace --service-account my-app \
  --role-arn arn:aws:iam::123456789012:role/my-app                           # role trusts pods.eks.amazonaws.com
aws eks list-pod-identity-associations --cluster-name prod --namespace my-namespace

aws eks describe-nodegroup --cluster-name prod --nodegroup-name workers --query 'nodegroup.[status,scalingConfig,releaseVersion,health.issues]'
aws eks update-nodegroup-version --cluster-name prod --nodegroup-name workers   # rolling AMI update to the cluster version
aws eks update-nodegroup-config --cluster-name prod --nodegroup-name workers --scaling-config minSize=3,maxSize=12,desiredSize=6
aws eks list-addons --cluster-name prod
aws eks describe-addon-versions --addon-name vpc-cni --kubernetes-version 1.34 --query 'addons[0].addonVersions[0].addonVersion' --output text
aws eks update-addon --cluster-name prod --addon-name vpc-cni --addon-version v1.20.0-eksbuild.1 --resolve-conflicts PRESERVE
aws eks describe-update --name prod --update-id "$UPDATE_ID"                    # progress of a cluster or nodegroup update
aws eks list-insights --cluster-name prod --query 'insights[?insightStatus.status!=`PASSING`].[name,insightStatus.status]' --output text   # upgrade blockers
```

Node problems that look like Kubernetes problems: a nodegroup `health.issues` entry of `Ec2SubnetInvalidConfiguration` or `InsufficientFreeAddresses` means the subnet is out of IPs, and pods stuck `ContainerCreating` with `failed to assign an IP address` is the same problem at the pod level (VPC CNI takes one address per pod unless prefix delegation is on). Check `aws ec2 describe-subnets --subnet-ids <id> --query 'Subnets[].AvailableIpAddressCount'`. A `kubectl` that fails with `the server has asked for the client to provide credentials` means `get-token` returned an identity with no access entry; `aws sts get-caller-identity` shows which one.

## Cost checks

```sh
aws ce get-cost-and-usage --time-period Start=2026-08-01,End=2026-09-01 \
  --granularity MONTHLY --metrics UnblendedCost --group-by Type=DIMENSION,Key=SERVICE \
  --query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn | head

aws ec2 describe-volumes --filters Name=status,Values=available --query 'Volumes[].[VolumeId,Size]' --output text
aws ec2 describe-addresses --query 'Addresses[?AssociationId==null].PublicIp' --output text
```

Each Cost Explorer API request is charged (USD 0.01 at review time). Unattached EBS volumes and unassociated Elastic IPs bill continuously; public IPv4 addresses are charged hourly whether or not they are attached.

```sh
# Daily spend for the last two weeks, to spot the day something changed
aws ce get-cost-and-usage --time-period Start="$(date -d '14 days ago' +%F)",End="$(date +%F)" --granularity DAILY --metrics UnblendedCost \
  --query 'ResultsByTime[].[TimePeriod.Start,Metrics.UnblendedCost.Amount]' --output text

# This month by a cost allocation tag (tag must be activated in Billing first)
aws ce get-cost-and-usage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --granularity MONTHLY --metrics UnblendedCost \
  --group-by Type=TAG,Key=Team --query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn

# Usage types inside one service: which EC2 line items are growing
aws ce get-cost-and-usage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --granularity MONTHLY --metrics UnblendedCost \
  --filter '{"Dimensions":{"Key":"SERVICE","Values":["Amazon Elastic Compute Cloud - Compute"]}}' \
  --group-by Type=DIMENSION,Key=USAGE_TYPE --query 'ResultsByTime[].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | sort -k2 -rn | head

# Forecast to month end
aws ce get-cost-forecast --time-period Start="$(date +%F)",End="$(date -d "$(date +%Y-%m-01) +1 month" +%F)" --metric UNBLENDED_COST --granularity MONTHLY --query Total.Amount --output text

# Rightsizing and idle recommendations
aws ce get-rightsizing-recommendation --service AmazonEC2 --query 'RightsizingRecommendations[].[CurrentInstance.ResourceId,RightsizingType,CurrentInstance.MonthlyCost]' --output text
aws compute-optimizer get-ec2-instance-recommendations --query 'instanceRecommendations[?finding==`Overprovisioned`].[instanceArn,recommendationOptions[0].instanceType]' --output text

# Savings Plans and RI coverage this month
aws ce get-savings-plans-utilization --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --query 'Total.Utilization.UtilizationPercentage' --output text
aws ce get-reservation-coverage --time-period Start="$(date +%Y-%m-01)",End="$(date +%F)" --query 'Total.CoverageHours.CoverageHoursPercentage' --output text

# Budgets and their current state
aws budgets describe-budgets --account-id "$(aws sts get-caller-identity --query Account --output text)" --query 'Budgets[].[BudgetName,BudgetLimit.Amount,CalculatedSpend.ActualSpend.Amount]' --output table
```

Cost Explorer data lags by up to 24 hours and `End` is exclusive. Snapshots (`aws ec2 describe-snapshots --owner-ids self`) and old AMIs are the other quiet bill; Data Lifecycle Manager or a tag-driven cleanup script keeps them bounded.

## Troubleshooting

| Symptom | Cause | Check |
| --- | --- | --- |
| `Unable to locate credentials` | No source in the chain answered | `aws configure list`; `aws sso login` |
| `The SSO session associated with this profile has expired` | Cached SSO token expired | `aws sso login --profile <name>` |
| `ExpiredToken` | Exported temporary credentials outlived their session | Unset `AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`, `AWS_SESSION_TOKEN` |
| Right command, wrong account | Environment variables override the profile | `aws sts get-caller-identity`; `env \| grep -c '^AWS_'` |
| `AccessDenied` with an encoded message | Missing allow, or a guardrail denies | `aws sts decode-authorization-message` |
| `AccessDenied` although the policy allows | SCP, RCP, permission boundary, resource policy or KMS key policy | Simulate; check the key policy for encrypted resources |
| Empty result, no error | Wrong region; most resources are regional | `aws configure get region`, `--region` |
| `Could not connect to the endpoint URL` | Invalid region name, proxy or no network path | `--debug 2>&1 \| grep -i endpoint` |
| `ThrottlingException` or `Rate exceeded` | API rate limit | `AWS_RETRY_MODE=adaptive`, `AWS_MAX_ATTEMPTS=10` |
| `SignatureDoesNotMatch` or `RequestTimeTooSkewed` | Local clock drift | `timedatectl`; sync NTP |
| `ssm start-session` fails with `TargetNotConnected` | SSM agent offline, no instance role, or no route to SSM endpoints | `aws ssm describe-instance-information` |
| `SessionManagerPlugin is not found` | Plugin not installed on the workstation | Install `session-manager-plugin`; `session-manager-plugin --version` |
| `kubectl` says `the server has asked for the client to provide credentials` | Identity from `aws eks get-token` has no access entry | `aws sts get-caller-identity`; `aws eks list-access-entries --cluster-name <name>` |
| Pods `ContainerCreating` with `failed to assign an IP address` | Subnet out of addresses (VPC CNI) | `aws ec2 describe-subnets --query 'Subnets[].[SubnetId,AvailableIpAddressCount]'` |
| `Not authorized to perform sts:AssumeRoleWithWebIdentity` | Trust policy `sub`/`aud` condition does not match the token | Decode the token claims; compare with the trust policy `Condition` |
| `AccessDenied` on `s3:ListBucket` while `GetObject` works | Policy grants object actions on `bucket/*` but not `ListBucket` on the bucket ARN | `aws iam simulate-principal-policy --action-names s3:ListBucket --resource-arns arn:aws:s3:::my-bucket` |
| `KMS.AccessDeniedException` reading an encrypted object or parameter | Key policy does not grant the principal `kms:Decrypt` | `aws kms get-key-policy --key-id <id> --policy-name default` |
| Logs Insights returns no rows | Time units wrong (seconds vs milliseconds), or the field is not auto-discovered | `--start-time` in seconds for `start-query`; use `parse` for unstructured lines |
| `InvalidParameterValue` on `--tag-specifications` or `--ip-permissions` | Shorthand syntax mismatch | Use `--generate-cli-skeleton` and pass `--cli-input-json file://` |
| Cost Explorer shows zero for a tag | Tag not activated as a cost allocation tag, or activated after the spend | Billing console tag activation; wait 24 h |
| `An error occurred (ValidationException) ... nodegroup` update stuck | Pod disruption budgets block node drain | `kubectl get pdb -A`; `aws eks describe-update` shows the error |

`--debug` prints every request, the endpoint and the credential provider that answered. It can include signed headers, so do not paste it into tickets unedited.

## Oneliners

```sh
# Confirm account and region before anything destructive
aws sts get-caller-identity --query '[Account,Arn]' --output text; aws configure get region

# Running instances in every enabled region
for r in $(aws ec2 describe-regions --query 'Regions[].RegionName' --output text); do aws ec2 describe-instances --region "$r" --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,InstanceType]' --output text | sed "s/^/$r /"; done

# Instance ID from a Name tag
aws ec2 describe-instances --filters 'Name=tag:Name,Values=web-01' --query 'Reservations[0].Instances[0].InstanceId' --output text

# Security groups with ingress open to the internet (IPv4 or IPv6)
aws ec2 describe-security-groups --query 'SecurityGroups[?IpPermissions[?IpRanges[?CidrIp==`0.0.0.0/0`] || Ipv6Ranges[?CidrIpv6==`::/0`]]].[GroupId,GroupName]' --output text

# Buckets without a full public access block
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do aws s3api get-public-access-block --bucket "$b" --query 'PublicAccessBlockConfiguration.[BlockPublicAcls,IgnorePublicAcls,BlockPublicPolicy,RestrictPublicBuckets]' --output text 2>/dev/null | grep -q False && echo "$b"; done

# Largest objects in a bucket
aws s3api list-objects-v2 --bucket my-bucket --query 'sort_by(Contents,&Size)[-10:].[Size,Key]' --output text

# Access keys older than 90 days
aws iam list-users --query 'Users[].UserName' --output text | tr '\t' '\n' | while read -r u; do aws iam list-access-keys --user-name "$u" --query "AccessKeyMetadata[?CreateDate<='$(date -d '90 days ago' +%Y-%m-%d)'].[UserName,AccessKeyId,CreateDate]" --output text; done

# Who acted on an instance
aws cloudtrail lookup-events --lookup-attributes AttributeKey=ResourceName,AttributeValue=i-0abc123 --query 'Events[].[EventTime,EventName,Username]' --output text

# Copy between buckets without downloading
aws s3 sync s3://src-bucket/prefix s3://dst-bucket/prefix --source-region us-east-1 --region ap-southeast-2

# Count of each value of a tag
aws ec2 describe-tags --filters Name=key,Values=Environment --query 'Tags[].Value' --output text | tr '\t' '\n' | sort | uniq -c

# Instances without an Owner tag
aws ec2 describe-instances --query 'Reservations[].Instances[?!not_null(Tags[?Key==`Owner`].Value|[0])].[InstanceId,LaunchTime]' --output text

# Instances still allowing IMDSv1
aws ec2 describe-instances --filters Name=metadata-options.http-tokens,Values=optional --query 'Reservations[].Instances[].InstanceId' --output text

# Private IP to instance name map for the whole account
aws ec2 describe-instances --query 'Reservations[].Instances[].[PrivateIpAddress, Tags[?Key==`Name`].Value|[0]]' --output text | sort -t. -k1,1n -k2,2n -k3,3n -k4,4n

# Unencrypted EBS volumes
aws ec2 describe-volumes --filters Name=encrypted,Values=false --query 'Volumes[].[VolumeId,Size,Attachments[0].InstanceId]' --output text

# Snapshots older than a year, with size
aws ec2 describe-snapshots --owner-ids self --query "Snapshots[?StartTime<='$(date -d '1 year ago' +%Y-%m-%d)'].[SnapshotId,VolumeSize,StartTime,Description]" --output text

# AMIs you own that no instance uses
comm -23 <(aws ec2 describe-images --owners self --query 'Images[].ImageId' --output text | tr '\t' '\n' | sort) <(aws ec2 describe-instances --query 'Reservations[].Instances[].ImageId' --output text | tr '\t' '\n' | sort -u)

# Security groups attached to nothing
comm -23 <(aws ec2 describe-security-groups --query 'SecurityGroups[?GroupName!=`default`].GroupId' --output text | tr '\t' '\n' | sort) <(aws ec2 describe-network-interfaces --query 'NetworkInterfaces[].Groups[].GroupId' --output text | tr '\t' '\n' | sort -u)

# Subnets running low on addresses
aws ec2 describe-subnets --query 'Subnets[?AvailableIpAddressCount<`20`].[SubnetId,CidrBlock,AvailableIpAddressCount,Tags[?Key==`Name`].Value|[0]]' --output text

# Total size of a bucket prefix without listing every object
aws s3 ls s3://my-bucket/logs/ --recursive --summarize | tail -2

# Bucket sizes from CloudWatch (free; updated daily), in GiB
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do printf '%s\t' "$b"; aws cloudwatch get-metric-statistics --namespace AWS/S3 --metric-name BucketSizeBytes --dimensions Name=BucketName,Value="$b" Name=StorageType,Value=StandardStorage --start-time "$(date -d '2 days ago' -u +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" --period 86400 --statistics Average --query 'Datapoints[-1].Average' --output text | awk '{printf "%.1f\n", $1/1073741824}'; done

# Buckets without versioning
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do [ "$(aws s3api get-bucket-versioning --bucket "$b" --query Status --output text)" = Enabled ] || echo "$b"; done

# Buckets without a lifecycle configuration
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do aws s3api get-bucket-lifecycle-configuration --bucket "$b" >/dev/null 2>&1 || echo "$b"; done

# Delete every version and delete marker under a prefix, 1000 at a time (irreversible; run the list-object-versions half alone first to review)
aws s3api list-object-versions --bucket my-bucket --prefix tmp/ --query '{Objects: [Versions[].{Key:Key,VersionId:VersionId}, DeleteMarkers[].{Key:Key,VersionId:VersionId}][] | [0:1000], Quiet: `true`}' --output json | aws s3api delete-objects --bucket my-bucket --delete file:///dev/stdin

# Roles nobody has used in 90 days
aws iam list-roles --query "Roles[?RoleLastUsed.LastUsedDate<='$(date -d '90 days ago' +%Y-%m-%d)' || !RoleLastUsed.LastUsedDate].[RoleName,RoleLastUsed.LastUsedDate]" --output text

# Users with console passwords but no MFA
aws iam generate-credential-report >/dev/null; sleep 5; aws iam get-credential-report --query Content --output text | base64 -d | awk -F, 'NR>1 && $4=="true" && $8=="false" {print $1}'

# Policies that grant Action "*" on Resource "*"
for arn in $(aws iam list-policies --scope Local --query 'Policies[].Arn' --output text); do v=$(aws iam get-policy --policy-arn "$arn" --query Policy.DefaultVersionId --output text); aws iam get-policy-version --policy-arn "$arn" --version-id "$v" --query 'PolicyVersion.Document.Statement[?Effect==`Allow` && (Action==`*` || contains(Action, `*`)) && Resource==`*`]' --output text | grep -q . && echo "$arn"; done

# Who is allowed to assume a role
aws iam get-role --role-name Deploy --query 'Role.AssumeRolePolicyDocument.Statement[].Principal' --output json

# Console sign-ins in the last day, by user
aws cloudtrail lookup-events --lookup-attributes AttributeKey=EventName,AttributeValue=ConsoleLogin --start-time "$(date -d '1 day ago' -u +%FT%TZ)" --query 'Events[].Username' --output text | tr '\t' '\n' | sort | uniq -c

# Every API call by one principal today
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=alice --start-time "$(date -u +%FT00:00:00Z)" --query 'Events[].[EventTime,EventSource,EventName]' --output text

# Log groups without retention, with stored size in GiB
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].[logGroupName,storedBytes]' --output text | awk '{printf "%s\t%.2f\n", $1, $2/1073741824}' | sort -k2 -rn

# Set 30-day retention on every log group that has none
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].logGroupName' --output text | tr '\t' '\n' | xargs -r -I{} aws logs put-retention-policy --log-group-name {} --retention-in-days 30

# Lambda functions on a deprecated runtime
aws lambda list-functions --query 'Functions[?starts_with(Runtime, `python3.8`) || starts_with(Runtime, `nodejs16`)].[FunctionName,Runtime]' --output text

# Lambda error rate in the last hour for one function
aws cloudwatch get-metric-statistics --namespace AWS/Lambda --metric-name Errors --dimensions Name=FunctionName,Value=my-fn --start-time "$(date -d '1 hour ago' -u +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" --period 3600 --statistics Sum --query 'Datapoints[0].Sum' --output text

# RDS instances that are publicly accessible or unencrypted
aws rds describe-db-instances --query 'DBInstances[?PubliclyAccessible || !StorageEncrypted].[DBInstanceIdentifier,PubliclyAccessible,StorageEncrypted]' --output text

# Wait for an RDS snapshot then print its ARN
aws rds wait db-snapshot-completed --db-snapshot-identifier pre-change && aws rds describe-db-snapshots --db-snapshot-identifier pre-change --query 'DBSnapshots[0].DBSnapshotArn' --output text

# EKS: cluster version and every nodegroup's release version (mismatch means a pending upgrade)
aws eks describe-cluster --name prod --query cluster.version --output text; for ng in $(aws eks list-nodegroups --cluster-name prod --query 'nodegroups[]' --output text); do aws eks describe-nodegroup --cluster-name prod --nodegroup-name "$ng" --query 'nodegroup.[nodegroupName,version,releaseVersion,status]' --output text; done

# EKS: access entries and their policies
for p in $(aws eks list-access-entries --cluster-name prod --query 'accessEntries[]' --output text); do echo "$p"; aws eks list-associated-access-policies --cluster-name prod --principal-arn "$p" --query 'associatedAccessPolicies[].[policyArn,accessScope.type]' --output text | sed 's/^/  /'; done

# SSM: managed instances whose agent has not checked in for a day
aws ssm describe-instance-information --query "InstanceInformationList[?LastPingDateTime<='$(date -d '1 day ago' -u +%FT%TZ)'].[InstanceId,PingStatus,LastPingDateTime]" --output text

# SSM: run one command on every instance with a tag and print output per host
CMD=$(aws ssm send-command --document-name AWS-RunShellScript --targets Key=tag:Environment,Values=prod --parameters 'commands=["uptime"]' --query Command.CommandId --output text); sleep 10; aws ssm list-command-invocations --command-id "$CMD" --details --query 'CommandInvocations[].[InstanceId,Status,CommandPlugins[0].Output]' --output text

# Parameter Store: every parameter under a path with its last modification
aws ssm get-parameters-by-path --path /my-app/ --recursive --query 'Parameters[].[Name,Type,Version,LastModifiedDate]' --output table

# Service quotas you are close to (EC2 vCPU example)
aws service-quotas get-service-quota --service-code ec2 --quota-code L-1216C47A --query 'Quota.Value' --output text; aws ec2 describe-instances --filters Name=instance-state-name,Values=running --query 'Reservations[].Instances[].CpuOptions.[CoreCount,ThreadsPerCore]' --output text | awk '{s+=$1*$2} END {print s " vCPUs in use"}'

# Regions where you have any EC2 resources at all
for r in $(aws ec2 describe-regions --query 'Regions[].RegionName' --output text); do n=$(aws ec2 describe-instances --region "$r" --query 'length(Reservations[].Instances[])' --output text); [ "$n" = 0 ] || echo "$r $n"; done

# Retry-hardened settings for a scripted session
export AWS_RETRY_MODE=adaptive AWS_MAX_ATTEMPTS=10 AWS_PAGER=""
```

## Scripts

Produce a one-page security posture report for an account: public buckets, open security groups, IMDSv1 instances, stale keys and roles, and log groups without retention.

```sh
#!/usr/bin/env bash
# account-audit.sh [profile]: read-only posture checks, one section per finding class
set -euo pipefail
export AWS_PAGER="" AWS_RETRY_MODE=adaptive AWS_MAX_ATTEMPTS=10
[ $# -eq 0 ] || export AWS_PROFILE=$1
acct=$(aws sts get-caller-identity --query '[Account,Arn]' --output text)
printf 'Account/identity: %s\nRegion: %s\n\n' "$acct" "$(aws configure get region || echo unset)"
section() { printf '== %s ==\n' "$1"; }

section "Buckets without full public access block"
for b in $(aws s3api list-buckets --query 'Buckets[].Name' --output text); do
  aws s3api get-public-access-block --bucket "$b" --query 'PublicAccessBlockConfiguration.[BlockPublicAcls,IgnorePublicAcls,BlockPublicPolicy,RestrictPublicBuckets]' --output text 2>/dev/null | grep -q False && echo "$b"
done || true

section "Security groups open to the world on ports other than 80/443"
aws ec2 describe-security-groups --query 'SecurityGroups[?IpPermissions[?(IpRanges[?CidrIp==`0.0.0.0/0`] || Ipv6Ranges[?CidrIpv6==`::/0`]) && !(FromPort==`80` || FromPort==`443`)]].[GroupId,GroupName]' --output text

section "Instances allowing IMDSv1"
aws ec2 describe-instances --filters Name=metadata-options.http-tokens,Values=optional Name=instance-state-name,Values=running --query 'Reservations[].Instances[].[InstanceId,Tags[?Key==`Name`].Value|[0]]' --output text

section "Unencrypted volumes"
aws ec2 describe-volumes --filters Name=encrypted,Values=false --query 'Volumes[].[VolumeId,Size]' --output text

section "Access keys older than 90 days"
cutoff=$(date -d '90 days ago' +%Y-%m-%d)
for u in $(aws iam list-users --query 'Users[].UserName' --output text); do
  aws iam list-access-keys --user-name "$u" --query "AccessKeyMetadata[?CreateDate<='$cutoff' && Status=='Active'].[UserName,AccessKeyId,CreateDate]" --output text
done

section "Roles unused for 90 days"
aws iam list-roles --query "Roles[?!starts_with(Path, '/aws-service-role/') && (RoleLastUsed.LastUsedDate<='$cutoff' || !RoleLastUsed.LastUsedDate)].RoleName" --output text | tr '\t' '\n'

section "Log groups without retention (GiB stored)"
aws logs describe-log-groups --query 'logGroups[?!retentionInDays].[logGroupName,storedBytes]' --output text | awk '{printf "%s\t%.2f\n", $1, $2/1073741824}' | sort -k2 -rn | head -20

section "CloudTrail"
aws cloudtrail describe-trails --query 'trailList[].[Name,IsMultiRegionTrail,LogFileValidationEnabled]' --output text
```

Tag-driven cleanup of EBS snapshots: delete snapshots older than a retention period unless tagged `Retain=true` or still referenced by an AMI, with a dry run by default.

```sh
#!/usr/bin/env bash
# snapshot-prune.sh DAYS [--apply]: delete own EBS snapshots older than DAYS unless retained or used by an AMI
set -euo pipefail
days=$1; apply=${2:-}
export AWS_PAGER=""
cutoff=$(date -d "$days days ago" +%Y-%m-%dT%H:%M:%SZ)
in_use=$(aws ec2 describe-images --owners self --query 'Images[].BlockDeviceMappings[].Ebs.SnapshotId' --output text | tr '\t' '\n' | sort -u)
count=0; bytes=0
while read -r id size start retain; do
  [ -n "$id" ] || continue
  [ "$retain" = true ] && continue
  grep -qx "$id" <<<"$in_use" && continue
  count=$((count + 1)); bytes=$((bytes + size))
  if [ "$apply" = --apply ]; then
    aws ec2 delete-snapshot --snapshot-id "$id" && echo "deleted $id ($size GiB, $start)"
  else
    echo "would delete $id ($size GiB, $start)"
  fi
done < <(aws ec2 describe-snapshots --owner-ids self --query "Snapshots[?StartTime<='$cutoff'].[SnapshotId,VolumeSize,StartTime,Tags[?Key=='Retain'].Value|[0]]" --output text)
printf '%d snapshots, %d GiB %s\n' "$count" "$bytes" "$([ "$apply" = --apply ] && echo deleted || echo 'to delete (pass --apply)')"
```

Poll a Logs Insights query to completion and print the result as a table, for use in runbooks and cron jobs.

```python
#!/usr/bin/env python3
"""Run a CloudWatch Logs Insights query and print a table.

Usage: insights.py LOG_GROUP MINUTES 'QUERY' [more log groups...]
Example: insights.py /aws/lambda/my-fn 60 'filter @message like /ERROR/ | stats count() by bin(5m)'
"""
import sys
import time

import boto3

group, minutes, query, *more = sys.argv[1:]
logs = boto3.client("logs")
now = int(time.time())
start = logs.start_query(logGroupNames=[group, *more], startTime=now - int(minutes) * 60, endTime=now, queryString=query)
qid = start["queryId"]
while True:
    res = logs.get_query_results(queryId=qid)
    if res["status"] in ("Complete", "Failed", "Cancelled", "Timeout"):
        break
    time.sleep(1)
if res["status"] != "Complete":
    sys.exit(f"query {res['status']}")
rows = [{c["field"]: c["value"] for c in r if c["field"] != "@ptr"} for r in res["results"]]
if not rows:
    sys.exit("no results")
cols = list(rows[0])
width = {c: max(len(c), *(len(r.get(c, "")) for r in rows)) for c in cols}
print("  ".join(c.ljust(width[c]) for c in cols))
for r in rows:
    print("  ".join(r.get(c, "").ljust(width[c]) for c in cols))
stats = res["statistics"]
print(f"\n{stats['recordsMatched']:.0f} matched, {stats['bytesScanned'] / 1e9:.2f} GB scanned", file=sys.stderr)
```

## Further reading

- [AWS CLI v2 user guide](https://docs.aws.amazon.com/cli/latest/userguide/) and [command reference](https://docs.aws.amazon.com/cli/latest/reference/)
- [Controlling command output (JMESPath)](https://docs.aws.amazon.com/cli/latest/userguide/cli-usage-output.html)
- [IAM policy evaluation logic](https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_evaluation-logic.html) and [condition keys](https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_condition-keys.html)
- [CloudWatch Logs Insights query syntax](https://docs.aws.amazon.com/AmazonCloudWatch/latest/logs/CWL_QuerySyntax.html)
- [EKS access entries](https://docs.aws.amazon.com/eks/latest/userguide/access-entries.html) and [Pod Identity](https://docs.aws.amazon.com/eks/latest/userguide/pod-identities.html)
- [S3 lifecycle configuration](https://docs.aws.amazon.com/AmazonS3/latest/userguide/object-lifecycle-mgmt.html)


