Tags, cost allocation and budgets
"The bill went up" is only actionable once you can say whose bill went up. On AWS that answer comes from tags: key=value labels on resources (team=orders, env=prod) that Cost Explorer can group by - if they are on the resources, spelled consistently, and activated for cost allocation. Then AWS Budgets turns "we noticed at the end of the month" into an email on day nine. This lesson builds that chain and shows where each link usually breaks.
Need to know: a tag groups costs only after it is an active cost allocation tag, and only from the activation on (a backfill re-processes up to 12 months). Cost Explorer groups by tag as team$orders; team$ is everything without the tag. Tag values are case-sensitive: Search and search are two teams. Tag policies (Organizations) define the allowed key spelling and values and report non-compliance - they do not make a tag mandatory. A budget watches a cost (all of it, or filtered: TagKeyValue = user:team$orders) and alerts on ACTUAL or FORECASTED spend to email or SNS.
Costs by team
$ M=$(date +%Y-%m-01); T=$(date -d tomorrow +%F)
$ aws ce get-cost-and-usage --time-period Start=$M,End=$T --granularity MONTHLY --metrics UnblendedCost --group-by Type=TAG,Key=team --query 'ResultsByTime[0].Groups[].[Keys[0],Metrics.UnblendedCost.Amount]' --output text | awk -F'\t' '{printf "%-16s %8.2f\n", $1, $2}'
team$ 54.95
team$Search 50.15
team$data 143.90
team$orders 201.62
team$payments 113.96
team$platform 265.25
$ aws ce list-cost-allocation-tags --status Active --query 'CostAllocationTags[].[TagKey,Type,Status]' --output text
team UserDefined Active
Three things to read in that table:
team$- costs with no team tag at all: shared things nobody tagged (the CloudWatch metrics bill, KMS keys, the hosted zone) and resources somebody forgot. Every account has this line; the goal is to keep it small and explained;team$Searchnext to the lower-case values - a typo that will never add up withsearch;- only tags listed as Active cost allocation tags group at all. A tag key nobody activated shows every cost under
key$, as if nothing were tagged.
Activation works from the moment you activate (it can take up to 24 hours to show); costs from before keep their old grouping until a backfill re-processes them with the current tags:
$ aws ce update-cost-allocation-tags-status --cost-allocation-tags-status TagKey=env,Status=Active
{
"Errors": []
}
$ aws ce start-cost-allocation-tag-backfill --backfill-from $(date +%Y-%m-01)T00:00:00Z
{
"BackfillRequest": {
"BackfillFrom": "2026-09-01T00:00:00Z",
"RequestedAt": "2026-09-22T20:00:04Z",
"BackfillStatus": "PROCESSING",
"LastUpdatedAt": "2026-09-22T20:00:04Z"
}
}
$ aws ce list-cost-allocation-tag-backfill-history --query 'BackfillRequests[0].[BackfillFrom,BackfillStatus]' --output text
2026-09-01T00:00:00Z PROCESSING
Finding tags across services
The Resource Groups Tagging API sees tags across services:
$ aws resourcegroupstaggingapi get-tag-values --key team
{
"TagValues": [
"Search",
"orders",
"payments",
"platform"
]
}
$ aws resourcegroupstaggingapi get-resources --tag-filters Key=team,Values=Search --query 'ResourceTagMappingList[].ResourceARN' --output text
arn:aws:ec2:eu-central-1:111122223333:instance/i-0f60718293a4b5c63
$ aws resourcegroupstaggingapi get-resources --resource-type-filters ec2:instance --query 'ResourceTagMappingList[?!not_null(Tags[?Key==`team`] | [0])].ResourceARN' --output text
arn:aws:ec2:eu-central-1:111122223333:instance/i-0a718293a4b5c6d74
The last one is the "who forgot" list: instances with no team tag (bastion-old). Before fixing anything by hand, make the spelling a rule.
Tag policies: the spelling, enforced organization-wide
A tag policy is an Organizations policy (this account is the delegated administrator for them, the next lesson) that says how a tag key is spelled and which values are allowed:
$ cat ~/oncall-lab/labs/aws/try3/tag-policy.json
{
"tags": {
"team": {
"tag_key": {
"@@assign": "team"
},
"tag_value": {
"@@assign": [
"orders",
"payments",
"platform",
"data",
"search"
]
}
}
}
}
$ P=$(aws organizations create-policy --type TAG_POLICY --name try-team-values --description "team: lower case, known teams" --content file://$HOME/oncall-lab/labs/aws/try3/tag-policy.json --query Policy.PolicySummary.Id --output text); echo $P
p-92f83332
$ aws organizations attach-policy --policy-id $P --target-id ou-f6g7-8m3tw1ab
$ aws organizations describe-effective-policy --policy-type TAG_POLICY --query EffectivePolicy.PolicyContent --output text | jq -c .
{"tags":{"team":{"tag_key":"team","tag_value":["orders","payments","platform","data","search"]}}}
$ aws resourcegroupstaggingapi get-resources --include-compliance-details --exclude-compliant-resources --query 'ResourceTagMappingList[].[ResourceARN,ComplianceDetails.KeysWithNoncompliantValues[0]]' --output text
arn:aws:ec2:eu-central-1:111122223333:instance/i-0f60718293a4b5c63 team
@@assign is the inheritance operator (a child OU's policy can replace or extend the parent's); the effective policy is what applies to this account after inheritance. The compliance report lists resources whose team value is not in the list - and says nothing about resources that have no team tag at all: a tag policy checks the tags that exist. Requiring a tag takes something else - an SCP with a Null condition on aws:RequestTag/team for create calls, or the rules of your infrastructure-as-code pipeline. With enforced_for, a tag policy also blocks non-compliant tagging operations for the listed resource types.
Fixing a tag is one call for up to 20 resources - and, like activation, it changes costs from now on, not the past:
$ aws resourcegroupstaggingapi tag-resources --resource-arn-list arn:aws:ec2:eu-central-1:111122223333:instance/i-0f60718293a4b5c63 --tags team=search
{
"FailedResourcesMap": {}
}
$ aws resourcegroupstaggingapi get-tag-values --key team --query 'TagValues' --output text
orders payments platform search
$ aws resourcegroupstaggingapi get-resources --include-compliance-details --exclude-compliant-resources --query 'length(ResourceTagMappingList)'
0
Budgets
A budget is a limit and a set of alerts:
$ cat ~/oncall-lab/labs/aws/try3/budget-orders.json
{
"BudgetName": "try-orders-monthly",
"BudgetLimit": {
"Amount": "250",
"Unit": "USD"
},
"CostFilters": {
"TagKeyValue": [
"user:team$orders"
]
},
"TimeUnit": "MONTHLY",
"BudgetType": "COST"
}
$ aws budgets create-budget --account-id 111122223333 --budget file://$HOME/oncall-lab/labs/aws/try3/budget-orders.json --notifications-with-subscribers file://$HOME/oncall-lab/labs/aws/try3/alerts.json
$ aws budgets describe-budget --account-id 111122223333 --budget-name try-orders-monthly --query 'Budget.[BudgetLimit.Amount,CalculatedSpend.ActualSpend.Amount,CalculatedSpend.ForecastedSpend.Amount]' --output text
250.0 201.619 276.727
$ aws budgets describe-notifications-for-budget --account-id 111122223333 --budget-name try-orders-monthly --query 'Notifications[].[NotificationType,Threshold,NotificationState]' --output text
ACTUAL 80 ALARM
FORECASTED 100 ALARM
$ grep '^Subject' ~/oncall-lab/labs/aws/pager.log | tail -2
Subject: AWS Budgets: try-orders-monthly has exceeded your alert threshold
Subject: AWS Budgets: try-orders-monthly has exceeded your alert threshold
Two alerts, two jobs: ACTUAL > 80% says "you have spent most of it" (late, but certain); FORECASTED > 100% says "at this rate you will overspend" (early, a projection). Budgets evaluates a few times a day, so an alert can arrive hours after the threshold is crossed - a budget is a safety net for the month, not an outage alarm. The filter syntax is the trap:
$ aws budgets create-budget --account-id 111122223333 --budget '{"BudgetName": "try-orders-typo", "BudgetLimit": {"Amount": "250", "Unit": "USD"}, "CostFilters": {"TagKeyValue": ["team$orders"]}, "TimeUnit": "MONTHLY", "BudgetType": "COST"}'
$ aws budgets describe-budget --account-id 111122223333 --budget-name try-orders-typo --query 'Budget.CalculatedSpend.ActualSpend'
{
"Amount": "0.0",
"Unit": "USD"
}
Without the user: prefix the filter matches no cost, the budget shows $0.0 spent all month and never alerts - with no error anywhere. After creating a budget, always check that its actual spend is the number you expected.
Beyond budgets, Cost Anomaly Detection learns each service's normal daily spend and alerts on deviations without a threshold you have to guess - a good second net for "the NAT gateway bill".
In an interview: "How do you make teams own their AWS costs?" - "Tag every resource with a team and an environment, activate those as cost allocation tags and backfill, keep the untagged share small with a tag policy for the spelling and IaC rules for the presence, give each team a monthly budget filtered on its tag with an actual and a forecasted alert, and review the cost by team in Cost Explorer every month."
You can now: group costs by a tag and read team$, activate and backfill cost allocation tags, find and fix tags with the tagging API, write and attach a tag policy and read its compliance report, and create a budget whose filter really matches - with alerts on actual and forecasted spend.