Observability Alerting — Two-Tier Cost-Free Design
Applies to: all Mantle instances · Convention checks: C144 (free-tier cap, HIGH), C145 (architecture guidance, MEDIUM)
The billing trap: why "aggregate" mode saves nothing
CloudWatch bills metric-math alarms per referenced metric, not per alarm object. The framework's observability module has an enable_per_function_lambda_alarms = false path that collapses N per-function alarms into 2 aggregate metric-math alarms — but each of those 2 alarms still references all N Lambda metrics. The billing line item is 2 × N alarm-metrics, identical to per-function mode.
Composite alarms make it worse: $0.50/alarm/month each, on top of the children they reference. Never composite.
The only genuine cost levers are:
- Keep the total alarm-metric count at or below 10 (always free per account).
- Route the bulk signal (Lambda errors) to a log subscription filter instead of alarms — filters are free.
Two-tier design
Tier A — free-tier alarm set (≤10 alarm-metrics, $0)
Generated by mantle generate infra in cost-optimized mode (default). Budget is spent on the highest-value metric-only signals:
| Signal | Alarm-metrics consumed | Why metric-only |
|---|---|---|
| API Gateway 5XX | 1 | No log line for synchronous failures |
| DLQ depth | 1 per DLQ | Queue-depth is a metric; no error log |
| Lambda Errors + Throttles | 2 per criticalFunction entry | Throttles are metric-only; errors also covered by Tier B |
Tier B — log-error notifier ($0)
One account-level CloudWatch Logs subscription filter (aws_cloudwatch_log_account_policy, policy_type = "SUBSCRIPTION_FILTER_POLICY") with filter pattern { $.level = "ERROR" }, delivering to a framework LogNotifier standalone Lambda → existing SNS email topic. The policy applies account-wide. SUBSCRIPTION_FILTER_POLICY supports selectionCriteria only as LogGroupName NOT IN [...] (exclusion — not LIKE/positive IN); we omit it because AWS automatically excludes a Lambda destination's own log group from account-level subscription filters (no recursion), and the filter pattern is the effective scope: only structured Powertools ERROR logs match. The handler also skips CloudWatch CONTROL_MESSAGE health-check frames (only DATA_MESSAGE is processed). Each instance runs in a dedicated AWS account.
Coverage: every Lambda's errors, including future functions auto-discovered on deploy. Delivery is event-driven (AWS-managed, sub-second latency). Cost: $0 — account-level filter is free (one per account/region), notifier invocations sit inside the 1M free tier, SNS publish and email delivery are free at this scale.
Alarm-metric count formula (C144)
alarmMetrics = 2 × lambda_function_names.length
+ sqs_dlq_names.length
+ (api_gateway_name ? 1 : 0)
+ 2 × eventbridge_rule_names.length
+ sqs_age_queue_names.length
+ custom_alarms.lengthThis formula is the single source of truth implemented in countAlarmMetrics() in packages/cli/src/generators/infra-observability.ts and imported by packages/cli/src/analysis/observability-alarm-count.ts. Two enforcement points:
- Generator throw (
mantle build/mantle generate infra): fires when the formula exceeds 10 during code generation. mantle check observability: parsesinfra/observability.tfon disk, catches ejected or hand-edited files that bypass the generator.
The 2 × lambda_function_names.length term holds regardless of enable_per_function_lambda_alarms — do not attempt to "optimize" this by looking at alarm object count.
Config surface
defineConfig({
observability: {
alerts: {
email: "ops@example.com",
// 'cost-optimized' (default): per-function alarms limited to criticalFunctions;
// all-function errors covered by Tier B log notifier.
// 'per-function': all discovered Lambdas get per-function alarms (same cap applies).
mode: "cost-optimized",
// Lambda names that get dedicated Errors + Throttles alarms.
// Each entry consumes 2 of the 10 free alarm-metric slots.
criticalFunctions: ["UserLogin", "FilesGet"],
// Enable the Tier B log-error notifier. Default: true in cost-optimized mode.
// Requires a LogNotifier standalone Lambda in src/lambdas/standalone/.
errorLogNotifier: true,
},
},
});Budget planning example (OMD): 5 DLQs + 1 API Gateway = 6 alarm-metrics consumed → 4 remaining → at most 2 criticalFunctions (2 × 2 = 4). Total: 10. Errors on all 20 functions covered by Tier B.
Coverage tradeoff (accepted)
Lambda throttles on non-critical functions are not covered. Throttles are metric-only — no log line is emitted, so Tier B cannot see them. Covering all N functions' throttles via metrics costs money. The accepted approach: cover throttles on the 1–2 busiest functions via criticalFunctions and accept the gap for low-traffic functions. If comprehensive throttle coverage is later needed, a scheduled watcher Lambda using GetMetricStatistics (free API tier) can close the gap.
Instance setup
- Add
src/lambdas/standalone/LogNotifier/index.ts— a thin handler that re-exportscreateLogSubscriptionNotifierfrom@j0nathan-ll0yd/observabilityviadefineLambda({ timeout: 30, bind: { ALERTS_TOPIC_ARN: ... } }). - Set
observability.alertsinmantle.config.tswithmode: 'cost-optimized'anderrorLogNotifier: true. - Run
npx mantle build— the generator emits the observability module block + theaws_cloudwatch_log_account_policy+aws_lambda_permission. If the alarm-metric count would exceed 10, the build throws with the count and a fix suggestion. - Run
npx mantle check observabilityto verify (also catches future hand-edits toinfra/observability.tf). - Deploy:
npx mantle deploy --stage staging. Confirm ≤10 alarms in the AWS console; trigger a test error and verify the notifier email arrives.
Related
- C144:
mantle check observability— the machine check (HIGH, blocking) - C145: Cost-free alerting architecture explainer (MEDIUM, stop-hook review)
- C109: CloudWatch cost guard (pre-existing, medium)
packages/cli/src/generators/infra-observability.ts—countAlarmMetrics()+generateObservabilityTf()packages/cli/src/analysis/observability-alarm-count.ts— file-parser for ejected TFmodules/observability/— the Terraform module (already supports selectivelambda_function_names)