Skip to content

Observability Alerting — Two-Tier Cost-Free Design ​

Applies to: all Mantle instances · Convention checks: C144 (free-tier cap, HIGH), C145 (architecture guidance, MEDIUM)


The billing trap: why "aggregate" mode saves nothing ​

CloudWatch bills metric-math alarms per referenced metric, not per alarm object. The framework's observability module has an enable_per_function_lambda_alarms = false path that collapses N per-function alarms into 2 aggregate metric-math alarms — but each of those 2 alarms still references all N Lambda metrics. The billing line item is 2 × N alarm-metrics, identical to per-function mode.

Composite alarms make it worse: $0.50/alarm/month each, on top of the children they reference. Never composite.

The only genuine cost levers are:

  1. Keep the total alarm-metric count at or below 10 (always free per account).
  2. Route the bulk signal (Lambda errors) to a log subscription filter instead of alarms — filters are free.

Two-tier design ​

Tier A — free-tier alarm set (≤10 alarm-metrics, $0) ​

Generated by mantle generate infra in cost-optimized mode (default). Budget is spent on the highest-value metric-only signals:

SignalAlarm-metrics consumedWhy metric-only
API Gateway 5XX1No log line for synchronous failures
DLQ depth1 per DLQQueue-depth is a metric; no error log
Lambda Errors + Throttles2 per criticalFunction entryThrottles are metric-only; errors also covered by Tier B

Tier B — log-error notifier ($0) ​

One account-level CloudWatch Logs subscription filter (aws_cloudwatch_log_account_policy, policy_type = "SUBSCRIPTION_FILTER_POLICY") with filter pattern { $.level = "ERROR" }, delivering to a framework LogNotifier standalone Lambda → existing SNS email topic. The policy applies account-wide. SUBSCRIPTION_FILTER_POLICY supports selectionCriteria only as LogGroupName NOT IN [...] (exclusion — not LIKE/positive IN); we omit it because AWS automatically excludes a Lambda destination's own log group from account-level subscription filters (no recursion), and the filter pattern is the effective scope: only structured Powertools ERROR logs match. The handler also skips CloudWatch CONTROL_MESSAGE health-check frames (only DATA_MESSAGE is processed). Each instance runs in a dedicated AWS account.

Coverage: every Lambda's errors, including future functions auto-discovered on deploy. Delivery is event-driven (AWS-managed, sub-second latency). Cost: $0 — account-level filter is free (one per account/region), notifier invocations sit inside the 1M free tier, SNS publish and email delivery are free at this scale.


Alarm-metric count formula (C144) ​

text
alarmMetrics = 2 × lambda_function_names.length
             + sqs_dlq_names.length
             + (api_gateway_name ? 1 : 0)
             + 2 × eventbridge_rule_names.length
             + sqs_age_queue_names.length
             + custom_alarms.length

This formula is the single source of truth implemented in countAlarmMetrics() in packages/cli/src/generators/infra-observability.ts and imported by packages/cli/src/analysis/observability-alarm-count.ts. Two enforcement points:

  • Generator throw (mantle build / mantle generate infra): fires when the formula exceeds 10 during code generation.
  • mantle check observability: parses infra/observability.tf on disk, catches ejected or hand-edited files that bypass the generator.

The 2 × lambda_function_names.length term holds regardless of enable_per_function_lambda_alarms — do not attempt to "optimize" this by looking at alarm object count.


Config surface ​

ts
defineConfig({
  observability: {
    alerts: {
      email: "ops@example.com",

      // 'cost-optimized' (default): per-function alarms limited to criticalFunctions;
      // all-function errors covered by Tier B log notifier.
      // 'per-function': all discovered Lambdas get per-function alarms (same cap applies).
      mode: "cost-optimized",

      // Lambda names that get dedicated Errors + Throttles alarms.
      // Each entry consumes 2 of the 10 free alarm-metric slots.
      criticalFunctions: ["UserLogin", "FilesGet"],

      // Enable the Tier B log-error notifier. Default: true in cost-optimized mode.
      // Requires a LogNotifier standalone Lambda in src/lambdas/standalone/.
      errorLogNotifier: true,
    },
  },
});

Budget planning example (OMD): 5 DLQs + 1 API Gateway = 6 alarm-metrics consumed → 4 remaining → at most 2 criticalFunctions (2 × 2 = 4). Total: 10. Errors on all 20 functions covered by Tier B.


Coverage tradeoff (accepted) ​

Lambda throttles on non-critical functions are not covered. Throttles are metric-only — no log line is emitted, so Tier B cannot see them. Covering all N functions' throttles via metrics costs money. The accepted approach: cover throttles on the 1–2 busiest functions via criticalFunctions and accept the gap for low-traffic functions. If comprehensive throttle coverage is later needed, a scheduled watcher Lambda using GetMetricStatistics (free API tier) can close the gap.


Instance setup ​

  1. Add src/lambdas/standalone/LogNotifier/index.ts — a thin handler that re-exports createLogSubscriptionNotifier from @j0nathan-ll0yd/observability via defineLambda({ timeout: 30, bind: { ALERTS_TOPIC_ARN: ... } }).
  2. Set observability.alerts in mantle.config.ts with mode: 'cost-optimized' and errorLogNotifier: true.
  3. Run npx mantle build — the generator emits the observability module block + the aws_cloudwatch_log_account_policy + aws_lambda_permission. If the alarm-metric count would exceed 10, the build throws with the count and a fix suggestion.
  4. Run npx mantle check observability to verify (also catches future hand-edits to infra/observability.tf).
  5. Deploy: npx mantle deploy --stage staging. Confirm ≤10 alarms in the AWS console; trigger a test error and verify the notifier email arrives.

  • C144: mantle check observability — the machine check (HIGH, blocking)
  • C145: Cost-free alerting architecture explainer (MEDIUM, stop-hook review)
  • C109: CloudWatch cost guard (pre-existing, medium)
  • packages/cli/src/generators/infra-observability.ts — countAlarmMetrics() + generateObservabilityTf()
  • packages/cli/src/analysis/observability-alarm-count.ts — file-parser for ejected TF
  • modules/observability/ — the Terraform module (already supports selective lambda_function_names)