Agent Skillsgrafana/skills › alerting-irm

alerting-irm

GitHub

用于配置Grafana告警、事件响应管理(IRM)及SLO的全流程技能。涵盖告警规则创建、通知路由、联系人设置、值班轮班及SLO定义,支持通过API或YAML进行基础设施即代码管理。

skills/grafana-core/alerting-irm/SKILL.md grafana/skills

Trigger Scenarios

配置告警规则 调试通知路由 设置值班轮班 管理 incidents 定义 SLO 用户要求报警或页面通知

Install

npx skills add grafana/skills --skill alerting-irm -g -y
More Options

Non-standard path

npx skills add https://github.com/grafana/skills/tree/main/skills/grafana-core/alerting-irm -g -y

Use without installing

npx skills use grafana/skills@alerting-irm

指定 Agent (Claude Code)

npx skills add grafana/skills --skill alerting-irm -a claude-code -g -y

安装 repo 全部 skill

npx skills add grafana/skills --all -g -y

预览 repo 内 skill

npx skills add grafana/skills --list

SKILL.md

Frontmatter
{
    "name": "alerting-irm",
    "license": "Apache-2.0",
    "description": "Configure Grafana Alerting, Incident Response Management (IRM), and SLOs end-to-end — provisions Grafana-managed and data-source-managed alert rules, contact points (Slack\/PagerDuty\/email\/webhook), notification policies with hierarchical matchers, silences, mute timings, on-call schedules and escalation chains, incident-management integrations, and SLOs with multi-window burn-rate alerts. Use when configuring alerts, debugging notification routing, setting up on-call rotations, declaring or managing incidents, defining SLOs, provisioning alerting via YAML or API, picking matchers for a notification policy, building a PagerDuty\/Slack webhook receiver, or troubleshooting why an alert isn't firing — even when the user says \"page me on errors\", \"alert me when X happens\", \"route this to the platform team\", or \"set up an SLO\" without naming Alerting or IRM."
}

Grafana Alerting & IRM

Docs: https://grafana.com/docs/grafana/latest/alerting.md

Common Workflows

Provisioning a new alert end-to-end

  1. Create contact points (where notifications go):

    curl -X POST https://grafana.example.com/api/v1/provisioning/contact-points \
      -H 'Authorization: Bearer <token>' -H 'Content-Type: application/json' \
      -d @contact-points.json
    

    Verify:

    curl https://grafana.example.com/api/v1/provisioning/contact-points \
      -H 'Authorization: Bearer <token>' | jq '.[].name'
    
  2. Add notification policies (which alerts go where) — see § Notification policies below for the matchers pattern.

  3. Write the alert rule — pick the type:

  4. Verify routing before going live:

    # Force-fire a test alert from the rule's UI, then check Alertmanager's view
    curl https://grafana.example.com/api/alertmanager/grafana/api/v2/alerts \
      -H 'Authorization: Bearer <token>' | jq '.[] | {alertname: .labels.alertname, receiver: .receivers}'
    

    The expected receiver should appear. If the wrong receiver appears, re-check the policy's matchers.

Routing alerts to IRM / on-call

  1. In IRM, create an Integration of type "Grafana Alerting webhook" → copy the integration URL
  2. Add a webhook contact point in Grafana Alerting pointing at that URL (full YAML in references/irm.md § Routing)
  3. Add a notification policy matcher routing the right severity to the new contact point
  4. Verify: trigger a test alert; it should appear in IRM within ~30s. Full debug procedure in references/irm.md § Verifying the IRM integration.

Defining an SLO

  1. Create the SLO via UI or API → Grafana auto-generates recording rules, dashboards, and burn-rate alerts (the generated YAML is in references/slo.md)
  2. Use multi-window burn-rate alerts, not single-window — see references/slo.md § Multi-window burn-rate alerts for why single-window fires on noise
  3. Verify with the 4-step pattern in references/slo.md § Validating SLO config

Contact Points (YAML provisioning)

# provisioning/alerting/contact_points.yaml
apiVersion: 1
contactPoints:
  - orgId: 1
    name: pagerduty-critical
    receivers:
      - uid: pd-receiver
        type: pagerduty
        settings:
          integrationKey: YOUR_PAGERDUTY_KEY
          severity: critical

  - orgId: 1
    name: slack-alerts
    receivers:
      - uid: slack-receiver
        type: slack
        settings:
          url: https://hooks.slack.com/services/YOUR/WEBHOOK/URL
          channel: '#alerts'

For email, webhook, Teams, Telegram, OnCall, and other receiver types, see references/alerting.md § Contact point receiver types.

Notification policies

Hierarchical routing tree with label matchers:

# provisioning/alerting/notification_policies.yaml
apiVersion: 1
policies:
  - orgId: 1
    receiver: default-receiver
    group_by: ['alertname', 'cluster', 'service']
    group_wait: 30s
    group_interval: 5m
    repeat_interval: 12h
    routes:
      # Critical alerts → PagerDuty
      - receiver: pagerduty-critical
        matchers:
          - severity = critical
        group_wait: 10s
        repeat_interval: 4h

      # Platform team → Slack, but page on critical
      - receiver: slack-alerts
        matchers:
          - team = platform
        routes:
          - receiver: pagerduty-critical
            matchers:
              - severity = critical

      # Everything else → email
      - receiver: email-alerts
        matchers:
          - severity =~ "warning|info"

Silences

Suppress notifications for matching alerts without stopping evaluation:

curl -X POST https://grafana.example.com/api/alertmanager/grafana/api/v2/silences \
  -H 'Authorization: Bearer <token>' \
  -H 'Content-Type: application/json' \
  -d '{
    "matchers": [
      {"name": "alertname", "value": "HighErrorRate", "isRegex": false},
      {"name": "env", "value": "staging", "isRegex": false}
    ],
    "startsAt": "2024-01-01T00:00:00Z",
    "endsAt": "2024-01-01T02:00:00Z",
    "comment": "Maintenance window",
    "createdBy": "admin"
  }'

# Verify it was created
curl https://grafana.example.com/api/alertmanager/grafana/api/v2/silences \
  -H 'Authorization: Bearer <token>' | jq '.[] | select(.status.state == "active")'

Alert rule states

State Description
Normal Condition not met
Pending Condition met, waiting for for duration
Firing Condition met for full for duration
NoData Query returned no data
Error Query/evaluation error
Recovering Was firing, condition no longer met

Provisioning directory layout

provisioning/alerting/
├── alert_rules.yaml          # Alert and recording rules
├── contact_points.yaml       # Notification destinations
├── notification_policies.yaml  # Routing tree
├── templates.yaml            # Message templates
└── mute_timings.yaml         # Recurring mute windows

API provisioning (keeps UI editable)

Add X-Disable-Provenance: true to keep resources editable in the UI after API provisioning:

curl -X PUT https://grafana.example.com/api/v1/provisioning/policies \
  -H 'Authorization: Bearer <token>' \
  -H 'X-Disable-Provenance: true' \
  -H 'Content-Type: application/json' \
  -d @policy.json

curl -X POST https://grafana.example.com/api/v1/provisioning/alert-rules \
  -H 'Authorization: Bearer <token>' \
  -H 'X-Disable-Provenance: true' \
  -H 'Content-Type: application/json' \
  -d @rule.json

References

  • references/alerting.md — full alert rule YAML (Grafana-managed / Prometheus / Loki) + notification templates
  • references/slo.md — generated SLO recording rules + multi-window burn-rate alert pattern + validation steps
  • references/irm.md — IRM capabilities, integration sources, Alerting → IRM routing + verification + common failure modes

Version History

  • b583762 Current 2026-07-06 00:35

Same Skill Collection

skills/grafana-app-sdk/admission-control/SKILL.md
skills/grafana-cloud/loki-label-analyzer/SKILL.md
skills/grafana-cloud/send-data/SKILL.md
skills/grafana-datasources/datasources-provisioning/SKILL.md
skills/grafana-lgtm/loki/SKILL.md
skills/grafana-lgtm/prometheus/SKILL.md
skills/grafana-plugins/audit-and-reduce-dependencies/SKILL.md
skills/grafana-plugins/check-npm/SKILL.md
template/SKILL.md
skills/grafana-app-sdk/app-sdk-concepts/SKILL.md
skills/grafana-app-sdk/cue-kind-definition/SKILL.md
skills/grafana-app-sdk/reconciler-logic/SKILL.md
skills/grafana-cloud/adaptive-metrics/SKILL.md
skills/grafana-cloud/admin/SKILL.md
skills/grafana-cloud/app-observability/SKILL.md
skills/grafana-cloud/assistant-mcp/SKILL.md
skills/grafana-cloud/cloud-integrations/SKILL.md
skills/grafana-cloud/cost-management/SKILL.md
skills/grafana-cloud/database-observability/SKILL.md
skills/grafana-cloud/dpm-finder/SKILL.md
skills/grafana-cloud/fleet-management/SKILL.md
skills/grafana-cloud/infrastructure/SKILL.md
skills/grafana-cloud/ml-ai/SKILL.md
skills/grafana-cloud/oncall-irm/SKILL.md
skills/grafana-cloud/private-connectivity/SKILL.md
skills/grafana-cloud/prometheus-cardinality-troubleshooter/SKILL.md
skills/grafana-cloud/prometheus-label-strategy/SKILL.md
skills/grafana-cloud/synthetic-monitoring-checks/SKILL.md
skills/grafana-cloud/testing/SKILL.md
skills/grafana-core/alloy/SKILL.md
skills/grafana-core/beyla/SKILL.md
skills/grafana-core/dashboarding/SKILL.md
skills/grafana-core/grafana-oss/SKILL.md
skills/grafana-core/opentelemetry/SKILL.md
skills/grafana-core/promql/SKILL.md
skills/grafana-core/skill-authoring/SKILL.md
skills/grafana-k6/k6-cloud-investigate-test/SKILL.md
skills/grafana-k6/k6-docs/SKILL.md
skills/grafana-k6/k6-manage/SKILL.md
skills/grafana-k6/k6-perf-test-website/SKILL.md
skills/grafana-k6/k6-test-maintenance/SKILL.md
skills/grafana-k6/k6-trend-analysis/SKILL.md
skills/grafana-k6/k6/SKILL.md
skills/grafana-lgtm/mimir/SKILL.md
skills/grafana-lgtm/pyroscope/SKILL.md
skills/grafana-lgtm/tempo/SKILL.md
skills/grafana-plugins/grafana-scenes/SKILL.md
skills/grafana-plugins/plugin-bundle-size/SKILL.md
skills/grafana-plugins/react-19-plugin-migration/SKILL.md

Metadata

Files
0
Version
80bb293
Hash
e23570f7
Indexed
2026-07-06 00:35

Главная - Вики-сайт
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-04 10:32
浙ICP备14020137号-1 $Гость$