ml-ai

GitHub

提供Grafana Cloud的AI/ML功能,包括自然语言转查询、动态异常检测告警、自动化根因分析及LLM插件集成,支持智能监控与故障排查。

skills/grafana-cloud/ml-ai/SKILL.md grafana/skills

Trigger Scenarios

用户希望使用自然语言生成PromQL或LogQL查询 需要配置基于机器学习的动态告警而非静态阈值 请求自动进行事故根因分析(RCA)或Sift调查 需要将LLM(如Claude/OpenAI)集成到Grafana中

Install

npx skills add grafana/skills --skill ml-ai -g -y
More Options

Non-standard path

npx skills add https://github.com/grafana/skills/tree/main/skills/grafana-cloud/ml-ai -g -y

Use without installing

npx skills use grafana/skills@ml-ai

指定 Agent (Claude Code)

npx skills add grafana/skills --skill ml-ai -a claude-code -g -y

安装 repo 全部 skill

npx skills add grafana/skills --all -g -y

预览 repo 内 skill

npx skills add grafana/skills --list

SKILL.md

Frontmatter
{
    "name": "ml-ai",
    "license": "Apache-2.0",
    "description": "Turn on AI + ML features in Grafana Cloud — Grafana Assistant (NL → PromQL\/LogQL\/TraceQL, dashboard build, incident investigation, MCP integration), Dynamic Alerting (Prophet forecasting + DBSCAN outlier detection), Sift (8-analysis automated root-cause), Knowledge Graph + RCA Workbench, and the LLM Plugin (OpenAI \/ Anthropic \/ Azure \/ Ollama \/ vLLM \/ LiteLLM). Use when you want anomaly alerts without static thresholds, natural-language querying, automated incident investigation, dashboards generated from a sentence, or a managed LLM proxy for plugins — even when the user says \"alert when something looks weird\", \"explain this PromQL\", \"find the root cause\", \"make this a dashboard\", or \"wire Claude into Grafana\" without naming any of these products."
}

Grafana Cloud AI & ML

Docs: https://grafana.com/docs/grafana-cloud/alerting-and-irm/machine-learning/

ML alerting + automated RCA + LLM-powered Assistant in one Grafana Cloud stack.

Prerequisites

  • Grafana Cloud stack (Pro / Advanced — most features GA, some in preview)
  • API token with plugins:write for ML / Sift / LLM-plugin endpoints
  • For Dynamic Alerting: at least 14 days (ideally 90d) of history for the metric you want to forecast

Common Workflows

1. Forecasting alert with Dynamic Alerting

# 1. Create forecast job (Prophet — learns daily/weekly seasonality)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/forecast \
  -H "Authorization: Bearer <token>" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "cpu-forecast",
    "metric": "avg(rate(node_cpu_seconds_total{mode=\"user\"}[5m]))",
    "datasourceId": 1,
    "interval": 300,
    "trainingWindow": "90d",
    "forecastWindow": "7d",
    "algorithm": { "name": "prophet", "config": {} }
  }'

# 2. Verify job is producing the predicted-value metric (may take a few minutes).
#    <datasourceId> must match the datasourceId used above (find it via
#    GET /api/datasources), or run the query from Explore instead.
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_forecast_upper{job="cpu-forecast"}' \
  | jq '.data.result | length'
# Expect > 0

# 3. Add an alert that fires when actual exceeds the upper bound
# expr:  avg(rate(node_cpu_seconds_total{mode="user"}[5m]))
#         > ml_forecast_upper{job="cpu-forecast"} * 1.1

2. Outlier alert — one service deviates from peers

# 1. Create outlier job (DBSCAN — groups peers, flags the odd one)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-ml-app/resources/ml/v1/outlier \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{
    "name": "service-error-outliers",
    "metric": "sum(rate(http_requests_total{status=~\"5..\"}[5m])) by (service)",
    "datasourceId": 1,
    "interval": 300,
    "algorithm": { "name": "dbscan", "sensitivity": 0.5, "config": { "epsilon": 0.5 } }
  }'

# 2. Verify the score metric exists (<datasourceId> must match the
#    datasourceId used above, or run the query from Explore instead)
curl -s -H "Authorization: Bearer <token>" \
  'https://<stack>.grafana.net/api/datasources/proxy/<datasourceId>/api/v1/query?query=ml_outlier_score{job="service-error-outliers"}' \
  | jq '.data.result | length'

# 3. Alert when ml_outlier_score{job="service-error-outliers"} > 0.8 for 5m

3. Run a Sift investigation

# 1. Trigger from API (or from Explore / Incident / OnCall)
curl -X POST https://<stack>.grafana.net/api/plugins/grafana-sift-app/resources/sift/v1/investigations \
  -H "Authorization: Bearer <token>" -H "Content-Type: application/json" \
  -d '{ "name":"checkout-spike","start":"2024-02-01T10:00:00Z","end":"2024-02-01T10:30:00Z",
        "filters":{"service":"checkout","namespace":"production"} }'

# 2. The response includes an investigation ID — open it in the UI:
#    https://<stack>.grafana.net/a/grafana-sift-app/investigations/<id>
# 3. Verify analyses ran — each of the 8 checks shows ✔ or ✖ with linked evidence.

See references/sift.md for the full 8-analysis table.

4. Wire up the LLM Plugin

# 1. Provision (provisioning/plugins/llm.yaml — see references/llm-and-graph.md)
apiVersion: 1
apps:
  - type: grafana-llm-app
    jsonData: { openAIUrl: https://api.openai.com, openAIModel: gpt-4o }
    secureJsonData: { openAIKey: sk-... }
# 2. Restart Grafana, then verify the health endpoint reports the configured provider
curl -s -H "Authorization: Bearer <token>" \
  https://<stack>.grafana.net/api/plugins/grafana-llm-app/health | jq
# Expect: {"status":"ok", ...}

# 3. Verify in a panel — open any panel, click the Assistant icon, ask "what does this query do?"

See references/llm-and-graph.md for Assistant capabilities, Knowledge Graph search syntax, and Adaptive Metrics recommendations.

Resources

Version History

  • b583762 Current 2026-07-06 00:35

Same Skill Collection

skills/grafana-app-sdk/admission-control/SKILL.md
skills/grafana-cloud/loki-label-analyzer/SKILL.md
skills/grafana-cloud/send-data/SKILL.md
skills/grafana-datasources/datasources-provisioning/SKILL.md
skills/grafana-lgtm/loki/SKILL.md
skills/grafana-lgtm/prometheus/SKILL.md
skills/grafana-plugins/audit-and-reduce-dependencies/SKILL.md
skills/grafana-plugins/check-npm/SKILL.md
template/SKILL.md
skills/grafana-app-sdk/app-sdk-concepts/SKILL.md
skills/grafana-app-sdk/cue-kind-definition/SKILL.md
skills/grafana-app-sdk/reconciler-logic/SKILL.md
skills/grafana-cloud/adaptive-metrics/SKILL.md
skills/grafana-cloud/admin/SKILL.md
skills/grafana-cloud/app-observability/SKILL.md
skills/grafana-cloud/assistant-mcp/SKILL.md
skills/grafana-cloud/cloud-integrations/SKILL.md
skills/grafana-cloud/cost-management/SKILL.md
skills/grafana-cloud/database-observability/SKILL.md
skills/grafana-cloud/dpm-finder/SKILL.md
skills/grafana-cloud/fleet-management/SKILL.md
skills/grafana-cloud/infrastructure/SKILL.md
skills/grafana-cloud/oncall-irm/SKILL.md
skills/grafana-cloud/private-connectivity/SKILL.md
skills/grafana-cloud/prometheus-cardinality-troubleshooter/SKILL.md
skills/grafana-cloud/prometheus-label-strategy/SKILL.md
skills/grafana-cloud/synthetic-monitoring-checks/SKILL.md
skills/grafana-cloud/testing/SKILL.md
skills/grafana-core/alerting-irm/SKILL.md
skills/grafana-core/alloy/SKILL.md
skills/grafana-core/beyla/SKILL.md
skills/grafana-core/dashboarding/SKILL.md
skills/grafana-core/grafana-oss/SKILL.md
skills/grafana-core/opentelemetry/SKILL.md
skills/grafana-core/promql/SKILL.md
skills/grafana-core/skill-authoring/SKILL.md
skills/grafana-k6/k6-cloud-investigate-test/SKILL.md
skills/grafana-k6/k6-docs/SKILL.md
skills/grafana-k6/k6-manage/SKILL.md
skills/grafana-k6/k6-perf-test-website/SKILL.md
skills/grafana-k6/k6-test-maintenance/SKILL.md
skills/grafana-k6/k6-trend-analysis/SKILL.md
skills/grafana-k6/k6/SKILL.md
skills/grafana-lgtm/mimir/SKILL.md
skills/grafana-lgtm/pyroscope/SKILL.md
skills/grafana-lgtm/tempo/SKILL.md
skills/grafana-plugins/grafana-scenes/SKILL.md
skills/grafana-plugins/plugin-bundle-size/SKILL.md
skills/grafana-plugins/react-19-plugin-migration/SKILL.md

Metadata

Files
0
Version
80bb293
Hash
446880b9
Indexed
2026-07-06 00:35

Главная - Вики-сайт
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-04 05:25
浙ICP备14020137号-1 $Гость$