773 free skills
Observability skills
Skills for observability — structured logging, distributed tracing, metrics dashboards, alerting, and incident response runbooks.
Sourced from real, public repositories — synced daily, never invented.
773 free skills
Skills for observability — structured logging, distributed tracing, metrics dashboards, alerting, and incident response runbooks.
Sourced from real, public repositories — synced daily, never invented.
15 tools across six categories
13 of them never send your data anywhere
Free · No signup · No trial clock
SEE THE DIRECTORY

reliability-engineering
miles990/claude-software-skills
SRE principles, observability, and incident management
incident-management
dralgorhythm/claude-agentic-framework
Handle production incidents effectively. Use when responding to outages, conducting post-mortems, or improving reliability. Covers incident response and blameless culture.
hedgefundmonitor
leonchaox/qinyan-academic-skills
Query the OFR (Office of Financial Research) Hedge Fund Monitor API for hedge fund data including SEC Form PF aggregated statistics, CFTC Traders in Financial Futures, FICC Sponsored Repo volumes, and FRB SCOOS dealer financing terms. Access time series data on hedge fund size, leverage, counterparties, liquidity, complexity, and risk management. No API key or registration required. Use when working with hedge fund data, systemic risk monitoring, financial stability research, hedge fund leverage or leverage ratios, counterparty concentration, Form PF statistics, repo market data, or OFR financial research data.
alterlab-hedgefund-monitor
alterlab-ieu/alterlab-academic-skills
Queries the OFR (Office of Financial Research) Hedge Fund Monitor API for time series on hedge fund size, leverage, counterparties, liquidity, complexity, and risk management, including SEC Form PF aggregated statistics, CFTC Traders in Financial Futures, FICC Sponsored Repo volumes, and FRB SCOOS dealer financing terms (no API key or registration required). Use when working with hedge fund data, systemic risk monitoring, financial stability research, hedge fund leverage or leverage ratios, counterparty concentration, Form PF statistics, repo market data, or OFR financial research data. Part of the AlterLab Academic Skills suite.
error-handling
pluginagentmarketplace/custom-plugin-nodejs
Implement robust error handling in Node.js with custom error classes, async patterns, Express middleware, and production monitoring
docs-decommission
laurigates/claude-plugins
Generate DECOMMISSION-<service>.md covering infra, data, access, DNS, dependencies, monitoring, and financial checklists. Use when decommissioning a service.
web-monitor
rogue-agent1/web-monitor
Monitor web pages for content changes and get alerts. Track URLs, detect updates, view diffs. Use when asked to watch a website, track changes on a page, monitor for new posts/content, set up page change alerts, or check if a site has been updated. Supports CSS selectors for targeted monitoring.
managing-incidents
ancoleman/ai-design-components
Guide incident response from detection to post-mortem using SRE principles, severity classification, on-call management, blameless culture, and communication protocols. Use when setting up incident processes, designing escalation policies, or conducting post-mortems.
monitoring-whale-activity
jeremylongshore/claude-code-plugins-plus-skills
Track large cryptocurrency transactions and whale wallet movements in
performing-paste-site-monitoring-for-credentials
mukul975/anthropic-cybersecurity-skills
Monitor paste sites like Pastebin and GitHub Gists for leaked credentials,
incident-response
hieutrtr/ai1-skills
>-
prometheus
julianobarbosa/claude-code-skills
Query and interact with Prometheus HTTP API for monitoring data. Use when Claude needs to query Prometheus metrics, execute PromQL queries, retrieve targets/alerts/rules status, access metadata about series/labels, manage TSDB operations, or troubleshoot monitoring infrastructure. Supports instant queries, range queries, metadata endpoints, admin APIs, and alerting information.
trace
simota/agent-skills
Analyzing session replays, extracting persona-based behavioral patterns, and storytelling UX issues. Reads the 'why' from real user operation logs. Works with Field/Echo for persona validation.
implementing-gcp-vpc-firewall-rules
mukul975/anthropic-cybersecurity-skills
Implements and audits GCP VPC firewall rules using gcloud, covering auditing overly permissive rules, creating restrictive ingress/egress rules, hierarchical firewall policies, and monitoring rule effectiveness with VPC Flow Logs. Use when deploying GCP workloads needing network access controls, auditing firewall configs, or responding to Security Command Center findings; not for Cloud Armor or DNS-based filtering.
implementing-privileged-session-monitoring
mukul975/anthropic-cybersecurity-skills
Implements privileged session monitoring and recording using PAM
ibkr-cli
fatwang2/ibkr-cli
Guide users through Interactive Brokers CLI operations — from installing IB Gateway/TWS and ibkr-cli itself, to trading stocks, monitoring accounts, retrieving market data, reading financial news, exploring options chains, screening stocks with market scanners, querying company fundamentals, viewing historical trades, checking P&L, reviewing fund transfers, and tracking dividends. Use this skill whenever the user mentions Interactive Brokers, IBKR, TWS, IB Gateway, stock trading via CLI, checking portfolios or positions, getting quotes, placing orders, reading stock news, options chain, greeks, stock screener, market scanner, top gainers, most active, company fundamentals, financial statements, balance sheet, income statement, cash flow, ownership, institutional holders, trade history, realized P&L, unrealized P&L, fund transfers, deposits, withdrawals, dividends, interest, or anything related to brokerage account management through a terminal. Even if the user doesn't say "ibkr" explicitly, trigger when they want to buy/sell stocks from the command line, check their brokerage account, read news about a stock, look up options data, screen for stocks, check company financials, view trade history, check profit and loss, review fund movements, or set up an API connection to a broker.
detecting-dnp3-protocol-anomalies
mukul975/anthropic-cybersecurity-skills
Detect anomalies in DNP3 communications used in SCADA/ICS systems by monitoring unauthorized control commands, firmware update attempts, protocol violations, and deviations from baseline traffic using deep packet inspection and machine learning approaches. Use when securing energy-sector or other OT/ICS networks, investigating suspicious DNP3 master/outstation activity, or building an anomaly-based IDS for industrial control traffic.
signals-scout-observability-gaps
posthog/ai-plugin
>
email-sdk
opencoredev/email-sdk
Use when adding, reviewing, or documenting Email SDK integrations in TypeScript/Bun apps. Dynamically refreshes the current Email SDK docs/source before implementation, then covers adapter selection, fallbacks, CLI smoke tests, hooks, and secret-safe observability.
cloudflare
heyvhuang/ship-faster
Infrastructure operations for Cloudflare: Workers, KV, R2, D1, Hyperdrive, observability, builds, audit logs. Triggers: worker/KV/R2/D1/logs/build/deploy/audit. Three permission tiers: Diagnose (read-only), Change (write requires confirmation), Super Admin (isolated environment). Write operations follow read-first, confirm, execute, verify pattern. MCP is optional — works with Wrangler CLI/Dashboard too.
paperclip-monitor
bbengamin/skills
Produce read-only Paperclip execution reports across issues, agents, heartbeats, approvals, activity, costs, and blocked work. Use when monitoring AFK loops or asking what needs operator attention.
performing-brand-monitoring-for-impersonation
mukul975/anthropic-cybersecurity-skills
Monitor for brand impersonation attacks across domains, social media,
monitoring-setup
eddiebe147/claude-settings
Expert guide for setting up monitoring dashboards, alerting, metrics collection, and observability. Use when implementing application monitoring, setting up alerts, or building dashboards.
china-portfolio-monitoring
jwangkun/claude-for-financial-services-cn
Track and report on A-share portfolio company performance for China-focused private equity funds. Monitors KPIs, financials, and strategic milestones. Adapted from the original portfolio-monitoring skill for Chinese portfolio companies. Triggers on "A股投后管理", "基金投后监控", "portfolio monitoring China", "monitor portfolio company", "投后报告", or "portfolio company review".
exploring-llm-clusters
posthog/ai-plugin
Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.
internal-dev-workbench
vercel/workflow
Spin up a portless + tmux dev session for the Workflow SDK that gives each git worktree isolated `<branch>.<name>.localhost` URLs for the Next.js workbench and the observability UI, plus a Claude statusline that surfaces those URLs. Use only when the user asks for a "portless dev session", a "tmux dev layout for workflow", "worktree-isolated dev URLs", or wants to wire workflow dev URLs into the Claude statusline. Do not activate for the generic "start the dev server" / "run pnpm dev" task.
track-and-trace
kishorkukreja/awesome-supply-chain
When the user wants to implement shipment tracking, product traceability, or supply chain visibility. Also use when the user mentions "tracking," "traceability," "visibility," "serialization," "lot tracking," "batch tracking," "chain of custody," "provenance," "track and trace," or "shipment monitoring." For control towers, see control-tower-design. For compliance, see compliance-management.
ai-product-canvas
mohitagw15856/pm-claude-skills
Structure AI and ML product decisions with the rigour of any product decision. Use when building AI-powered features, evaluating LLM integrations, designing AI products, or assessing AI readiness. Produces a complete AI product canvas covering problem definition, model approach, data requirements, evaluation framework, UX design, responsible AI checklist, and launch monitoring plan.
fiber-logging-and-project-structure
oimiragieo/agent-studio
Applies best practices for logging (zerolog/zap), project structure (cmd/internal/pkg), middleware registration order, and environment configuration in Go Fiber v2/v3 applications.
holmesgpt
julianobarbosa/claude-code-skills
Guide for implementing HolmesGPT - an AI agent for troubleshooting cloud-native environments. Use when investigating Kubernetes issues, analyzing alerts from Prometheus/AlertManager/PagerDuty, performing root cause analysis, configuring HolmesGPT installations (CLI/Helm/Docker), setting up AI providers (OpenAI/Anthropic/Azure), creating custom toolsets, or integrating with observability platforms (Grafana, Loki, Tempo, DataDog).
great-expectations
majesticlabs-dev/majestic-marketplace
Data validation using Great Expectations. Expectation suites, checkpoints, and data docs for pipeline monitoring.
mcp-cloudflare
heyvhuang/ship-faster
Manage Workers/KV/R2/D1/Hyperdrive via Cloudflare MCP, perform observability/build troubleshooting/audit/container sandbox operations. Triggers: worker/KV/R2/D1/logs/build/deploy/screenshot/audit/sandbox. Three permission tiers: Diagnose (read-only), Change (write requires confirmation), Super Admin (isolated environment). Write operations must follow read-first, user confirmation, post-execution verification.
pr-media-monitoring
asgard-ai-platform/skills
Set up and conduct media monitoring to track brand mentions, sentiment, and share of voice across news, social, and online channels. Use this skill when the user needs to track what's being said about their brand, monitor competitors' media presence, detect emerging PR issues early, or measure campaign reach — even if they say 'what are people saying about us', 'monitor our brand mentions', 'track competitor PR', or 'set up media alerts'.
dibbla
dibbla-agents/skills
Dibbla CLI for scaffolding projects, deploying apps, managing apps/databases/secrets and multi-service `dibbla.yaml` manifests, authoring and operating workflows (slim YAML; `wf execute` sync/`--async`/`--follow`; `wf logs <runId>`, `wf runs list/output` for run monitoring), building Go workers via `github.com/dibbla-agents/sdk-go`, and running `dibbla-task.yaml` pipelines (`dibbla run`).
blogwatcher
alphaonedev/openclaw-graph
Content monitoring: new post detection, RSS feeds, web scraping triggers, change alerts
runbook
julianobarbosa/claude-code-skills
Create or load an operational runbook for a given topic. Searches `runbooks/` for an existing match; if none, scaffolds a new one from the standard template (Purpose / Prerequisites / Steps / Verification / Troubleshooting / Last Tested). Use when asked to "create a runbook", "load runbook for X", document a procedure, or look up an SOP.
robusta-dev
julianobarbosa/claude-code-skills
Robusta Kubernetes observability and alert automation platform. USE WHEN installing Robusta OR configuring playbooks OR setting up notification sinks OR troubleshooting Kubernetes alerts OR creating custom actions OR integrating with Prometheus/AlertManager OR automating incident remediation.
RobustaDev
julianobarbosa/claude-code-skills
Robusta Kubernetes observability and alert automation platform. USE WHEN installing Robusta OR configuring playbooks OR setting up notification sinks OR troubleshooting Kubernetes alerts OR creating custom actions OR integrating with Prometheus/AlertManager OR automating incident remediation.
holmesgpt-skill
julianobarbosa/claude-code-skills
Guide for implementing HolmesGPT - an AI agent for troubleshooting cloud-native environments. Use when investigating Kubernetes issues, analyzing alerts from Prometheus/AlertManager/PagerDuty, performing root cause analysis, configuring HolmesGPT installations (CLI/Helm/Docker), setting up AI providers (OpenAI/Anthropic/Azure), creating custom toolsets, or integrating with observability platforms (Grafana, Loki, Tempo, DataDog).
gke-observability
googlecloudplatform/gke-mcp
Workflows for setting up and auditing observability (logging, monitoring, tracing) on GKE.
signals-scout-ai-observability
posthog/ai-plugin
>
add-analytics
tushaarmehtaa/tushar-skills
Set up PostHog analytics, Sentry error tracking, and health endpoints for web apps. Use when adding or auditing product analytics, monitoring, or uptime checks.
pubnub-scale
pubnub/skills
Scale PubNub applications for high-volume real-time events using channel groups, wildcard subscriptions, sharding, and large-event readiness. Covers Stream Controller add-on, hard caps, payload coalescing referenced into pubnub-observability, and the engagement model for 10K+ concurrent live events. Persistence/history is owned by pubnub-history.
aeo-scorecard
majesticlabs-dev/majestic-marketplace
Measurement framework for Answer Engine Optimization (AEO). Provides AI visibility metrics, share of voice tracking, citation monitoring, and referral demand measurement. Use when discussing AEO/GEO metrics or AI visibility performance.
airflow
alphaonedev/openclaw-graph
Open-source platform for authoring, scheduling, and monitoring data pipelines programmatically.
launch
fcakyon/phd-skills
Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup. Use when the user asks to launch, kick off, start, restart, or kill a training run, or mentions launching a multi-hour or multi-day GPU job (python train, accelerate launch, torchrun, deepspeed, sbatch, tmux training).
monitor
whawkinsiv/solo-founder-skills
Use this skill when the user needs to set up production monitoring, track app health, configure error alerts, or respond to incidents. Also use when the user says 'my app went down,' 'how do I know if something breaks,' 'set up alerts,' 'is my app healthy,' or 'I found out from a user that my site was down.' Covers error tracking, uptime monitoring, performance metrics, and incident response for SaaS applications.
sentry-setup-ai-monitoring
getsentry/sentry-for-claude
Setup Sentry AI Agent Monitoring in any project. Use when asked to monitor LLM calls, track AI agents, track conversations, or instrument OpenAI/Anthropic/Vercel AI/LangChain/Google GenAI/Pydantic AI/Laravel AI. Detects installed AI SDKs and configures appropriate integrations.
build-dashboard
openai/role-specific-plugins
Build source-backed dashboards for monitoring performance, exploring drivers, or acting on product and business metrics. Use when the task needs a dashboard, scorecard, or monitoring view with clear metrics, filters, source definitions, and QA.
didit-transaction-monitoring
didit-protocol/skills
>
agent-tools
blogic-cz/agent-tools
LOAD THIS SKILL when: using CLI wrapper tools (gh-tool, observability-tool, db-tool, k8s-tool, az-tool, azdo-tool, logs-tool, session-tool), working with observability, databases, GitHub PRs, Kubernetes, Azure platform resources, Azure DevOps, or application logs. Contains tool overview, usage patterns, and project-specific aliases.
algo-mfg-spc
asgard-ai-platform/skills
Implement Statistical Process Control charts to monitor production process stability. Use this skill when the user needs to detect process shifts, set control limits, or distinguish common cause from special cause variation — even if they say 'process monitoring', 'control chart', or 'is our process in control'.
implementing-ot-incident-response-playbook
mukul975/anthropic-cybersecurity-skills
Develops OT-specific incident response playbooks using a SANS PICERL-based Python
deliverability-incident-response
growthenginenowoslawski/coldoutboundskills
Triage playbook for when cold email deliverability breaks. Decision-tree guidance for "I landed in spam", "bounce rate spiked", "domain blacklisted", "inbox blocked in warmup", "reply rate dropped". Tells you what to check first, what to fix, and how long the fix takes. Pair with /email-deliverability-audit for diagnosis.
cloudflare-workers-otel-utels
mizchi/skills
Cloudflare Worker telemetry at the fetch boundary — OTLP traces / metrics / logs + utels error tracking + D1 Proxy that emits slow-query warnings. Use when adding observability to a Worker without touching handler code.
incident-response
nickcrew/claude-cortex
Incident triage, cascade prevention, and postmortem methodology. Use when handling production incidents, designing resilience patterns, or conducting chaos engineering exercises.
runbook-writer
mohitagw15856/pm-claude-skills
Write an operational runbook for a service, incident type, or deployment procedure. Use when asked to write a runbook, create an ops guide, document an operational procedure, or prepare an incident response playbook. Produces a runbook with overview, prerequisites, step-by-step procedures, rollback steps, troubleshooting table, and escalation paths.
callingly
membranedev/application-skills
|
kafka-dlq-review
lensesio/agentic-engineering-for-apache-kafka
Review dead letter queue implementations for completeness using the Lenses MCP server. Checks DLQ topic existence, configuration, monitoring, metadata preservation, retry logic, reprocessing paths and connector DLQ alignment. Use when user says "review dead letter queues", "check DLQ setup", "DLQ audit" or asks about error handling, message failures or reprocessing. Do NOT use for reprocessing DLQ messages or managing consumer offsets.
android-performance-observability
krutikjain/android-agent-skills
Measure startup, rendering, memory, jank, vitals, logs, and crash signals for Android apps with actionable traces.