499 free skills
Data Engineering skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
499 free skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
15 tools across six categories
13 of them never send your data anywhere
Free · No signup · No trial clock
SEE THE DIRECTORY

sparkpost
membranedev/application-skills
|
airflow-operator-creator
jeremylongshore/claude-code-plugins-plus-skills
Create airflow operator creator operations. Auto-activating skill for
fp-pipe-ref
sickn33/agentic-awesome-skills
Quick reference for pipe and flow. Use when user needs to chain functions, compose operations, or build data pipelines in fp-ts.
data-pipeline-architect
wyattowalsh/agents
>-
whitespark-tool
garrettjsmith/localseoskills
When the user wants citation gap analysis, managed citation building, review generation campaigns, or local rank tracking. Trigger on "Whitespark," "citation finder," "where are my competitors listed," "citation gap," "build citations," "review generation tool," or "get more reviews." Note that Whitespark does NOT have an MCP server — it's dashboard-driven with limited API.
polars
silvainfm/claude-skills
Lightning-fast DataFrame library written in Rust for high-performance data manipulation and analysis. Use when user wants blazing fast data transformations, working with large datasets, lazy evaluation pipelines, or needs better performance than pandas. Ideal for ETL, data wrangling, aggregations, joins, and reading/writing CSV, Parquet, JSON files.
monte-carlo-push-ingestion
sickn33/agentic-awesome-skills
Expert guide for pushing metadata, lineage, and query logs to Monte Carlo from any data warehouse.
nu-shell
knoopx/pi
Reads, filters, transforms, and manipulates CSV/TSV files using Nushell's structured data pipeline. Use when working with tabular data, data cleaning, CSV validation, or batch processing spreadsheet-like files.
sales-rudderstack
sales-skills/sales
RudderStack platform help — warehouse-native CDP, Event Streams, Reverse ETL, Transformations, Profiles, 200+ destinations, open-source. Use when setting up RudderStack SDKs or HTTP API, events not reaching destinations, Reverse ETL syncs failing or returning stale data, transformations throwing errors or dropping events, visitor profiles not merging across channels, choosing between RudderStack and Segment, or working with the RudderStack API. Do NOT use for general CDP comparison (use /sales-cdp) or CRM data dedup without RudderStack (use /sales-data-hygiene).
data-pipeline-spec
mohitagw15856/pm-claude-skills
Design an ETL/ELT data pipeline specification. Use when asked to design a data pipeline, spec an ETL or ELT process, document a data ingestion workflow, or plan a data integration. Produces a complete pipeline spec with sources, transforms, destinations, SLAs, error handling, and data quality rules.
tech-data-pipeline
asgard-ai-platform/skills
Design data pipelines covering ETL vs ELT architectures, data source integration, scheduling, quality checks, and warehouse design. Use this skill when the user needs to move data between systems, build a data warehouse, automate data processing, or improve data reliability — even if they say 'move data from X to Y', 'build an ETL pipeline', 'our data is a mess', or 'set up a data warehouse'.
signals-scout-data-warehouse
posthog/ai-plugin
>
spark-job-creator
jeremylongshore/claude-code-plugins-plus-skills
|
managed-airflow-dag-troubleshooting
google/skills
>-
dbt-test-creator
jeremylongshore/claude-code-plugins-plus-skills
Create dbt test creator operations. Auto-activating skill for Data Pipelines.
machine-learning-ops-ml-pipeline
rmyndharis/antigravity-skills
Design and implement a complete ML pipeline for: $ARGUMENTS
data-engineering
pluginagentmarketplace/custom-plugin-data-engineer
Data pipeline architecture, ETL/ELT patterns, data modeling, and production data platform design
universal-scraping-architect
alirezarezvani/claude-skills
Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.
etl-tools
pluginagentmarketplace/custom-plugin-data-engineer
Apache Airflow, dbt, Prefect, Dagster, and modern data orchestration for production data pipelines
analytics-data-engineer
daemon-blockint-tech/agentic-enteprises-skill
|
huawei-cloud-mrs-spark-sql-check
huaweicloud/huaweicloud-skills
|
spark-app-template
github/copilot-plugins
Comprehensive guidance for building web apps with opinionated defaults for tech stack, design system, and code standards. Use when user wants to create a new web application, dashboard, or interactive interface. Provides tech choices, styling guidance, project structure, and design philosophy to get users up and running quickly with a fully functional, beautiful web app.
etl-pipelines
claude-dev-suite/claude-dev-suite
|
dataset-curator
nickcrew/claude-cortex
Use this skill when designing, cleaning, deduplicating, or documenting datasets for model training and evaluation including schema design, class imbalance handling, and train/val/test splits. Not for running model training or hyperparameter tuning. Not for real-time data pipeline engineering.
pp-setlist-fm
mvanhorn/printing-press-library
Every Setlist.fm endpoint, plus offline analytics no API call can return — tour shape, song frequency, what's overdue, setlist prediction. Trigger phrases: `predict the setlist`, `what songs are overdue for`, `look up setlist for`, `how often does X play Y`, `compare these two tours`, `use setlist-fm`, `run setlist-fm`.
databricks-spark-structured-streaming
databricks-solutions/ai-dev-kit
Comprehensive guide to Spark Structured Streaming for production workloads. Use when building streaming pipelines, working with Kafka ingestion, implementing Real-Time Mode (RTM), configuring triggers (processingTime, availableNow), handling stateful operations with watermarks, optimizing checkpoints, performing stream-stream or stream-static joins, writing to multiple sinks, or tuning streaming cost and performance.
spark-version-upgrade
openhands/extensions
Upgrade Apache Spark applications between major versions (2.x→3.x, 3.x→4.x). Covers build files, deprecated APIs, configuration changes, SQL/DataFrame updates, and test validation.
app-studio-dataset-etl-gen-demo
stahura/domo-ai-vibe-rules
Vertical App Studio micro-demo with optional dataset creation (domo-data-generator), ETL (magic-etl-cli or magic-etl), then multi-page layout using advanced-app-studio + domo-app-theme. Picks industry packs from references/*.md. Use when building an App Studio walkthrough that needs realistic data pipelines and themed pages per vertical.
data-engineering-data-pipeline
aaaaqwq/agi-super-team
You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
ml-pipeline-workflow
rmyndharis/antigravity-skills
Build end-to-end MLOps pipelines from data preparation through model training, validation, and production deployment. Use when creating ML pipelines, implementing MLOps practices, or automating model training and deployment workflows.
jcl-migration-analyzer
dauquangthanh/hanoi-rainbow
Analyzes legacy JCL (Job Control Language) scripts to assist with migration to modern workflow orchestration and batch processing systems. Extracts job flows, step sequences, data dependencies, conditional logic, and program invocations. Generates migration reports and creates implementation strategies for Spring Batch, Apache Airflow, or shell scripts. Use when working with mainframe job migration, JCL analysis, batch workflow modernization, or when users mention JCL conversion, analyzing .jcl/.JCL files, working with job steps, procedures, or planning workflow orchestration from JCL jobs.
valohai-design-pipelines
valohai/valohai-skills
Analyze ML project structure and design Valohai pipelines to orchestrate multi-step workflows. Use this skill when a user wants to create a Valohai pipeline, connect multiple ML steps into an automated workflow, identify pipeline opportunities in their codebase, design data flow between preprocessing/training/evaluation/inference steps, or add conditional logic and parallel execution to pipelines. Triggers on mentions of pipelines, workflows, DAGs, orchestration, multi-step ML, or connecting Valohai steps.
migrate-airflow-kestra
kestra-io/agent-skills
Migrate an Airflow DAG to a production-ready Kestra flow. Extracts Python task logic into namespace files, maps DAG dependencies to Kestra tasks, and preserves parallel execution structure.
spark-optimization
full-stack-skills/devops-skills
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.
iceberg
gordonmurray/data-engineering-skills
Design, migrate, tune, and maintain Apache Iceberg tables across query engines. Use for Iceberg schema and partition evolution, v2 to v3 upgrades, catalog selection, time travel and rollback, row-level deletes, snapshot expiry, orphan file cleanup, small-file compaction, or slow queries and metadata bloat on a lakehouse table. Covers Spark, Flink, Trino, Athena, Snowflake, REST catalogs, Polaris, Nessie, Glue, and Hive.
reverse-etl-activation
swan-gtm/gtm-skills
Use this skill when you need to turn a scored, segmented audience sitting in the warehouse (BigQuery, Snowflake, Redshift) into live GTM motion, syncing accounts, contacts, and computed traits out to the CRM, ad platforms, and sequencing tools. It encodes the judgment that keeps reverse ETL from quietly polluting every downstream system: identity resolution before any write, a change-data-capture diff so only real deltas move, field-level mapping with type and format guards, suppression and consent gates, sync scheduling matched to each destination's rate limits, and a mandatory dry run plus row-count reconciliation before and after the live push. Every threshold (match-confidence floor, batch size, sync cadence, max-delete guard, suppression rules) is a tunable default. Each run produces a validated sync plan, a dry-run diff of adds, updates, and suppressions with counts, and a reconciliation report after the live sync.
building-data-pipelines
legout/data-agent-skills
Build production batch data pipelines with Polars, DuckDB, and PyArrow. Covers ETL patterns, medallion architecture, partitioning, and CRUD operations. Use when designing or implementing data ingestion, transformation, and loading workflows in Python.
kafka
mjunaidca/mjs-agent-skills
|
designing-data-storage
legout/data-agent-skills
File formats and lakehouse table formats for data lakes: Parquet, Arrow, Lance, Zarr, Avro, ORC, Delta Lake, Apache Iceberg, and Apache Hudi. Covers compression, partitioning, ACID transactions, schema evolution, and format selection.
airflow-dag-patterns
rmyndharis/antigravity-skills
Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
spark-python-data-source
databricks-solutions/ai-dev-kit
Build custom Python data sources for Apache Spark using the PySpark DataSource API — batch and streaming readers/writers for external systems. Use this skill whenever someone wants to connect Spark to an external system (database, API, message queue, custom protocol), build a Spark connector or plugin in Python, implement a DataSourceReader or DataSourceWriter, pull data from or push data to a system via Spark, or work with the PySpark DataSource API in any way. Even if they just say "read from X in Spark" or "write DataFrame to Y" and there's no native connector, this skill applies.
spark-advisor
yaooqinn/spark-history-cli
Diagnose, compare, and optimize Apache Spark applications and SQL queries using Spark History Server data. Use this skill whenever the user wants to understand why a Spark app is slow, compare two benchmark runs or TPC-DS results, find performance bottlenecks (skew, GC pressure, shuffle spill, straggler tasks), get tuning recommendations, or optimize Spark/Gluten configurations. Also trigger when the user mentions 'diagnose', 'compare runs', 'why is this query slow', 'tune my Spark job', 'benchmark comparison', 'performance regression', or asks about executor skew, shuffle overhead, AQE effectiveness, or Gluten offloading issues.
spark-history-cli
yaooqinn/spark-history-cli
Query a running Apache Spark History Server from Copilot CLI. Use this whenever the user wants to inspect SHS applications, jobs, stages, executors, SQL executions, environment details, or event logs, especially when they mention Spark History Server, SHS, event log history, benchmark runs, or application IDs.
spark-optimization
rmyndharis/antigravity-skills
Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.
assuring-data-pipelines
legout/data-agent-skills
Data quality validation and observability for data pipelines. Combines Great Expectations and Pandera for data validation with OpenTelemetry and Prometheus for monitoring and alerting.
sqldw-cli
microsoft/skills-for-fabric
Author, query and diagnose Fabric Warehouse, Lakehouse SQL endpoints and Mirrored Databases: DDL/DML and COPY INTO ingestion, read-only T-SQL SELECT and row counts over lakehouse tables, and queryinsights performance triage. Fabric SQL database (OLTP) is sqldb-*-cli. Triggers:query warehouse,count rows lakehouse,SELECT lakehouse,create warehouse table,COPY INTO,warehouse MERGE,slowest warehouse queries,queryinsights CPU
netlify-deploy
jetbrains/skills
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
paimon
gordonmurray/data-engineering-skills
Design, ingest into, tune, and operate Apache Paimon tables for streaming lakehouses. Use for Paimon primary-key or append-only table design, bucket sizing, changelog producer choice, Flink CDC ingestion, compaction backlog, lookup join performance, PyPaimon, Spark reads, Iceberg compatibility, or streaming writes that produce too many small files.
databricks-iceberg
databricks-solutions/ai-dev-kit
Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS Spark, external engine access and credential vending. Use when creating Iceberg tables, enabling External Iceberg Reads (uniform) on Delta tables (including Streaming Tables and Materialized Views via compatibility mode), configuring external engines to read Databricks tables via Unity Catalog IRC, integrating with Snowflake catalog to read Foreign Iceberg tables
spark-best-practices
oleanderhq/skills
>-
orchestrating-data-pipelines
legout/data-agent-skills
Pipeline orchestration and workflow management with Prefect, Dagster, and dbt. Covers scheduling, dependency management, retries, and integration patterns.
mashup
nickcrew/claude-cortex
Force-fit patterns from other domains to spark novel concepts.
pipeline
h-mmer/pentest-agents
Prepare the battlefield — recon, scanning, and surface ranking. Stops before hunting. Run /hunt or /autopilot after. Usage: /pipeline or /pipeline <target>
stream-chain
natea/fitfinder
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
etl-lineage-explainer
sisodiabhumca/agent-skills
Vendor-neutral skill for extracting and summarizing table-level lineage from SQL-based ETL jobs.
pp-netlify
mvanhorn/printing-press-library
Every Netlify API endpoint, plus a local SQLite mirror of your whole account so you can search, diff, and audit across all sites at once — offline and agent-native. Trigger phrases: `list my netlify sites`, `check netlify deploys`, `search netlify form submissions`, `audit netlify dns`, `compare netlify env vars`, `use netlify`, `run netlify`.
data-pipelines
booklib-ai/booklib
>
airflow-dag-patterns
fefogarcia/approved-skills
Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
occupancy-map
isaac-sim/isaacsim
Generate ROS-compatible occupancy maps from USD scenes for robot navigation and perception training. Covers obstacle extraction, 2D grid projection, height filtering, dilation buffers, and Nav2/MobilityGen/A* path planning integration.
pipeline
elliotjlt/claude-skill-potions
|