499 free skills
Data Engineering skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
499 free skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
15 tools across six categories
13 of them never send your data anywhere
Free · No signup · No trial clock
SEE THE DIRECTORY

data-pipelines
hacho55/absolutelyskilled
>
data-quality-frameworks
karim-bhalwani/agent-skills-collection
Specialist in data quality validation frameworks—Great Expectations, dbt tests, data contracts. Builds comprehensive data quality gates into pipelines for reliability and trust.
data-analysis
flonat/flonat-research
Deliver an end-to-end analysis pipeline: EDA, estimation, or publication output. Use when the user requests an end-to-end analysis pipeline: EDA, estimation, or publication output.
data-eng-pipeline-architect
scanady/nexus-skills
Expert data engineer skill for designing and building production-grade data pipelines, ETL/ELT systems, data models, and DataOps workflows. Use when designing data architecture, building batch or streaming pipelines, implementing data quality checks, modeling warehouse schemas, or applying DataOps practices with dbt, Airflow, Spark, Kafka, Snowflake, BigQuery, or Databricks.
data-pipeline-engineer
karim-bhalwani/agent-skills-collection
Specialist in hands-on ETL/ELT design, implementation, and optimization using Airflow, Spark, Kafka, dbt, and data quality frameworks. Use when building data pipelines, orchestrating data workflows, optimizing data processing, implementing streaming solutions, or ensuring data quality throughout pipelines.
flame-test-data
rk234/flame-cli
>
stream-chain
agenticsorg/hackathon-tv5
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
stream-chain
bcasci/boswell-backoffice
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
stream-chain
iluminatto1970/ruflo
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
stream-chain
jkappers/claude-flow
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
stream-chain
sparkling/ruflo
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
test-pipeline
ahrav/gossip-rs
Use when implementing a new feature and assessing coverage gaps, during periodic test hygiene, when test suites feel bloated, or before merging code that changes coordination or hot paths. Two-phase assess-then-improve testing pipeline.
dev-etl-designer
khalilbenaz/claude-skills-collection
Conception de processus ETL/ELT pour l'intégration de données. Se déclenche avec "ETL", "ELT", "extraction", "transformation", "load", "data integration", "data warehouse", "SSIS", "Talend", "Informatica". Also triggers on "ETL process", "ELT design", "data integration job".
airflow-dag-patterns
coppermare/skillverse
Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling batch jobs.
spark-optimization
karim-bhalwani/agent-skills-collection
Specialist in Apache Spark performance optimization—partitioning strategies, memory tuning, shuffle reduction, and job profiling for production systems. Use when optimizing Spark jobs, tuning performance, reducing costs, profiling applications, or scaling Spark workloads.
warehouse-optimization
nrakow/ae-skills-dev
Optimize warehouse performance and cost through clustering, partitioning, materialization strategies, and query tuning. Use when queries are slow, compute costs are high, or a model needs to be optimized for production scale. Triggers: 'optimize warehouse', 'query performance', 'slow queries', 'clustering', 'partitioning', 'cost optimization', 'warehouse cost', 'query tuning', 'performance tuning'.
data-pipeline
librefang/librefang-registry
Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality
data-engineering-data-pipeline
rootcastleco/rei-skills
You are a data pipeline architecture expert specializing in scalable, reliable, and cost-effective data pipelines for batch and streaming data processing.
airflow
joavirtudes19/agents
Queries, manages, and troubleshoots Apache Airflow using the af CLI. Covers listing DAGs, triggering runs, reading task logs, diagnosing failures, debugging DAG import errors, checking connections, variables, pools, and monitoring health. Also routes to sub-skills for writing DAGs, debugging, deploying, and migrating Airflow 2 to 3. Use when user mentions "Airflow", "DAG", "DAG run", "task log", "import error", "parse error", "broken DAG", or asks to "trigger a pipeline", "debug import errors", "check Airflow health", "list connections", "retry a run", or any Airflow operation. Do NOT use for warehouse/SQL analytics on Airflow metadata tables — use analyzing-data instead.
databricks-sql
mats16/briclaude
|
data-pipeline
mk-knight23/ai-agent-nanobot
Builds ETL pipelines for CSV, JSON, Parquet, or database sources. Describe your transformation in natural language; Nanobot generates a Python pipeline using pandas/polars, runs it, validates the output schema, and writes PIPELINE_REPORT.md. Use when you need to transform, clean, or migrate data between formats or systems. Works with local files or database connections (Postgres, SQLite, Supabase).
netlify-deploy
firecrawl/openai-skills
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
databricks-iceberg
databricks-solutions/lakebase-online-ml
Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS Spark, external engine access and credential vending. Use when creating Iceberg tables, enabling External Iceberg Reads (uniform) on Delta tables (including Streaming Tables and Materialized Views via compatibility mode), configuring external engines to read Databricks tables via Unity Catalog IRC, integrating with Snowflake catalog to read Foreign Iceberg tables
databricks-spark-structured-streaming
databricks-solutions/lakebase-online-ml
Comprehensive guide to Spark Structured Streaming for production workloads. Use when building streaming pipelines, implementing real-time data processing, handling stateful operations, or optimizing streaming performance.
spark-python-data-source
databricks-solutions/lakebase-online-ml
Use when building custom Spark data source connectors for external systems (databases, APIs, message queues), implementing batch/streaming readers/writers, or creating data source plugins for systems without native Spark support. Triggers - "build Spark data source", "create Spark connector", "implement Spark reader/writer", "connect Spark to [system]", "streaming data source
04-conformed-dimensions
databricks-solutions/vibe-coding-workshop-template
Enterprise integration patterns for Gold layer dimensional models. Covers conformed dimensions, the enterprise data warehouse bus matrix, shrunken/rollup dimensions, conformed facts, and drill-across query patterns. Use when planning dimensions shared across multiple fact tables, creating a bus matrix for enterprise integration, designing rollup dimensions, or enabling cross-process analytics. Triggers on "conformed dimension", "bus matrix", "drill-across", "shrunken dimension", "rollup", "enterprise integration", "cross-process".
scala
dkbnull/hello-skill
Scala开发专家助手。当用户需要进行Scala函数式编程、大数据开发、Akka并发、Play框架或Spark开发时调用。
synapseml-local-setup
microsoft/synapseml
Set up and validate SynapseML locally in WSL or Linux. Use when an agent needs SynapseML working locally, runs sbt compile/test, sees Java 21, Scala 2.12 compiler-bridge, bad constant pool index, Spark, or local validation failures.
airflow-workflow
datus-ai/datus-agent
Execution guide for Airflow scheduled jobs — troubleshooting, updating, conn_id conventions, and cron references
airflow
dkbnull/hello-skill
Airflow开发专家助手。当用户需要进行Airflow工作流调度、DAG开发、数据管道编排、任务依赖管理或ETL流程自动化时调用。
netlify-deploy
abanoub-ashraf/manus-skills-import
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
netlify
ericrisco/rsc-harness
Use when deploying or operating a site on Netlify — writing or fixing netlify.toml, authoring Functions or Edge Functions, redirects, rewrites and headers, env vars per deploy context, and shipping via the Netlify CLI. NOT deploying to Vercel (that is `vercel`).
netlify-deploy
lidge-jun/cli-jaw-skills
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
netlify-deploy
jackeyunjie/skillscodex
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
deploy
victorvianaverbo/avantik-site
Use when the user wants to publish, deploy, upload the site, put it online, or host it. Covers Git, GitHub via gh CLI, Netlify init with GitHub OAuth integration, Deploy Previews via PR to save build minutes, and pre-deploy checks.
forms
victorvianaverbo/avantik-site
Use when creating or modifying contact forms, lead capture forms, or any form with a phone field. Includes intl-tel-input with masks, email validation, Netlify Forms integration with AJAX submit, redirect with URL params forwarding, and thank you page.
netlify-deploy
rylaispirit/rylai-codex-hermes-skills
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
nodeboot-server-netlify
nodejs-boot/node-boot
Use when the user wants to deploy a Node-Boot application to Netlify Functions, and needs the catch-all `netlify/functions/api.ts` pattern built around `NetlifyServer.getHandler()` plus the `netlify.toml` rewrite that sends `/api/*` into the function.
netlify-publish
mz038197/vanscoding-skills
>-
netlify-api
cuzic/clipship
Netlify API を使用したコードを書く際に使用。ZIPデプロイ、サイト作成、ky/fflate/zod を使った実装パターンを提供。
kafka-engineer
xgaisystems/claude-supercode-skills-404kidwiz
Expert in Apache Kafka, Event Streaming, and Real-time Data Pipelines. Specializes in Kafka Connect, KSQL, and Schema Registry.
etl-pipeline-design
shinzoxd/knackbox
Design ETL and ELT pipelines with clear contracts, quality checks,
data-pipeline
tryboy869/dojutsu-for-ai
Apply when building ETL pipelines, data processing workflows, batch jobs, or streaming data systems. Covers: chunking, idempotency, error handling, progress tracking, schema validation. Trigger for: ETL, pipeline, batch, data processing, CSV, transform, extract, load.
perf-sampling-parser
qualcomm/easywos
Analyze per-process CPU usage from ETL log files and export SpeedScope flame graphs. Keywords - ETL, flame graph, CPU usage, speedscope.
etl-generator
qualcomm/easywos
Run a target program under a CPU-sampling profiler and produce an ETL trace file for later performance analysis. Keywords - collect ETL, profile a program, generate etl, capture CPU trace, record ETL.
architecture-boundaries
unit8co/unit8-public-project-starter
Preserve clean starter architecture boundaries between shared config, runtime wiring, API, CLI, agentic modules, and ETL modules. Use when moving code or introducing new packages in src.
data-analytics
t786279007/yhct_huatu
Create data analytics and data pipeline diagrams using PlantUML syntax with analytics/database stencil icons. Best for ETL pipelines, data lakes, real-time streaming, data warehousing, and BI dashboards. NOT for simple flowcharts (use mermaid) or general cloud infra (use cloud skill).
data-analytics
95369149/openclaw
Create data analytics and data pipeline diagrams using PlantUML syntax with analytics/database stencil icons. Best for ETL pipelines, data lakes, real-time streaming, data warehousing, and BI dashboards. NOT for simple flowcharts (use mermaid) or general cloud infra (use cloud skill).
etl-data-sources
nielskschjoedt/groen-trepart-tracker
>
flyte-sdk-data
flyteorg/flyte-agent-plugins
Handles data engineering patterns: ETL pipelines, data processing, data quality checks, fanout/map tasks, conditions, dynamic workflows, and batch data transformations. Use when the user wants to build ETL pipelines, process large datasets, run data quality checks, fan out data processing tasks, or handle batch data transformations. Trigger words: "ETL", "data pipeline", "data processing", "fanout", "map", "transform", "data quality", "parquet", "CSV", "batch", "extract", "load", "validate", "schema".
data-analytics
happyrust/skills
Create data analytics and data pipeline diagrams using PlantUML syntax with analytics/database stencil icons. Best for ETL pipelines, data lakes, real-time streaming, data warehousing, and BI dashboards. NOT for simple flowcharts (use mermaid) or general cloud infra (use cloud skill).
python-data-engineering
dtsong/my-claude-setup
Use this skill when writing Python code for data pipelines or transformations. Covers Polars, Pandas, PySpark DataFrames, dbt Python models, API extraction scripts, and data validation with Pydantic or Pandera. Common phrases: \"Polars vs Pandas\", \"PySpark DataFrame\", \"validate this data\", \"Python extraction script\". Do NOT use for SQL-based dbt models (use dbt-transforms) or integration architecture (use data-integration).
data-processing
pedrohbo/opencode-config-skills
Practical data processing using Pandas, Polars, and NumPy. Use for data cleaning, transformation, and analysis tasks in Python data pipelines.
ds-shopper-spark
ingeleyton/spark-skill
Playbook técnico de Apache Spark para diagnóstico, tuning y troubleshooting sobre RDD/Core Spark, Spark SQL/DataFrames, PySpark, MLlib/Spark ML, Streaming y operación en clúster. Usar cuando Codex deba analizar jobs lentos o inestables, reducir shuffles/OOM/GC, optimizar joins y particionamiento, evaluar UDFs y conversiones entre APIs, ajustar persistencia/checkpointing/configuración, diseñar pruebas distribuidas, o entregar planes de mejora con métricas de validación.
spark-memory
memcoai/spark-teams-cli-skills
|
spark-analyzer
l4nternnn/spark-analyzer
|
tt-spark-ads
0-shiv/secondstep-claude-skills
Analyzes and optimizes TikTok Spark Ads -- creator content amplification, authorization codes, organic vs paid lift, and performance comparison against standard in-feed ads.
spark-savings-plugin
mig-pre/plugin-store
Spark Savings - earn Sky Savings Rate (SSR) on USDS via the sUSDS yield-bearing vault. Deposit USDS or upgrade DAI 1:1, redeem any time, no collateral, no liquidation. Supports Ethereum (ERC-4626 vault), Base & Arbitrum (Spark PSM).
gpt-engineer-spark
symbaiex/skills
Lead end-to-end engineering work with a capable main agent and a parallel fleet of model-pinned GPT-5.3-Codex-Spark subagents. Use when the user asks for a Spark fleet, ultra-fast parallel coding agents, rapid repository exploration, many bounded implementation shards, or low-latency independent verification while retaining architecture, integration, and final acceptance with the main agent.
spark-optimization
sarovios/dotfiles
Analyze and optimize Apache Spark jobs for performance