499 free skills
Data Engineering skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
499 free skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
15 tools across six categories
13 of them never send your data anywhere
Free · No signup · No trial clock
SEE THE DIRECTORY

netlify-deploy
firecrawl/skills
Deploy web projects to Netlify using the Netlify CLI (`npx netlify`). Use when the user asks to deploy, host, publish, or link a site/repo on Netlify, including preview and production deploys.
data-pipeline-engineering
hack23/riksdagsmonitor
ETL workflow design, automated data fetching, version tracking, and pipeline orchestration expertise
data-pipelines
excalimate/excalimate
>
data-pipeline-patterns
vibeeval/vibecosystem
ETL/ELT patterns, batch vs streaming, idempotency, data quality framework, and pipeline orchestration
etl-pipeline
inbharatai/claude-skills
Build end-to-end ETL pipelines — extract from APIs/databases, transform, validate, and load into data warehouses.
deploy-netlify
ominou5/funnel-architect-plugin
>
spark-submit
oleanderhq/skills
>-
spark-cli-knowledge-sharing
memcoai/spark-cli-skills
|
spark-lineage
oleanderhq/skills
>-
spark-builder
inbharatai/claude-skills
Write PySpark and Spark SQL jobs — RDDs, DataFrames, streaming, MLlib, and cluster configuration.
gcp-spark
gemini-cli-extensions/data-agent-kit-starter-pack
|
gcp-data-pipelines
gemini-cli-extensions/data-agent-kit-starter-pack
Primary entry point for building, managing, and orchestrating data pipelines
spark-operations-cli
microsoft/skills-for-fabric
>
using-data-engineering-agent-skills
vaquarkhan/data-engineering-agent-skills
Helps agents classify data engineering work, choose the right preset and skill bundle, and pick the safest next command. Use when starting a session, triaging an ambiguous request, or deciding how to proceed.
spark-benchmark
sapid-labs/spark-skills
Use when the user wants to run, measure, or submit a How To Spark community benchmark ("bench card") for a recipe on a DGX Spark — e.g. "benchmark this recipe", "measure decode speed on the Spark", "submit throughput numbers for <recipe>", or "report a bench card". Requires the howtospark MCP server and an API key from howtospark.com/profile.
spark-eval
sapid-labs/spark-skills
Use when the user wants to run, measure, or submit a How To Spark community eval score (accuracy/quality benchmark, e.g. GSM8K, HumanEval) for a recipe served on a DGX Spark — e.g. "run the missing evals for <recipe>", "score this recipe on gsm8k", "submit an eval result". Requires the howtospark MCP server and an API key from howtospark.com/profile.
use-spark
drewburchfield/spark-cli-skills
>-
stream-chain
dnyoussef/ai-chrome-extension
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
data-reverse-etl
j4flmao/agent-skills
>
data-etl-pipeline
j4flmao/agent-skills
>
fp-pipe-ref
whatiskadudoing/fp-ts-skills
Quick reference for pipe and flow. Use when user needs to chain functions, compose operations, or build data pipelines in fp-ts.
enterprise-etl-and-data-integration-modernization
vaquarkhan/data-engineering-agent-skills
Guides agents through operating, hardening, and modernizing enterprise ETL and integration stacks such as Informatica, Talend, DataStage, SSIS, and Matillion. Use when legacy mappings, job orchestration, migration, or coexistence with modern lakehouse patterns must be handled safely.
airflow
kilo-org/kilo-marketplace
>-
airflow-and-workflow-orchestration
vaquarkhan/data-engineering-agent-skills
Guides agents through workflow orchestration design and operation across Airflow-style DAGs, cloud-native schedulers, and event-driven pipeline control planes. Use when building or modifying workflow dependencies, retries, triggers, sensors, SLAs, or cross-system pipeline coordination.
deploying-airflow
kilo-org/kilo-marketplace
>-
adf-master
kilo-org/kilo-marketplace
>-
azure-synapse-analytics
kilo-org/kilo-marketplace
>-
modeling-warehouse-foundations
posthog/ai-plugin
>
stream-chain
ruvnet/agentic-flow
Stream-JSON chaining for multi-agent pipelines, data transformation, and sequential workflows
airflow-translations
astronomer/airflow
>
databricks-spark-structured-streaming
kilo-org/kilo-marketplace
>-
etl-elt-and-modernization-strategy
vaquarkhan/data-engineering-agent-skills
Guides agents through ETL, ELT, and transformation-modernization decisions. Use when choosing execution boundaries, redesigning transformation layers, or moving from legacy ETL estates to warehouse- or lakehouse-centered ELT patterns.
reverse-etl-and-operational-data-serving
vaquarkhan/data-engineering-agent-skills
Guides agents through reverse ETL and operational data serving workflows. Use when sending curated warehouse data to business systems, SaaS tools, APIs, activation layers, or operational applications that rely on stable downstream contracts.
data-pipeline
claude-office-skills/skills-hub
Data pipeline and ETL automation - extract, transform, load workflows for data integration and analytics
hbase
kilo-org/kilo-marketplace
Apache HBase wide-column store on Hadoop. Use for big data.
spark-engineer
kilo-org/kilo-marketplace
>-
moai-lang-scala
modu-ai/cc-plugins
>
warehouse-performance-and-cost-optimization
vaquarkhan/data-engineering-agent-skills
Guides agents through warehouse performance and cost decisions. Use when optimizing BigQuery, Snowflake, Redshift, Athena, Synapse, or lakehouse query patterns, storage layout, and workload isolation.
etl-retry-backoff-simulator
sisodiabhumca/agent-skills
Simulate retry and exponential backoff strategies against a failure-rate model to estimate expected runtime and cost (vendor-neutral).
warehouse-and-schema-design
vaquarkhan/data-engineering-agent-skills
Guides agents through data warehouse and schema design. Use when defining fact and dimension models, keys, grain, normalization versus denormalization, serving-layer schema boundaries, and downstream-friendly table design.
polygres-data-pipeline
evokoa/polygres-skills
Set up or extend Polygres from either a short request such as "Help me set up Polygres" or a detailed ingestion, memory, graph, embedding, synchronization, or retrieval specification. Also use for questions such as "What can I do with Polygres?" by scanning the accessible current workspace and project read-only and giving a personalized recommendation without changing anything. Ask one short direction question first when a setup request identifies neither a source nor an outcome; otherwise inspect the user's accessible data and application, resolve only critical unknowns, design the smallest useful schema and retrieval setup, generate source-specific ingestion and retrieval code, configure the selected project after one consolidated approval, verify a small vertical slice, and optionally connect capture and recall to the user's agent. Use whenever the user intends to make their data usable through Polygres, even if they do not say "data pipeline.
sql-server-table-reconciliation
github/awesome-copilot
Use when: comparing SQL Server tables across instances, data migration validation, ETL verification, row mismatch detection, schema drift, reconciliation report, production vs staging comparison. Uses mssql-python driver with Apache Arrow for fast columnar data transfer and comparison.
airflow-new-sdk
astronomer/airflow
>
scala-data-engineering-on-jvm-runtimes
vaquarkhan/data-engineering-agent-skills
Guides agents through Scala-based data engineering on JVM runtimes. Use when building Spark, Flink, Kafka Streams, or other Scala data jobs that require explicit build, packaging, type, and runtime discipline.
stock-research-publish
samyakjain0606/awesome-stock-skills
Publish a generated equity research report as either a professional PDF or a magazine-quality HTML page, then optionally deploy the HTML to Netlify and return a live URL. Use after the stock-research-pipeline has produced NLM query answers (Business Model, Industry, Management, Financials, Growth Triggers / VP, Risks / Bull-Base-Bear) for a SYMBOL, OR when the user says "publish report for [SYMBOL]", "deploy report to netlify", "html version of [SYMBOL] report", "ship the [SYMBOL] research", "make the [SYMBOL] report shareable", or "/stock-research-publish [SYMBOL]". Asks the user upfront whether they want PDF, HTML, or both.
accelerated-computing-cudf
practicalswan/agent-skills
Official NVIDIA-authored guidance for NVIDIA cuDF GPU DataFrames, pandas acceleration, dask-cuDF, ETL, joins, groupby, CSV/Parquet I/O, nullable semantics, and multi-GPU DataFrame workloads.
spark-and-distributed-processing
vaquarkhan/data-engineering-agent-skills
Guides agents through batch and distributed data processing design. Use when implementing or reviewing Spark-based pipelines, or managed distributed runtimes such as Glue and EMR.
spark
zrr1999/skills
>
airflow-java-sdk
astronomer/airflow
>
data-pipeline-check
gonglitian/agent-skills
Validate data pipeline integrity, check dataset format compatibility, verify data quality, and diagnose data loading issues. Trigger when user mentions data problems, format conversion, dataset validation, or says "检查数据", "数据格式", "data check", "数据质量", "数据对不对", "格式兼容吗", "数据结构", "数据长什么样", "action space对得上吗", "验证数据".
ray-data
firecrawl/ai-research-skills
Scalable data processing for ML workloads. Streaming execution across CPU/GPU, supports Parquet/CSV/JSON/images. Integrates with Ray Train, PyTorch, TensorFlow. Scales from single machine to 100s of nodes. Use for batch inference, data preprocessing, multi-modal data loading, or distributed ETL pipelines.
spark-structured-streaming
databricks-solutions/ai-dev-kit
Comprehensive guide to Spark Structured Streaming for production workloads. Use when building streaming pipelines, implementing real-time data processing, handling stateful operations, or optimizing streaming performance.
kafka-engineer
belokonm/claude-supercode-skills
Expert in Apache Kafka, Event Streaming, and Real-time Data Pipelines. Specializes in Kafka Connect, KSQL, and Schema Registry.
data-pipeline-quality
hollandkevint/data-product-operator
>
ETL Pipeline
claude-office-skills/skills-hub
Design and automate Extract, Transform, Load data pipelines for data integration and analytics
data-pipelines
zlstas/skills
>
multi-omics-pipeline
awslabs/hcls-agent-skills
>
blog-spark
bsbbera/ideaverse-skills
Write an SEO-optimized teaser blog for an ebook-weaver topic — opens curiosity loops, teases 2-3 cited facts, drives readers to the ebook, and never resolves the questions it raises. Trigger the user types /blog-spark "<topic>" or /blog-spark <topic-folder-path>.
unified-data-warehouse
kangarooking/ai-for-everyone-skill
|
data-pipeline
zhouziyue233/great-econometrics
End-to-end data pipeline for empirical research: fetch economic data from APIs (FRED, World Bank, IMF, BLS, OECD, Yahoo Finance), clean and transform raw data, construct strategy-specific variables, and validate panel structure. Use when asked to fetch data, download data, clean data, merge datasets, prepare analysis-ready data.