499 free skills
Data Engineering skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
499 free skills
Skills for data engineering — ETL pipelines, Airflow DAGs, Spark jobs, and data warehouse modeling.
Sourced from real, public repositories — synced daily, never invented.
15 tools across six categories
13 of them never send your data anywhere
Free · No signup · No trial clock
SEE THE DIRECTORY

data-engineer
daffy0208/ai-dev-standards
Expert in data pipelines, ETL processes, and data infrastructure
spark-recipe-stakeholder-brief
readdle/spark-cli-skills
>-
spark-persona-exec-assistant
readdle/spark-cli-skills
>-
data-engineering
rohitg00/awesome-claude-code-toolkit
Data engineering patterns for ETL pipelines, data warehousing, Apache Spark, and data quality validation
spark-recipe-topic-timeline
readdle/spark-cli-skills
>-
alibabacloud-emr-spark-manage
aliyun/alibabacloud-aiops-skills
>
spark-recipe-newsletter-cleanup
readdle/spark-cli-skills
>-
tenero
aibtcdev/skills
Tenero (formerly STXTools) market analytics — token info, market stats, top gainers/losers, wallet holdings and trades, trending DEX pools, whale trades, holder distribution, and search. Covers Stacks, Spark, and SportsFun chains. No API key required.
airflow-expert
personamanagmentlayer/pcl
Expert-level Apache Airflow orchestration, DAGs, operators, sensors, XComs, task dependencies, and scheduling
spark-recipe-inbox-zero
readdle/spark-cli-skills
>-
spark-recipe-shared-inbox-status
readdle/spark-cli-skills
>-
spark-recipe-invitation-manager
readdle/spark-cli-skills
>-
spark-recipe-schedule-meeting
readdle/spark-cli-skills
>-
python-pipeline
jamditis/claude-skills-journalism
Python data pipelines with modular architecture. Use for content workflows, batch jobs, or Google Sheets/Drive integration.
aibtc-agents
aibtcdev/skills
Community registry of agent configurations for the AIBTC platform — browse reference configs for arc0btc, spark0btc, iris0btc, loom0btc, and forge0btc, or copy the template to bootstrap a new agent.
spark-recipe-label-organize
readdle/spark-cli-skills
>-
spark-recipe-team-workload
readdle/spark-cli-skills
>-
data-warehouse-experimentation
rampstackco/claude-skills
Running experiments out of the data warehouse instead of via dedicated experiment platforms. SQL-based assignment, exposure logging discipline, metric definitions in dbt models, statistical analysis in SQL or Python, variance reduction with CUPED, sequential testing, and the operational tradeoffs vs platforms like Statsig and Optimizely. Triggers on warehouse-native experimentation, run experiments in BigQuery, run experiments in Snowflake, dbt experiments, SQL t-test, CUPED variance reduction, exposure log, sample ratio mismatch, sequential testing, mSPRT, doubly robust estimation, build vs buy experimentation. Also triggers when the team is choosing between platform and warehouse, building warehouse-native experiment infrastructure, auditing one, or running an experiment with a custom metric the platform cannot handle.
spark-recipe-priority-tuning
readdle/spark-cli-skills
>-
spark-recipe-notification-hygiene
readdle/spark-cli-skills
>-
spark-persona-founder
readdle/spark-cli-skills
>-
spark-recipe-draft-batch
readdle/spark-cli-skills
>-
spark-recipe-new-sender-review
readdle/spark-cli-skills
>-
spark-persona-freelancer
readdle/spark-cli-skills
>-
spark-recipe-vacation-catchup
readdle/spark-cli-skills
>-
spark-persona-project-manager
readdle/spark-cli-skills
>-
spark-recipe-delegate-and-track
readdle/spark-cli-skills
>-
databricks-execution-compute
databricks/databricks-agent-skills
Execute code and manage compute on Databricks: run Python/Scala/SQL/R via serverless, classic, or interactive clusters, and create/resize/delete clusters and SQL warehouses.
spark-recipe-meeting-followup
readdle/spark-cli-skills
>-
spark-persona-meeting-manager
readdle/spark-cli-skills
>-
segment
membranedev/application-skills
|
spark-recipe-shared-inbox-triage
readdle/spark-cli-skills
>-
databricks-spark-structured-streaming
databricks/databricks-agent-skills
Comprehensive guide to Spark Structured Streaming for production workloads. Use when building streaming pipelines, working with Kafka ingestion, implementing Real-Time Mode (RTM), configuring triggers (processingTime, availableNow), handling stateful operations with watermarks, optimizing checkpoints, performing stream-stream or stream-static joins, writing to multiple sinks, or tuning streaming cost and performance.
spark-persona-support-agent
readdle/spark-cli-skills
>-
spark-persona-team-lead
readdle/spark-cli-skills
>-
databricks-ml-training
databricks/databricks-agent-skills
Train ML models on Databricks. Use for: classification/regression/deep-learning (XGBoost, scikit-learn, LightGBM, PyTorch) with Optuna, @prod/@challenger aliases, batch scoring (spark_udf for plain models, fe.score_batch for feature-store-backed), custom PyFunc, custom ResponsesAgent (LangGraph + UC Function/Vector Search); UC feature tables + FeatureLookup + point-in-time joins + Lakebase online store; declarative Feature Views (create_feature, DeltaTableSource, RollingWindow/SlidingWindow/TumblingWindow, materialize_features, streaming Kafka features). NOT for: endpoint ops (databricks-model-serving), MLflow evaluation (databricks-mlflow-evaluation).
netlify
membranedev/application-skills
|
spark-persona-sales-rep
readdle/spark-cli-skills
>-
spark-recipe-unsubscribe-audit
readdle/spark-cli-skills
>-
ml-data-pipeline-architecture
terrylica/cc-skills
Patterns for efficient ML data pipelines using Polars, Arrow, and ClickHouse. TRIGGERS - data pipeline, polars vs pandas, arrow format
databricks-iceberg
databricks/databricks-agent-skills
Apache Iceberg tables on Databricks — Managed Iceberg tables, External Iceberg Reads (fka Uniform), Compatibility Mode, Iceberg REST Catalog (IRC), Iceberg v3, Snowflake interop, PyIceberg, OSS Spark, external engine access and credential vending. Use when creating Iceberg tables, enabling External Iceberg Reads (uniform) on Delta tables (including Streaming Tables and Materialized Views via compatibility mode), configuring external engines to read Databricks tables via Unity Catalog IRC, integrating with Snowflake catalog to read Foreign Iceberg tables
labs64-netlicensing
membranedev/application-skills
|
pipeliner
membranedev/application-skills
|
dbt-model-index
warpdotdev/oz-skills
Provide a lookup index of dbt models (BigQuery tables) to guide query writing against a data warehouse. Use when you need to query, analyze, or look up data in a dbt-powered data warehouse, or when resolving a vague data question into the right BigQuery tables to query.
azure-synapse-analytics
microsoftdocs/agent-skills
Expert knowledge for Azure Synapse Analytics development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations & coding patterns, and deployment. Use when using Synapse SQL pools, serverless SQL, Spark pools, Synapse Link, or PolyBase/ELT data loading, and other Azure Synapse Analytics related development tasks. Not for Azure Data Factory (use azure-data-factory), Azure Data Explorer (use azure-data-explorer), Azure HDInsight (use azure-hdinsight), Azure Databricks (use azure-databricks).
vfx
drawcall-ai/skills
Add impact and feedback effects to a Three.js game with lightweight particle bursts — muzzle flashes, hit sparks, blood/dust, explosions, pickups — using pooled additive sprites or Points. Use when actions need visible punch (firing, impacts, deaths, explosions) or for ambient effects (smoke, embers, dust).
apache-superset
membranedev/application-skills
|
syncfusion-wpf-sparkline
syncfusion/wpf-ui-components-skills
Comprehensive guide for implementing Syncfusion WPF Sparkline (SfSparkline) controls in Windows Presentation Foundation applications. Use this when working with sparklines, mini charts, or trend visualization. This skill covers sparkline types (line, column, area, WinLoss), markers, track ball, range bands, axis controls, and segment customization for compact data visualization in WPF applications.
azure-hdinsight
microsoftdocs/agent-skills
Expert knowledge for Azure HDInsight development including troubleshooting, best practices, decision making, architecture & design patterns, limits & quotas, security, configuration, integrations & coding patterns, and deployment. Use when working with HDInsight Spark/Hive/Kafka/HBase clusters, Ambari/Oozie pipelines, or Azure-integrated storage/BI, and other Azure HDInsight related development tasks. Not for Azure Databricks (use azure-databricks), Azure Synapse Analytics (use azure-synapse-analytics), Azure Stream Analytics (use azure-stream-analytics).
smoke-test-ml-pipeline
probabl-ai/skills
>
migrating-dbt-project-across-platforms
dbt-labs/dbt-agent-skills
Use when migrating a dbt project from one data platform or data warehouse to another (e.g., Snowflake to Databricks, Databricks to Snowflake) using dbt Fusion's real-time compilation to identify and fix SQL dialect differences.
data-pipeline
rightnow-ai/openfang
Data pipeline expert for ETL, Apache Spark, Airflow, dbt, and data quality
social-selling
gtmagents/gtm-agents
Use when engaging prospects through LinkedIn, communities, and social channels to spark warm conversations and meetings.
magic-etl
stahura/domo-ai-vibe-rules
Create, update, and execute Magic ETL dataflows programmatically via API and CLI. Covers DAG-based JSON dataflow definitions, input/transform/output node wiring, join operations, and execution lifecycle.
pipeline
jwilger/agent-skills
>-
discover-data
rand/cc-polymath
Automatically discover data pipeline and ETL skills when working with ETL, data pipelines, streaming, batch processing, data validation, or pipeline orchestration. Activates for data development tasks.
magic-etl-cli
stahura/domo-ai-vibe-rules
Magic ETL dataflows via community-domo-cli — list, get-definition, create, update, run, execution status; JSON DAG actions, transforms, joins. Use when automating dataflows with the community Domo CLI end-to-end. For REST/Java-CLI–first flows or mixed API patterns, use magic-etl instead.
json-and-csv-data-transformation
besoeasy/open-skills
Transform data between JSON, CSV, and other formats with filtering, mapping, and flattening. Use when: (1) Converting API responses to CSV, (2) Processing data pipelines, (3) Extracting specific fields, or (4) Flattening nested structures.
configuring-airflow-language-sdks
astronomer/agents
Configures Airflow to run language SDK tasks (Java, Go, and future native SDKs) — register a coordinator, map a queue to it, ensure the runtime/artifact on workers, and tune coordinator options. Use when the user wants Airflow to route a queue to a native-language coordinator, asks about the `[sdk]` `coordinators`/`queue_to_coordinator` settings, `AIRFLOW__SDK__COORDINATORS`, `jars_root`, `executables_root` or other coordinator `kwargs`, `task_startup_timeout`, or why their native tasks aren't being picked up. Covers the shared routing mechanism plus per-coordinator options (e.g. JavaCoordinator, ExecutableCoordinator).
authoring-go-sdk-tasks
astronomer/agents
Writes Airflow task logic in Go using the Airflow Go SDK. Use when the user wants to implement Airflow tasks in Go, asks about `BundleProvider`/`RegisterDags`, the `bundlev1` Registry/Dag interfaces, registering Go tasks (`AddTask`/`AddTaskWithName`), dependency injection by parameter type (`context.Context`, `sdk.TIRunContext`, `*slog.Logger`, `sdk.Client`), or reading connections/variables/XComs from Go. This skill covers the Go-specific native API; the shared Python-stub pattern and conceptual model live in authoring-language-sdk-tasks. For building/packing/shipping the bundle see deploying-go-sdk-bundles; for coordinator config see configuring-airflow-language-sdks.