spark-optimization

29.9k
wshobsonwshobson

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

192 days ago

senior-data-engineer

21.8k
davila7davila7

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

192 days ago

generate-sparkle-appcast

19.1k
CaldisCaldis

Generate Mos Sparkle appcast.xml from the latest build zip and recent git changes (since a given commit), then sync to docs/ for publishing.

192 days ago

spark-optimization

18.0k
sickn33sickn33

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

192 days ago

data-engineer

18.0k
sickn33sickn33

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms. Use PROACTIVELY for data pipeline design, analytics infrastructure, or modern data stack implementation.

192 days ago

scala-pro

18.0k
sickn33sickn33

Master enterprise-grade Scala development with functional programming, distributed systems, and big data processing. Expert in Apache Pekko, Akka, Spark, ZIO/Cats Effect, and reactive architectures.

192 days ago

senior-data-engineer

2.2k
alirezarezvanialirezarezvani

Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, implementing data governance, or troubleshooting data issues.

192 days ago

BreezClaw

1.8k
openclawopenclaw

Self-custodial Bitcoin and Lightning wallet for AI agents. Send and receive sats via Lightning Network, Spark, or on-chain Bitcoin. Use when: checking bitcoin balance, sending/receiving payments, generating Lightning invoices, managing wallet operations. Requires the BreezClaw plugin and a Breez API key.

192 days ago

george

1.8k
openclawopenclaw

Automate George online banking (Erste Bank / Sparkasse Austria): login/logout, list accounts, and fetch transactions via Playwright.

192 days ago

spark-engineer

1.8k
openclawopenclaw

Use when building Apache Spark applications, distributed data processing pipelines, or optimizing big data workloads. Invoke for DataFrame API, Spark SQL, RDD operations, performance tuning, streaming analytics.

192 days ago

spark-sql-optimizer

1.5k
jeremylongshorejeremylongshore

Optimize Spark SQL optimizer operations. Auto-activating skill for Data Pipelines. Triggers on: Spark SQL optimizer, Spark SQL optimizer Part of the Data Pipelines skill category. Use when working with Spark SQL optimizer functionality. Trigger with phrases like "Spark SQL optimizer", "Spark optimizer", "Spark".

192 days ago

databricks-performance-tuning

1.5k
jeremylongshorejeremylongshore

Optimize Databricks cluster and query performance. Use when jobs are running slowly, optimizing Spark configurations, or improving Delta Lake query performance. Trigger with phrases like "databricks performance", "spark tuning", "databricks slow", "optimize databricks", "cluster performance".

192 days ago

Spark Job Creator

1.5k
jeremylongshorejeremylongshore

Spark Job Creator - Auto-activating skill for Data Pipelines. Triggers on: spark job creator, spark job creator Part of the Data Pipelines skill category.

192 days ago

databricks-common-errors

1.5k
Jeremylongshore Claude Code Plugins Plus Skills Databricks Common ErrorsJeremylongshore Claude Code Plugins Plus Skills Databricks Common Errors

Diagnose and fix Databricks common errors and exceptions. Use when encountering Databricks errors, debugging failed jobs, or troubleshooting cluster and notebook issues. Trigger with phrases like "databricks error", "fix databricks", "databricks not working", "debug databricks", "spark error".

192 days ago

moai-lang-scala

776
modu-aimodu-ai

Scala 3.4+ development specialist covering Akka, Cats Effect, ZIO, and Spark patterns. Use when building distributed systems, big data pipelines, or functional programming applications.

192 days ago

Apache Spark Optimizer

376
a5c-aia5c-ai

Analyzes and optimizes Apache Spark jobs for performance, cost, and resource utilization

192 days ago

data-lineage-mapper

376
a5c-aia5c-ai

Extracts and maps data lineage from various sources including SQL, dbt, Airflow, and Spark, generating comprehensive lineage graphs for impact analysis.

192 days ago

macos-sparkle-config

376
a5c-aia5c-ai

Configure Sparkle framework for macOS auto-updates with appcast, delta updates, and code signing

macossparkleautoupdate+2
192 days ago

auto-cdc

300
Databricks Cli Auto CdcDatabricks Cli Auto Cdc

Apply Change Data Capture (CDC) with apply_changes API in Spark Declarative Pipelines. Use when user needs to process CDC feeds from databases, handle upserts/deletes, maintain slowly changing dimensions (SCD Type 1 and Type 2), synchronize data from operational databases, or process merge operations.

192 days ago

streaming-data

296
ancolemanancoleman

Build event streaming and real-time data pipelines with Kafka, Pulsar, Redpanda, Flink, and Spark. Covers producer/consumer patterns, stream processing, event sourcing, and CDC across TypeScript, Python, Go, and Java. When building real-time systems, microservices communication, or data integration pipelines.

192 days ago

youtube-title

111
kenneth-liaokenneth-liao

Generate optimized YouTube video titles that maximize click-through rates by sparking curiosity and complementing thumbnails. This skill should be used when the user asks to create, improve, or brainstorm YouTube video titles, or when working on YouTube content that requires title optimization.

192 days ago

social-selling

96
gtmagentsgtmagents

Use when engaging prospects through LinkedIn, communities, and social channels to spark warm conversations and meetings.

192 days ago

displaying-streamlit-data

93
streamlitstreamlit

Displaying charts, dataframes, and metrics in Streamlit. Use when visualizing data, configuring dataframe columns, or adding sparklines to metrics. Covers native charts, Altair, and column configuration.

192 days ago

releasing-macos-apps

86
jamesrochabrunjamesrochabrun

Create notarized macOS app releases with Sparkle auto-updates, DMG installers, and GitHub releases. Use when releasing macOS apps, creating DMG files, notarizing apps, or setting up Sparkle updates. Handles version updates, code signing, notarization, and distribution.

192 days ago

spark-app-template

83
githubgithub

Comprehensive guidance for building web apps with opinionated defaults for tech stack, design system, and code standards. Use when user wants to create a new web application, dashboard, or interactive interface. Provides tech choices, styling guidance, project structure, and design philosophy to get users up and running quickly with a fully functional, beautiful web app.

192 days ago

data-pipeline-engineer

43
curiositechcuriositech

Expert data engineer for ETL/ELT pipelines, streaming, data warehousing. Activate on: data pipeline, ETL, ELT, data warehouse, Spark, Kafka, Airflow, dbt, data modeling, star schema, streaming data, batch processing, data quality. NOT for: API design (use api-architect), ML training (use ML skills), dashboards (use design skills).

etlsparkkafka+2
192 days ago

apache-spark-data-processing

40
manutejmanutej

Complete guide for Apache Spark data processing including RDDs, DataFrames, Spark SQL, streaming, MLlib, and production deployment

sparkbig-datadistributed-computing+3
192 days ago

senior-data-engineer

28
hainamchunghainamchung

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

192 days ago

dbt-migration-hive

28
sfc-gh-dflipposfc-gh-dflippo

Convert Hive/Spark/Databricks DDL to dbt models compatible with Snowflake. This skill should be used when converting views, tables, or UDFs from Hive, Spark, or Databricks to dbt code, generating schema.yml files with tests and documentation, or migrating HiveQL to follow dbt best practices.

192 days ago

PRD v0.1 Problem Framing

24
Mattgierhart Prd Driven Context Engineering Prd V01 Problem FramingMattgierhart Prd Driven Context Engineering Prd V01 Problem Framing

Transform vague product ideas into evidence-anchored problem statements for PRD v0.1 Spark. Triggers on starting new products/features, validating market opportunities, drafting PRD Why sections, or requests like "frame the problem", "define pain points", "write problem statement", "start v0.1", "what problem are we solving". Outputs structured problem tables with CFD evidence IDs.

192 days ago

prd-v01-user-value-articulation

24
mattgierhartmattgierhart

Transform validated pain points into articulated user value statements for PRD v0.1 Spark. Triggers on completing problem framing, defining user outcomes, articulating value propositions, or requests like "what value do users get", "define outcomes", "articulate the benefit", "finish v0.1", "pain to value", "what do they gain". Outputs CFD- entries tagged as value hypotheses with evidence tiers. Follows Problem Framing skill in workflow.

192 days ago

prd-v02-competitive-landscape-mapping

24
mattgierhartmattgierhart

Map the competitive landscape before positioning your product for PRD v0.2 Market Definition. Triggers on completing v0.1 Spark, analyzing competitors, researching market, or requests like "competitive analysis", "who else solves this", "market landscape", "what alternatives exist", "competitor research", "feature comparison". Outputs CFD- entries for competitive intelligence and BR- entries for positioning rules.

192 days ago

wdk

14
tethertotetherto

Tether Wallet Development Kit (WDK) for building non-custodial multi-chain wallets. Use when working with @tetherto/wdk-core, wallet modules (wdk-wallet-btc, wdk-wallet-evm, wdk-wallet-evm-erc-4337, wdk-wallet-solana, wdk-wallet-spark, wdk-wallet-ton, wdk-wallet-tron, ton-gasless, tron-gasfree), and protocol modules including swap (wdk-protocol-swap-velora-evm), bridge (wdk-protocol-bridge-usdt0-evm), and lending (wdk-protocol-lending-aave-evm). Covers wallet creation, transactions, token transfers, DEX swaps, cross-chain bridges, and DeFi lending/borrowing.

192 days ago

mashup

11
NickCrewNickCrew

Force-fit patterns from other domains to spark novel concepts.

192 days ago

elevator-pitch-techniques

11
mike-coulbournmike-coulbourn

Provides elevator pitch and verbal brand communication frameworks including Donald Miller's StoryBrand (SB7), Nancy Duarte's Sparkline, Chris Westfall's CLARITY, Andy Raskin's Strategic Narrative, Simon Sinek's Golden Circle, and time-based pitch structures (10s, 30s, 60s). Auto-activates during elevator pitch creation, one-liner development, brand pitch refinement, and verbal communication work. Use when discussing elevator pitches, one-liners, brand intros, verbal pitches, pitch coaching, spoken brand messages, or pitch variations.

192 days ago

apache-spark

9
TerminalSkillsTerminalSkills

Process large-scale data with Apache Spark. Use when a user asks to process big data, run distributed computations, build ETL pipelines, perform data analysis at scale, or use PySpark for data engineering.

192 days ago

Spark

8
simotasimota

Propose new features that leverage existing data and logic as a Markdown specification. Use when you need feature ideation, product planning, or feature proposals. Does not write code.

192 days ago

shopify-polaris-viz

6
toilahuonggtoilahuongg

Guide for creating data visualizations in Shopify Apps using the Polaris Viz library. Use this skill when building charts, graphs, dashboards, or any data visualization components that need to integrate with the Shopify Admin aesthetic. Covers BarChart, LineChart, DonutChart, SparkLineChart, and theming.

192 days ago

data-engineering

5
pluginagentmarketplacepluginagentmarketplace

ETL pipelines, Apache Spark, data warehousing, and big data processing. Use for building data pipelines, processing large datasets, or data infrastructure.

192 days ago

spark-optimization

5
agent-skills-hubagent-skills-hub

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

192 days ago

data-engineer

5
agent-skills-hubagent-skills-hub

Build scalable data pipelines, modern data warehouses, and real-time streaming architectures. Implements Apache Spark, dbt, Airflow, and cloud-native data platforms. Use PROACTIVELY for data pipeline design, analytics infrastructure, or modern data stack implementation.

192 days ago

scala-pro

5
agent-skills-hubagent-skills-hub

Master enterprise-grade Scala development with functional programming, distributed systems, and big data processing. Expert in Apache Pekko, Akka, Spark, ZIO/Cats Effect, and reactive architectures. Use PROACTIVELY for Scala system design, performance optimization, or enterprise integration.

192 days ago

spark-optimization

4
ngxtmngxtm

Optimize Apache Spark jobs with partitioning, caching, shuffle optimization, and memory tuning. Use when improving Spark performance, debugging slow jobs, or scaling data processing pipelines.

192 days ago

senior-data-engineer

4
ngxtmngxtm

World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt, Kafka, and modern data stack. Includes data modeling, pipeline orchestration, data quality, and DataOps. Use when designing data architectures, building data pipelines, optimizing data workflows, or implementing data governance.

192 days ago

spark-engineer

4
ngxtmngxtm

Use when building Apache Spark applications, distributed data processing pipelines, or optimizing big data workloads. Invoke for DataFrame API, Spark SQL, RDD operations, performance tuning, streaming analytics.

192 days ago

defi-lending-guide

3
SperaxSperax

Comprehensive guide to DeFi lending — protocol comparison, supply/borrow mechanics, health factor management, liquidation risks, and yield optimization. Covers Aave V3, Compound V3, Spark, and Radiant. Use when helping users lend, borrow, or manage lending positions.

192 days ago

make-fire

2
pjt222pjt222

Start and maintain a fire using friction, spark, and solar methods. Covers site selection, material grading (tinder/kindling/fuel), fire lay construction (teepee, log cabin, platform), ignition techniques (ferro rod, flint & steel, bow drill), flame nurturing, and Leave No Trace extinguishing. Use when needing warmth, light, or a signal in a wilderness setting, when boiling water for purification, when cooking foraged food, or in an emergency survival situation requiring heat or morale support.

192 days ago

make-fire

2
pjt222pjt222

Start and maintain a fire using friction, spark, and solar methods. Covers site selection, material grading (tinder/kindling/fuel), fire lay construction (teepee, log cabin, platform), ignition techniques (ferro rod, flint & steel, bow drill), flame nurturing, and Leave No Trace extinguishing. Use when needing warmth, light, or a signal in a wilderness setting, when boiling water for purification, when cooking foraged food, or in an emergency survival situation requiring heat or morale support.

192 days ago