Skip to content
Skillv1.0.0

senior-data-engineer

Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or tr

by borghei(0) 0 installs
Free
Sign in to install

Free account. Installing gives you the manifest plus copy-paste snippets.

See reviews

About

Imported from borghei/claude-skills (engineering/senior-data-engineer/SKILL.md). Install upstream with npx skills add borghei/claude-skills --skill senior-data-engineer. Copyright stays with the author (MIT + Commons Clause).

Senior Data Engineer

Generate pipeline configurations (Airflow, Prefect, Dagster), validate data quality with profiling and anomaly detection, and optimize SQL/Spark performance with actionable recommendations.

Core Capabilities

  • Pipeline generation — Airflow/Prefect/Dagster DAG code for batch and incremental loads, with DAG validation.
  • Data quality — schema validation, profiling, anomaly detection, data contracts, and Great Expectations suite generation.
  • ETL/ELT optimization — SQL and Spark analysis, partition strategy, and query cost estimation per warehouse.
  • Architecture decisions — batch vs streaming and warehouse vs lakehouse trade-off frameworks.
  • Reliability patterns — incremental watermarks, dead letter queues, freshness checks, and schema-drift detection.

When to Use

  • Designing a data architecture or choosing batch vs streaming / warehouse vs lakehouse.
  • Building or generating Airflow/Spark/dbt pipelines.
  • Adding data-quality checks or data contracts.
  • Optimizing slow ETL/ELT queries or troubleshooting pipeline failures.

Clarify First

Before generating pipelines, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Orchestrator — Airflow / Prefect / Dagster (--type; changes the generated DAG code)
  • Source, destination & load mode — systems involved and batch vs incremental (--source/--destination/--mode; shapes the pipeline)
  • Data-quality expectations — the schema and contracts to enforce (drives the Great Expectations suite generation)

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Quick Start

# Generate an Airflow DAG for incremental PostgreSQL -> Snowflake
python scripts/pipeline_orchestrator.py generate \
  --type airflow --source postgres --destination snowflake \
  --tables orders,customers --mode incremental --schedule "0 5 * * *"

# Validate data quality against a schema
python scripts/data_quality_validator.py validate data.csv \
  --schema schema.json --detect-anomalies --json

# Profile a dataset
python scripts/data_quality_validator.py profile data.csv --json

# Optimize a slow SQL query
python scripts/etl_performance_optimizer.py analyze-sql query.sql \
  --warehouse snowflake --json

# Estimate query cost
python scripts/etl_performance_optimizer.py estimate-cost query.sql \
  --warehouse bigquery --stats data_stats.json --json

Tools

Tool Subcommands Purpose
pipeline_orchestrator.py generate, validate, template Generate Airflow/Prefect/Dagster pipeline code, validate DAGs
data_quality_validator.py validate, profile, generate-suite, contract, schema Schema validation, profiling, anomaly detection, Great Expectations
etl_performance_optimizer.py analyze-sql, analyze-spark, optimize-partition, estimate-cost, template SQL/Spark optimization, partition strategy, cost estimation

All subcommands support --json for machine-readable output and --output for file writing.

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

Integration Points

Skill Integration
senior-data-scientist Feature engineering consumes curated mart data
senior-ml-engineer ML pipelines depend on feature store tables
senior-devops CI/CD for dbt, Airflow deployment, container orchestration
senior-architect Architecture reviews for lakehouse vs warehouse decisions
code-reviewer Pipeline code reviews for DAGs, dbt models, Spark jobs

Use it

Copy one of these into your project. Installing also returns the manifest and these snippets.

yaml
targets:
  - https://api.opensmartroute.ai/api/v1/registry/borghei-claude-skills-senior-data-engineer/manifest   # or paste the manifest below

Manifest

An Open Capability Manifest: the router reads it to know what this does, what it costs and when to pick it.

borghei-claude-skills-senior-data-engineer.ocm.jsonjson
{
  "ocm": "1",
  "id": "borghei-claude-skills-senior-data-engineer",
  "kind": "skill",
  "name": "senior-data-engineer",
  "description": "Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or troubleshooting pipeline failures.",
  "publisher": "borghei",
  "version": "1.0.0",
  "capabilities": {
    "domains": [
      "general"
    ],
    "tags": [
      "skill-md",
      "airflow",
      "spark",
      "data-pipelines",
      "warehousing",
      "etl",
      "skills-sh"
    ],
    "languages": [
      "en"
    ]
  },
  "quality_prior": 0.6,
  "examples": [
    "Data engineering for batch and streaming pipelines with Airflow, dbt, Spark, and Kafka. Use when designing data architectures, building pipelines, adding data-quality checks, optimizing ETL/ELT, or troubleshooting pipeline failures."
  ],
  "primary": false,
  "metadata": {
    "source": {
      "provider": "skills.sh",
      "repository": "https://github.com/borghei/claude-skills",
      "path": "engineering/senior-data-engineer/SKILL.md",
      "ref": "HEAD",
      "url": "https://github.com/borghei/claude-skills/blob/HEAD/engineering/senior-data-engineer/SKILL.md",
      "key": "borghei/claude-skills/engineering/senior-data-engineer/SKILL.md"
    },
    "license": "MIT + Commons Clause"
  },
  "instructions": "# Senior Data Engineer\n\nGenerate pipeline configurations (Airflow, Prefect, Dagster), validate data quality with profiling and anomaly detection, and optimize SQL/Spark performance with actionable recommendations.\n\n## Core Capabilities\n\n- **Pipeline generation** — Airflow/Prefect/Dagster DAG code for batch and incremental loads, with DAG validation.\n- **Data quality** — schema validation, profiling, anomaly detection, data contracts, and Great Expectations suite generation.\n- **ETL/ELT optimization** — SQL and Spark analysis, partition strategy, and query cost estimation per warehouse.\n- **Arc",
  "cost": {
    "context_tokens": 1210
  }
}

Fetch it by URL: GET /api/v1/registry/borghei-claude-skills-senior-data-engineer/manifest?version=1.0.0

Reviews

Star ratings from people who tried it. One review per account; edit yours any time.

No reviews yet. Install it, try it, and be the first to rate it.