Skip to content
OpenSmartRoute
Skillv1.0.0

large-data-with-dask

Specific optimization strategies for Python scripts working with larger-than-memory datasets via Dask.

by oimiragieo(0) 0 installs
Free
Sign in to install

Free account. Installing gives you the manifest plus copy-paste snippets.

See reviews

About

Imported from oimiragieo/agent-studio (.claude/skills/large-data-with-dask/SKILL.md). Install upstream with npx skills add oimiragieo/agent-studio --skill large-data-with-dask. Copyright stays with the author.

Large Data With Dask Skill

  • Consider using dask for larger-than-memory datasets.

Iron Laws

  1. ALWAYS call dask.compute() only once at the end of a pipeline — multiple intermediate compute() calls break the lazy evaluation graph and eliminate Dask's ability to fuse and parallelize operations.
  2. NEVER use df.apply(lambda ...) with Dask DataFrames for element-wise operations — Pandas-style apply forces row-by-row Python execution that bypasses Dask's vectorized C extensions and is slower than single-threaded Pandas.
  3. ALWAYS specify partition sizes explicitly when reading large datasets (blocksize= for CSV, chunksize= for Parquet) — auto-detected partition sizes frequently produce thousands of tiny partitions (slow scheduler overhead) or a single giant partition (no parallelism).
  4. NEVER call len(df) or df.shape on a Dask DataFrame without wrapping in compute() — these trigger immediate full dataset computation and negate lazy evaluation.
  5. ALWAYS use dask.distributed.Client for multi-machine or CPU-bound workloads — the default threaded scheduler serializes Python-heavy operations due to the GIL; the distributed scheduler bypasses this.

Anti-Patterns

Anti-Pattern Why It Fails Correct Approach
Multiple compute() calls in pipeline Breaks lazy graph; forces data to materialize and re-partition at each call Build complete computation graph first; call compute() once at the end
df.apply(lambda ...) on large DataFrames Row-by-row Python; GIL contention; slower than equivalent Pandas on single core Use vectorized Dask operations (map_partitions, assign, arithmetic operators)
Default blocksize on large CSV files 128MB default creates thousands of partitions for 100GB files; scheduler overhead dominates Set blocksize="256MB" or blocksize="1GB" for large files; profile optimal size
len(df) without compute() Triggers full dataset read and count; defeats lazy evaluation Use df.shape[0].compute() explicitly; only compute when size is truly needed
Threaded scheduler for CPU-bound work Python GIL serializes CPU computation across threads; no true parallelism Use dask.distributed.LocalCluster() or process-based scheduler for CPU tasks

Memory Protocol (MANDATORY)

Before starting:

cat .claude/context/memory/learnings.md

After completing: Record any new patterns or exceptions discovered.

ASSUME INTERRUPTION: Your context may reset. If it's not in memory, it didn't happen.

Use it

Copy one of these into your project. Installing also returns the manifest and these snippets.

yaml
targets:
  - https://api.opensmartroute.ai/api/v1/registry/oimiragieo-agent-studio-large-data-with-dask/manifest   # or paste the manifest below

Manifest

An Open Capability Manifest: the router reads it to know what this does, what it costs and when to pick it.

oimiragieo-agent-studio-large-data-with-dask.ocm.jsonjson
{
  "ocm": "1",
  "id": "oimiragieo-agent-studio-large-data-with-dask",
  "kind": "skill",
  "name": "large-data-with-dask",
  "description": "Specific optimization strategies for Python scripts working with larger-than-memory datasets via Dask.",
  "publisher": "oimiragieo",
  "version": "1.0.0",
  "capabilities": {
    "domains": [
      "coding"
    ],
    "tags": [
      "skill-md",
      "dask",
      "python",
      "parallel",
      "big-data",
      "dataframe",
      "skills-sh"
    ],
    "languages": [
      "en"
    ]
  },
  "quality_prior": 0.6,
  "examples": [
    "Specific optimization strategies for Python scripts working with larger-than-memory datasets via Dask."
  ],
  "primary": false,
  "metadata": {
    "source": {
      "provider": "skills.sh",
      "repository": "https://github.com/oimiragieo/agent-studio",
      "path": ".claude/skills/large-data-with-dask/SKILL.md",
      "ref": "HEAD",
      "url": "https://github.com/oimiragieo/agent-studio/blob/HEAD/.claude/skills/large-data-with-dask/SKILL.md",
      "key": "oimiragieo/agent-studio/.claude/skills/large-data-with-dask/SKILL.md"
    }
  },
  "instructions": "# Large Data With Dask Skill\n\n<identity>\nYou are a coding standards expert specializing in large data with dask.\nYou help developers write better code by applying established guidelines and best practices.\n</identity>\n\n<capabilities>\n- Review code for guideline compliance\n- Suggest improvements based on best practices\n- Explain why certain patterns are preferred\n- Help refactor code to meet standards\n</capabilities>\n\n<instructions>\nWhen reviewing or writing code, apply these guidelines:\n\n- Consider using dask for larger-than-memory datasets.\n  </instructions>\n\n<examples>\nExample usage:\n```\nUse",
  "cost": {
    "context_tokens": 934
  }
}

Fetch it by URL: GET /api/v1/registry/oimiragieo-agent-studio-large-data-with-dask/manifest?version=1.0.0

Reviews

Star ratings from people who tried it. One review per account; edit yours any time.

No reviews yet. Install it, try it, and be the first to rate it.