Skip to content

Marketplace

Everything your AI needs, in one place.

Ready-made agents, skills, personas, prompts, templates and tools. Each one is checked before it goes live, works with any model, and installs in a click. Rate what you use so the best rises to the top.

146.7K
listings
1
installs
0
reviews
40.4K
publishers
70 results
Skill

phoenix-observability

Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production A

by orchestra-researchskills.sh
Not rated yet
Free
Skill

evaluating-cosmos-policy

Evaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting up cosmos-policy for robot manipulation evaluation, running headless GPU evaluations with EGL rendering,

by orchestra-researchskills.sh
Not rated yet
Free
Skill

advanced-evaluation

Design and operate LLM-as-a-Judge evaluation systems using direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment. Use

by shipshitdevskills.sh
Not rated yet
Free
Skill

evaluation

Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choices, or measuring improvements over time.

by shipshitdevskills.sh
Not rated yet
Free
Skill

skill-comply

Measure whether agents actually follow a skill, rule, command, or agent definition by deriving expected behaviors, running representative scenarios, and comparing observed action timelines against the

by shipshitdevskills.sh
Not rated yet
Free
Skill

eval-designer

Use this skill when building evaluation frameworks to measure LLM quality, safety, accuracy, or alignment including test suites, human eval rubrics, automated evals, and metrics design. Not for traini

by nickcrewskills.sh
Not rated yet
Free
Skill

evaluator-optimizer

Iterative refinement workflow for polishing code, documentation, or designs through systematic evaluation and improvement cycles. Use when refining drafts into production-grade quality.

by nickcrewskills.sh
Not rated yet
Free
Skill

rag-eval

NVIDIA RAG Blueprint evaluation guidance for measuring retrieval and answer quality with stable datasets, baselines, and reproducible scoring workflows.

by practicalswanskills.sh
Not rated yet
Free
Skill

recommender-evaluation

Evaluate recommender system quality for the CSX4207 Vinyl Record Store. Use whenever you must compute or report recommender metrics (Precision@k, Recall@k, HitRate@k, MRR, MAP@k, NDCG@k, coverage, div

by practicalswanskills.sh
Not rated yet
Free
Skill

assess

Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions

by yonatangrossskills.sh
Not rated yet
Free
Skill

golden-dataset

Golden dataset lifecycle patterns for curation, versioning, quality validation, and CI integration. Use when building evaluation datasets, managing dataset versions, validating quality scores, or inte

by yonatangrossskills.sh
Not rated yet
Free
Skill

testing-llm

LLM and AI testing patterns — mock responses, evaluation with DeepEval/RAGAS, structured output validation, and agentic test patterns (generator, healer, planner). Use when testing AI features, valida

by yonatangrossskills.sh
Not rated yet
Free
Skill

agent-evaluation

LLM-as-judge evaluation framework with 5-dimension rubric (accuracy, groundedness, coherence, completeness, helpfulness) for scoring AI-generated content quality with weighted composite scores and evi

by oimiragieoskills.sh
Not rated yet
Free
Skill

pipeline-evaluator

Evaluates completed agent pipelines across 5 scoring dimensions and produces a composite verdict with actionable recommendations

by oimiragieoskills.sh
Not rated yet
Free
Skill

evaluating-llms-harness

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training p

by ovachieverskills.sh
Not rated yet
Free
Skill

stockfish-analyzer

国际象棋引擎分析工具,提供最佳走法推荐、局面评估和多种走法选择分析。支持FEN字符串直接输入分析。

by aiskillstoreskills.sh
Not rated yet
Free
Skill

clawvard-agent-eval

Take the Clawvard entrance exam, report the result, and optionally save the agent identity token with explicit user confirmation.

by okxskills.sh
Not rated yet
Free
Skill

genkit

Route Firebase AI feature work into either direct app/client Firebase AI Logic SDK integration or a server-owned Genkit workflow. Use when a web, mobile, backend, or full-stack feature needs model cal

by akillnessskills.sh
Not rated yet
Free
Skill

langsmith

Route LangSmith work into one workflow packet before touching SDK code. Use when the user needs LangSmith tracing, offline evals, annotation/review queues, prompt-registry decisions, audit/gap review,

by akillnessskills.sh
Not rated yet
Free
Skill

opik

Run Comet's Opik — open-source LLM observability, evaluation, and optimization — from one routing-first skill: install the Python/TypeScript SDK, stand up a server (Comet.com cloud, Docker Compose via

by akillnessskills.sh
Not rated yet
Free
Skill

scientific-llm-benchmarks

A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery.

by akillnessskills.sh
Not rated yet
Free
Skill

arize

You are an expert in Arize and its open-source Phoenix library for AI observability. You help developers monitor LLM applications with tracing, evaluation, embedding analysis, drift detection, and ret

by terminalskillsskills.sh
Not rated yet
Free
Skill

braintrust

You are an expert in Braintrust, the evaluation and observability platform for AI applications. You help developers run systematic evaluations, compare model versions, track experiments, log productio

by terminalskillsskills.sh
Not rated yet
Free
Skill

deepeval

Expert guidance for DeepEval, the open-source framework for unit testing LLM applications. Helps developers write test cases, define custom metrics, and integrate LLM quality checks into CI/CD pipelin

by terminalskillsskills.sh
Not rated yet
Free
1

Find

Search or browse by kind. Every card shows who made it, how many people installed it and what they think.

2

Install

One click. You get a manifest the router understands, plus copy-paste snippets for the CLI, Python and YAML.

3

Rate and publish

Leave a star rating after you have used it. Made something useful? Publish it - free listings go live immediately.

Prefer the terminal? osr stack apply registry://starter installs the starter template.