Skip to content

Marketplace

Everything your AI needs, in one place.

Ready-made agents, skills, personas, prompts, templates and tools. Each one is checked before it goes live, works with any model, and installs in a click. Rate what you use so the best rises to the top.

146.7K
listings
1
installs
0
reviews
40.4K
publishers
70 results
Skill

langfuse

You are an expert in Langfuse, the open-source LLM engineering platform. You help developers trace LLM calls, evaluate output quality, manage prompts, track costs and latency, run experiments, and bui

by terminalskillsskills.sh
Not rated yet
Free
Skill

langsmith

Monitor, trace, debug, and evaluate LLM applications with LangSmith. Use when a user asks to trace LLM calls, debug chain executions, evaluate AI output quality, set up LLM observability, monitor agen

by terminalskillsskills.sh
Not rated yet
Free
Skill

langtrace

You are an expert in Langtrace, the open-source observability platform for LLM applications built on OpenTelemetry. You help developers trace LLM calls, RAG pipelines, agent tool use, and chain execut

by terminalskillsskills.sh
Not rated yet
Free
Skill

market-evaluation

Evaluate market opportunities before building — assess demand, competition, margins, and timing using Personal MBA's proven 10-factor framework. Use when: evaluating a new business idea, deciding whic

by terminalskillsskills.sh
Not rated yet
Free
Skill

prompt-tester

Design, test, and iterate on AI prompts systematically using structured evaluation criteria. Use when building AI features, optimizing agent instructions, comparing prompt variants, or evaluating outp

by terminalskillsskills.sh
Not rated yet
Free
Skill

promptfoo

Test and evaluate LLM prompts systematically with Promptfoo — open-source eval framework. Use when someone asks to "test my prompts", "evaluate LLM output", "Promptfoo", "prompt regression testing", "

by terminalskillsskills.sh
Not rated yet
Free
Skill

ragas

Expert guidance for Ragas, the framework for evaluating Retrieval-Augmented Generation pipelines. Helps developers measure and improve the quality of their RAG systems across retrieval accuracy, answe

by terminalskillsskills.sh
Not rated yet
Free
Skill

weave

You are an expert in Weave, the lightweight toolkit by Weights & Biases for tracking and evaluating AI applications. You help developers trace LLM calls, evaluate outputs, compare model versions, trac

by terminalskillsskills.sh
Not rated yet
Free
Skill

uxui-principles

Evaluate interfaces against 168 research-backed UX/UI principles, detect antipatterns, and inject UX context into AI coding sessions.

by JantonioFCGitHub
Not rated yet
Free
Skill

uxui-principles

Evaluate interfaces against 168 research-backed UX/UI principles, detect antipatterns, and inject UX context into AI coding sessions.

by DorianGalloGitHub
Not rated yet
Free
Skill

uxui-principles

Evaluate interfaces against 168 research-backed UX/UI principles, detect antipatterns, and inject UX context into AI coding sessions.

by inskillflowGitHub
Not rated yet
Free
Skill

uxui-principles

Evaluate interfaces against 168 research-backed UX/UI principles, detect antipatterns, and inject UX context into AI coding sessions.

by eery1677-labGitHub
Not rated yet
Free
Skill

uxui-principles

Evaluate interfaces against 168 research-backed UX/UI principles, detect antipatterns, and inject UX context into AI coding sessions.

by iamanacarolinarezendeGitHub
Not rated yet
Free
Skill

rag-eval

Filesystem RAG benchmarks: corpus/, train.json, evaluate_rag.py (RAGAS quality). Not for prod monitoring, latency/throughput benchmarking (use rag-perf), or evals outside this repo layout.

by SolizardkingGitHub
Not rated yet
Free
Skill

langsmith-observability

LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems, or building systemati

by majiayu000GitHub
Not rated yet
Free
Skill

evaluating-code-models

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language suppor

by John-Wang-0809GitHub
Not rated yet
Free
Skill

uxui-evaluator

Evaluate interface descriptions against 168 research-backed UX/UI principles. Returns structured findings with severity, remediation, and business impact. API key optional — enriched output requires u

by majiayu000GitHub
Not rated yet
Free
Skill

nemo-evaluator-sdk

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud p

by ihatesea69GitHub
Not rated yet
Free
Skill

evaluating-code-models

评估HumanEval、MBPP、MultiPL-E和15+基准的代码生成模型,并使用pass@k 度量衡。当设定代码模型的基准时,可以比较编码能力,测试多种语言的支持,或者测量代码生成质量。HuggingFace导板所使用的BigCode Project的工业标准。

by lxhb2GitHub
Not rated yet
Free
Skill

nemo-evaluator-sdk

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud p

by NzettodessGitHub
Not rated yet
Free
Skill

eval-designer

Use this skill when building evaluation frameworks to measure LLM quality, safety, accuracy, or alignment including test suites, human eval rubrics, automated evals, and metrics design. Not for traini

by xcrrrGitHub
Not rated yet
Free
Skill

skill-creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill

by felinicsGitHub
Not rated yet
Free
1

Find

Search or browse by kind. Every card shows who made it, how many people installed it and what they think.

2

Install

One click. You get a manifest the router understands, plus copy-paste snippets for the CLI, Python and YAML.

3

Rate and publish

Leave a star rating after you have used it. Made something useful? Publish it - free listings go live immediately.

Prefer the terminal? osr stack apply registry://starter installs the starter template.