Skip to content

Marketplace

Everything your AI needs, in one place.

Ready-made agents, skills, personas, prompts, templates and tools. Each one is checked before it goes live, works with any model, and installs in a click. Rate what you use so the best rises to the top.

143.8K
listings
1
installs
0
reviews
38.7K
publishers
32 results
Skill

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving models with limited GP

by wang030327-botGitHub
(0)
0Free
Skill

gguf-quantization

GGUF format and llama.cpp quantization for efficient CPU/GPU inference. Use when deploying models on consumer hardware, Apple Silicon, or when needing flexible quantization from 2-8 bit without GPU re

by OpenCovenGitHub
(0)
0Free
Skill

hqq-quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deployin

by tianhao909GitHub
(0)
0Free
Skill

hqq-quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deployin

by qcmuuGitHub
(0)
0Free
Skill

hqq-quantization

Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deployin

by tadod12GitHub
(0)
0Free
Skill

gptq

Post-training 4-bit quantization for LLMs with minimal accuracy loss. Use for deploying large models (70B, 405B) on consumer GPUs, when you need 4× memory reduction with <2% perplexity degradation, or

by asadbekXodjayevGitHub
(0)
0Free
Skill

vllm

Use when self-hosting an open-weight LLM for high-throughput concurrent serving with vLLM — running an OpenAI-compatible endpoint, splitting a model across GPUs with tensor or pipeline parallelism, lo

by ericriscoGitHub
(0)
0Free
Skill

deploy-edge-ai-model

使用 Google AI Edge Gallery、TensorFlow Lite、ONNX Runtime 和 MediaPipe 将机器学习模型部署到边缘设备。涵盖模型量化(INT8/INT4)、使用 Gemma 4 模型的设备端推理、通过 AI Edge Gallery 进行 Android/iOS 部署、硬件代理 选择(GPU/NPU/DSP),以及在受限设备上的性能基准测试。在因延迟、成

by pjt222GitHub
(0)
0Free
1

Find

Search or browse by kind. Every card shows who made it, how many people installed it and what they think.

2

Install

One click. You get a manifest the router understands, plus copy-paste snippets for the CLI, Python and YAML.

3

Rate and publish

Leave a star rating after you have used it. Made something useful? Publish it - free listings go live immediately.

Prefer the terminal? osr stack apply registry://starter installs the starter template.