Skip to content

Marketplace

Everything your AI needs, in one place.

Ready-made agents, skills, personas, prompts, templates and tools. Each one is checked before it goes live, works with any model, and installs in a click. Rate what you use so the best rises to the top.

146.7K
listings
1
installs
0
reviews
40.4K
publishers
55 results
Skill

multimodal-llm

Vision, audio, video generation, and multimodal LLM integration patterns. Use when processing images, transcribing audio, generating speech, generating AI video (Kling v3, Sora 2, Veo 3.1 std/lite/fas

by yonatangrossskills.sh
Not rated yet
Free
Skill

clip

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content m

by ovachieverskills.sh
Not rated yet
Free
Skill

llama-factory

Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support

by ovachieverskills.sh
Not rated yet
Free
Skill

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal supp

by ovachieverskills.sh
Not rated yet
Free
Skill

llava

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, vi

by ovachieverskills.sh
Not rated yet
Free
Skill

sentence-transformers

Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and m

by ovachieverskills.sh
Not rated yet
Free
Skill

whisper

OpenAI's general-purpose speech recognition model. Supports 99 languages, transcription, translation to English, and language identification. Six model sizes from tiny (39M params) to large (1550M par

by ovachieverskills.sh
Not rated yet
Free
Skill

detecting-indirect-prompt-injection

Detect and defend against indirect prompt injection hidden in web pages, documents, and images consumed by an agent, via content extraction (HTML/PDF/OCR), normalization, and scanning with LLM Guard's

by mukul975skills.sh
Not rated yet
Free
Skill

glm4v-analyze-image

智谱AI的视觉语言模型,用于图像分析、内容识别和视觉问答

by aiskillstoreskills.sh
Not rated yet
Free
Skill

god-tibo-imagen

Generate images using Codex's ChatGPT backend with zero production dependencies. Reuses existing local Codex authentication (~/.codex/auth.json) — no new credentials needed. Supports CLI (gti command)

by akillnessskills.sh
Not rated yet
Free
Skill

google-ai-studio

Google AI Studio and Gemini API for multimodal AI. Use when you need multimodal AI (text + image + video + audio), long context up to 1M tokens, code generation with Gemini, grounding with Google Sear

by terminalskillsskills.sh
Not rated yet
Free
Skill

openai-realtime

Build voice-enabled AI applications with the OpenAI Realtime API. Use when a user asks to implement real-time voice conversations, stream audio with WebSockets, build voice assistants, or integrate Op

by terminalskillsskills.sh
Not rated yet
Free
Skill

llama-factory

Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support

by chen-yu-haoGitHub
Not rated yet
Free
Skill

llama-factory

Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support

by gabrielmoreiraGitHub
Not rated yet
Free
Skill

youtube-learn

Phân tích video (YouTube, LinkedIn, Facebook, X, TikTok) theo hướng Belief Archaeology, kết hợp Multimodal phân tích hình ảnh và thế giới quan của người nói.

by dotanminhGitHub
Not rated yet
Free
Skill

stable-diffusion-image-generation

State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or bu

by majiayu000GitHub
Not rated yet
Free
Skill

stable-diffusion-image-generation

State-of-the-art text-to-image generation with Stable Diffusion models via HuggingFace Diffusers. Use when generating images from text prompts, performing image-to-image translation, inpainting, or bu

by Mishit18GitHub
Not rated yet
Free
Skill

segment-anything-model

Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks

by majiayu000GitHub
Not rated yet
Free
Skill

segment-anything-model

Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks

by tamaguskoGitHub
Not rated yet
Free
Skill

segment-anything-model

Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks

by nota-americaGitHub
Not rated yet
Free
Skill

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal supp

by diegosouzapwGitHub
Not rated yet
Free
Skill

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal supp

by nota-americaGitHub
Not rated yet
Free
Skill

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal supp

by synthetic-sciencesGitHub
Not rated yet
Free
Skill

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal supp

by wangyw07GitHub
Not rated yet
Free
1

Find

Search or browse by kind. Every card shows who made it, how many people installed it and what they think.

2

Install

One click. You get a manifest the router understands, plus copy-paste snippets for the CLI, Python and YAML.

3

Rate and publish

Leave a star rating after you have used it. Made something useful? Publish it - free listings go live immediately.

Prefer the terminal? osr stack apply registry://starter installs the starter template.