Skip to content
Skillv1.0.0

mlx-vlm

Run Vision Language Models locally on Apple Silicon Macs using MLX. Use when: installing mlx-vlm, running VLM inference (image + text → response), fine-tuning vision models on custom datasets, batch p

by terminalskills(0) 0 installs
Free
Sign in to install

Free account. Installing gives you the manifest plus copy-paste snippets.

See reviews

About

Imported from terminalskills/skills (skills/mlx-vlm/SKILL.md). Install upstream with npx skills add terminalskills/skills --skill mlx-vlm. Copyright stays with the author (Apache-2.0).

MLX-VLM — Vision Language Models on Apple Silicon

Overview

mlx-vlm runs vision-language models natively on Apple Silicon using the MLX framework. It supports inference and fine-tuning with unified memory — no GPU server needed.

Repo: Blaizzy/mlx-vlm
Requirements: macOS 14+, Apple Silicon (M1/M2/M3/M4), Python 3.10+

Installation

# Create virtual environment (recommended)
python3 -m venv ~/.venvs/mlx-vlm
source ~/.venvs/mlx-vlm/bin/activate

# Install
pip install mlx-vlm

For development:

git clone https://github.com/Blaizzy/mlx-vlm.git
cd mlx-vlm && pip install -e .

Supported Models

Model HuggingFace ID Best For
Pixtral mistral-community/pixtral-12b-240910 General vision, multi-image
Qwen2-VL Qwen/Qwen2-VL-7B-Instruct OCR, document understanding
Phi-3-Vision microsoft/Phi-3.5-vision-instruct Lightweight, fast inference
LLaVA-1.6 llava-hf/llava-v1.6-mistral-7b-hf Conversation about images
Llama-3.2-Vision meta-llama/Llama-3.2-11B-Vision-Instruct Strong general reasoning

Inference

CLI

# Single image analysis
python -m mlx_vlm.generate \
  --model mlx-community/pixtral-12b-240910-4bit \
  --image path/to/image.jpg \
  --prompt "Describe this image in detail" \
  --max-tokens 512

# Multi-image comparison
python -m mlx_vlm.generate \
  --model mlx-community/pixtral-12b-240910-4bit \
  --image img1.jpg img2.jpg \
  --prompt "Compare these two images"

Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model_path = "mlx-community/pixtral-12b-240910-4bit"
model, processor = load(model_path)

prompt = apply_chat_template(
    processor,
    config=model.config,
    prompt="What objects are in this image?",
    images=["product.jpg"],
)

output = generate(
    model, processor, prompt,
    images=["product.jpg"],
    max_tokens=512,
    temperature=0.7,
)
print(output)

Batch Processing

import os, csv
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("mlx-community/pixtral-12b-240910-4bit")
image_dir = "images/"

results = []
for filename in os.listdir(image_dir):
    if not filename.lower().endswith((".jpg", ".png", ".webp")):
        continue
    path = os.path.join(image_dir, filename)
    prompt = apply_chat_template(
        processor, config=model.config,
        prompt="Describe this product photo. Include: category, color, condition, key features.",
        images=[path],
    )
    desc = generate(model, processor, prompt, images=[path], max_tokens=256)
    results.append({"file": filename, "description": desc})

with open("descriptions.csv", "w", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=["file", "description"])
    writer.writeheader()
    writer.writerows(results)

Fine-Tuning

Prepare Dataset

Create JSONL with image paths and conversations:

{"image": "train/001.jpg", "conversations": [{"role": "user", "content": "Classify this product"}, {"role": "assistant", "content": "Category: Electronics, Subcategory: Headphones, Condition: New"}]}
{"image": "train/002.jpg", "conversations": [{"role": "user", "content": "Classify this product"}, {"role": "assistant", "content": "Category: Clothing, Subcategory: T-Shirt, Condition: Used - Good"}]}

Run Fine-Tuning (LoRA)

python -m mlx_vlm.lora \
  --model mlx-community/pixtral-12b-240910-4bit \
  --data ./dataset \
  --train-file train.jsonl \
  --valid-file val.jsonl \
  --num-layers 8 \
  --batch-size 1 \
  --epochs 3 \
  --lr 1e-5 \
  --adapter-path ./adapters

Inference with Fine-Tuned Adapter

python -m mlx_vlm.generate \
  --model mlx-community/pixtral-12b-240910-4bit \
  --adapter-path ./adapters \
  --image test.jpg \
  --prompt "Classify this product"

Cloud API Comparison

Factor mlx-vlm (Local) Cloud APIs (GPT-4V, Claude)
Cost $0 after hardware $0.01-0.04 per image
Privacy Data stays local Data sent to provider
Speed ~2-8s per image (M3 Max) ~1-3s per image
Offline Yes No
Custom models LoRA fine-tuning Limited / expensive
Quality Good (7-12B models) Excellent (frontier models)

Performance Tips

  • Use 4-bit quantized models (4bit in name) for 2-3x speedup with minimal quality loss
  • M3 Max / M4 Pro with 36GB+ RAM can run 12B models comfortably
  • For M1/M2 with 16GB, stick to 7B 4-bit models
  • Set MLX_METAL_JIT=1 for potential speedup on first run
  • Close memory-heavy apps before inference — unified memory is shared with system

Use it

Copy one of these into your project. Installing also returns the manifest and these snippets.

yaml
targets:
  - https://api.opensmartroute.ai/api/v1/registry/terminalskills-skills-mlx-vlm/manifest   # or paste the manifest below

Manifest

An Open Capability Manifest: the router reads it to know what this does, what it costs and when to pick it.

terminalskills-skills-mlx-vlm.ocm.jsonjson
{
  "ocm": "1",
  "id": "terminalskills-skills-mlx-vlm",
  "kind": "skill",
  "name": "mlx-vlm",
  "description": "Run Vision Language Models locally on Apple Silicon Macs using MLX. Use when: installing mlx-vlm, running VLM inference (image + text → response), fine-tuning vision models on custom datasets, batch processing images with local AI, comparing local VLM to cloud APIs (GPT-4V, Claude Vision), or working with LLaVA, Phi-3-Vision, Qwen2-VL, Pixtral, Llama-3.2-Vision on Mac.",
  "publisher": "terminalskills",
  "version": "1.0.0",
  "capabilities": {
    "domains": [
      "general"
    ],
    "tags": [
      "skill-md",
      "mlx",
      "vision",
      "apple-silicon",
      "skills-sh"
    ],
    "languages": [
      "en"
    ]
  },
  "quality_prior": 0.6,
  "examples": [
    "Run Vision Language Models locally on Apple Silicon Macs using MLX. Use when: installing mlx-vlm, running VLM inference (image + text → response), fine-tuning vision models on custom datasets, batch processing images with local AI, comparing local VLM to cloud APIs (GPT-4V, Claude Vision), or working with LLaVA, Phi-3-Vision, Qwen2-VL, Pixtral, Llama-3.2-Vision on Mac."
  ],
  "primary": false,
  "metadata": {
    "source": {
      "provider": "skills.sh",
      "repository": "https://github.com/terminalskills/skills",
      "path": "skills/mlx-vlm/SKILL.md",
      "ref": "HEAD",
      "url": "https://github.com/terminalskills/skills/blob/HEAD/skills/mlx-vlm/SKILL.md",
      "key": "terminalskills/skills/skills/mlx-vlm/SKILL.md"
    },
    "compatibility": "macOS 14+, Apple Silicon, Python 3.10+",
    "license": "Apache-2.0"
  },
  "instructions": "# MLX-VLM — Vision Language Models on Apple Silicon\n\n## Overview\n\nmlx-vlm runs vision-language models natively on Apple Silicon using the MLX framework. It supports inference and fine-tuning with unified memory — no GPU server needed.\n\n**Repo:** `Blaizzy/mlx-vlm`  \n**Requirements:** macOS 14+, Apple Silicon (M1/M2/M3/M4), Python 3.10+\n\n## Installation\n\n```bash\n# Create virtual environment (recommended)\npython3 -m venv ~/.venvs/mlx-vlm\nsource ~/.venvs/mlx-vlm/bin/activate\n\n# Install\npip install mlx-vlm\n```\n\nFor development:\n```bash\ngit clone https://github.com/Blaizzy/mlx-vlm.git\ncd mlx-vlm && ",
  "cost": {
    "context_tokens": 1192
  }
}

Fetch it by URL: GET /api/v1/registry/terminalskills-skills-mlx-vlm/manifest?version=1.0.0

Reviews

Star ratings from people who tried it. One review per account; edit yours any time.

No reviews yet. Install it, try it, and be the first to rate it.