Imported from Yuling-Tian/mineru-doc-parser-skill (
AGENTS.md). Install upstream withnpx skills add Yuling-Tian/mineru-doc-parser-skill. Copyright stays with the author.
MinerU Document Parser
This repository contains a document parsing skill/tool that converts PDF, Word, PPT, and image files into clean Markdown using the MinerU cloud OCR API.
Key Files
mineru-doc-parser/scripts/parse_doc.py— Main parser script (Python CLI)mineru-doc-parser/SKILL.md— Skill definition with full workflow instructions
How to Use
When the user wants to parse a document (PDF/DOCX/PPTX/PNG/JPG) to Markdown, run the parser script:
python mineru-doc-parser/scripts/parse_doc.py "<file_path>" --lang en
Workflow
1. Detect page count (PDF only)
pip install PyMuPDF -q 2>/dev/null
python -c "import fitz; d=fitz.open('<path>'); print(d.page_count); d.close()"
2. Choose API tier
- ≤ 20 pages → Agent API (no token needed)
- > 20 pages → Ask user: truncate first 20 pages or full parse?
3. Large PDFs (>20 pages)
For full parse, check for a persisted token:
python mineru-doc-parser/scripts/parse_doc.py --check-token
- Exit 0 → token exists, proceed
- Exit 1 → ask user for token (get one free at https://mineru.net/apiManage)
4. Run parser
python mineru-doc-parser/scripts/parse_doc.py "<file>" \
--lang <en|ch|korean|japan|latin> \
[--page-range "1-20"] \
[--token "<token>"] \
[--save-token] \
[--ocr] [--no-table] [--no-formula]
5. Report results
Tell the user: output path, character count, API tier used, whether token was persisted.
API Reference
| Tier | Endpoint | Limit | Token |
|---|---|---|---|
| Agent | https://mineru.net/api/v1/agent |
≤ 20 pages | No |
| Precise | https://mineru.net/api/v4 |
Unlimited | Yes |
Config
- Token persistence:
~/.mineru/config.json(or~/.claude/mineru_config.jsonfor Claude Code users) - Env var override:
MINERU_API_TOKEN
Requirements
- Python ≥ 3.8
requests(pip install requests)PyMuPDF(pip install PyMuPDF, optional — for PDF page count detection)
