Imported from grahama1970/agent-stack-public (
skills/ingest-website/SKILL.md). Install upstream withnpx skills add grahama1970/agent-stack-public --skill ingest-website. Copyright stays with the author.
STOP. READ THIS ENTIRE SKILL.MD BEFORE CALLING ANY ENDPOINT.
Ingest Website
Single command to ingest a website into /memory as a RAG resource:
/fetcher (crawl) → /extractor (structured extraction) → /doc2qra (QRA extraction) → /taxonomy (bridge tags) → /memory (store)
Quick Start
# Ingest a website into memory
./run.sh ingest https://sparta.aerospace.org/resources --scope sparta
# Ingest with image download
./run.sh ingest https://example.com/docs --scope research --images
# Ingest specific URLs (not crawl)
./run.sh ingest --urls urls.txt --scope research
# Save fetched content to local files
./run.sh ingest https://example.com/docs --scope research --output-dir ./reference/
# Dry run (fetch + save locally, no memory storage)
./run.sh ingest https://example.com/docs --dry-run --output-dir ./docs/
Commands
ingest - Fetch website and store as RAG resource
| Option | Description |
|---|---|
URL |
Base URL to crawl (follows same-domain links) |
--urls FILE |
File with one URL per line (skip crawl) |
--scope NAME |
Memory scope (default: research) |
--images |
Download images (PNG, JPG, SVG) to output dir |
--output-dir DIR |
Save fetched pages as local markdown/images |
--max-pages N |
Max pages to fetch (default: 50) |
--depth N |
Max crawl depth from base URL (default: 2) |
--dry-run |
Fetch + save locally, skip memory storage |
--no-qra |
Store raw text chunks, skip QRA extraction |
--delay MS |
Delay between requests in ms (default: 500) |
Pipeline
- Crawl (
/fetcher): Fetch base URL, extract same-domain links, fetch linked pages up to--depthand--max-pages. - Images (optional): Download referenced images to
--output-dir. - Extract (
/extractor): Run structured extraction on fetched HTML via HTMLProvider. Produces clean markdown from raw HTML. - Save (optional): Write extracted content as markdown files to
--output-dir. - QRA Extract (
/doc2qra): Extract Q&A pairs from each page. doc2qra handles taxonomy + embedding internally. - Taxonomy (
/taxonomy): Extract federated bridge tags (Precision, Resilience, etc.) — used on--no-qrapath. - Store (
/memory): Batch learn QRA pairs with bridge tags and source tracking tags. - Post-hooks: Trigger
/embeddingvectors and edge proposal for graph traversal.
Output
Fetched: 13 pages from sparta.aerospace.org
Images: 4 downloaded to ./reference/images/
Stored: 88 QRA pairs in scope 'sparta'
Saved: 13 markdown files to ./reference/
Integration
/dogpilecan use this to turn research findings into permanent RAG/ux-labuses this for reference product documentation (SPARTA, ATT&CK)/learn-datalakecan compose this for web source ingestion