Imported from wikisqueaks/edmontons-lost-waterways (
.positai/AGENTS.md). Install upstream withnpx skills add wikisqueaks/edmontons-lost-waterways --skill .positai. Copyright stays with the author.
R Research Project Template
This is a GitHub template repository for reproducible research in R. It is not an R package — it is a project scaffold that users clone (via "Use this template") to start new research projects.
Core Technologies
targets— pipeline orchestration and caching (_targets.Rdefines the pipeline)- Quarto — exploratory notebooks (
notebooks/) and manuscript (manuscript/)
Project Structure
project/
├── _targets.R # Pipeline definition; sources all functions in R/
├── data/
│ ├── raw/ # Immutable source data (git-ignored)
│ ├── interim/ # Intermediate data (git-ignored)
│ └── derived/ # Final processed data (git-ignored)
├── R/ # Promoted analysis functions (no library() calls, no top-level code)
├── notebooks/ # Exploratory Quarto notebooks (freeze: auto)
├── make-outputs/
│ ├── figures/ # make-fig-[description].R scripts
│ └── tables/ # make-tbl-[description].R scripts
├── outputs/ # Pipeline-generated artifacts (git-ignored)
│ ├── figures/
│ ├── tables/
│ └── models/
├── manuscript/ # Quarto manuscript project
│ ├── manuscript.qmd
│ ├── pre-render.R # Copies outputs/figures/ and outputs/tables/ before render
│ ├── figures/ # Auto-copied at render time (git-ignored; never edit manually)
│ ├── tables/ # Auto-copied at render time (git-ignored; never edit manually)
│ ├── rendered_notebooks/ # Rendered HTML notebooks linked from manuscript
│ ├── references.bib
│ └── styles.css
├── references/ # Data dictionaries, literature, manuals
├── docs/ # Project and developer documentation
└── README.md
Workflow (6 stages)
Understand this workflow so you can guide the user through it and avoid helping them work against it.
Stage 1 — Explore in notebooks
Each investigation gets a .qmd file in notebooks/. No structure required at this stage. Notebooks use freeze: auto so they only re-execute when changed. This is where the user thinks freely.
Stage 2 — Promote mature logic to R/
When analysis stabilizes, extract core logic into a named function in R/. Functions must:
- Accept inputs as arguments and return a value
- Write files and return their output path(s) for file-based targets
- Contain no top-level executing code and no
library()calls
Stage 3 — Register in _targets.R
Add the function as a target. Use format = "file" for file-based inputs and outputs so targets tracks file contents and reruns downstream steps on change. Run with targets::tar_make().
Stage 4 — Consume targets in notebooks
Once promoted, replace computations in the notebook with targets::tar_load(target_name). The notebook becomes a consumer, not a compute engine.
Stage 5 — Produce outputs with make-outputs/
Generate final figures and tables using scripts in make-outputs/figures/ and make-outputs/tables/. Naming convention:
make-fig-[description].R— writes tooutputs/figures/make-tbl-[description].R— writes tooutputs/tables/
Each script reads from data/ or outputs/models/ and writes to exactly one location.
Stage 6 — Write the manuscript
The manuscript lives in manuscript/ and is a Quarto manuscript project. It does not run analysis — it consumes static outputs only. Before each render, pre-render.R copies everything from outputs/figures/ and outputs/tables/ into local manuscript/figures/ and manuscript/tables/. Reference figures and tables using those local paths. The manuscript/figures/ and manuscript/tables/ directories are git-ignored — never place files there manually.
Naming Conventions
Targets (_targets.R) use descriptive verb_noun snake_case that reads like a summary of what the target does:
- Input file targets: noun phrases describing the data (
ab_wildfire,study_boundary) - Processing targets:
verb_noundescribing the transformation (rasterize_wildfire,compute_stand_age,mask_to_green_zone) - Avoid generic names like
data,result,output, or names that only make sense with surrounding context
The pipeline list in _targets.R should read like a plain-language description of the workflow from top to bottom.
Functions in R/ follow a verb-first convention: build_*, process_*, create_*, extract_*, rasterize_*, mask_*. File names in R/ should match or clearly describe their primary function (e.g., build_stand_age_rasters.R).
Key Rules
outputs/is the single source of truth for generated artifacts.manuscript/figures/andmanuscript/tables/are always regenerated at render time — never place files there manually.- All data directories and
outputs/are git-ignored.
Publishing the Manuscript
- Ensure figures and tables are up to date in
outputs/figures/andoutputs/tables/. If not, run the relevantmake-outputs/scripts first. - Render exploratory notebooks so they are available for linking:
Rendered HTML is deposited intocd notebooks quarto rendermanuscript/rendered_notebooks/. - Publish from
manuscript/. Thepre-render.Rscript runs automatically:cd manuscript quarto publish gh-pages --no-prompt - To render locally without publishing:
quarto renderfrommanuscript/.
rendered_notebooks/** is included as a Quarto resource passthrough so notebooks are bundled with the published site.
Key Commands
| Command | Purpose |
|---|---|
targets::tar_make() |
Run the pipeline |
targets::tar_visnetwork() |
Visualize the dependency graph |
targets::tar_load(name) |
Load a cached target into the environment |
targets::tar_outdated() |
See which targets need to rerun |
quarto render (in notebooks/) |
Render exploratory notebooks to manuscript/rendered_notebooks/ |
quarto render (in manuscript/) |
Render manuscript locally |
quarto publish gh-pages --no-prompt (in manuscript/) |
Publish manuscript to GitHub Pages |
Setup for New Projects
After cloning this template:
- Replace
README.mdwithPROJECT_README.mdcontent (customize for the project).
