Imported from michaeldgreenphd/website (
scripts/AGENTS.md). Install upstream withnpx skills add michaeldgreenphd/website --skill scripts. Copyright stays with the author.
scripts/ — the data pipeline
Read together with the root AGENTS.md. These five scripts are run by
.github/workflows/update-scholar.yml; its schedule, pins and commit step
are described in .github/workflows/AGENTS.md.
fetch_scholar.py: four attempts of 90 seconds each (direct, then free-proxy rotation), each in a subprocess that is killed on timeout. A Google block ends in a::warning::and keeps the cached numbers — routine, not a failure. scholarly'sMaxTriesExceededExceptionandDOSException, timeouts, and any other error raised on a Google block page (consent, "unusual traffic", captcha) or a page that is not Scholar's (a free proxy's junk) count as blocks; the page is the last one scholarly received, kept by a hook on itsNavigator._get_soup. An error on a page Scholar served (recognised by its site-widegs_/gsc_markup or "Google Scholar" title, so a profile redesign still counts) or before any page arrived exits 1 as a real bug. A successful fetch warns if the page check stops recognising Scholar's page; then updateBLOCK_PAGE_MARKERS/SCHOLAR_MARKUP_RE. Refusing an empty or zero payload is the only sanity guard; do not add a never-lower or percentage rule (settled decision, rootAGENTS.md). BumpCHART_STYLE_VERSIONwhenever the chart's palette or styling changes; otherwise the PNGs are re-rendered only when the data changes.fetch_substack.py: keeps the cached posts on a transient failure but fails loudly when the cache is older thanMAX_CACHE_AGE_DAYS(30). A successful fetch refreshes the cache'supdateddate weekly (REFRESH_AFTER_DAYS) even when the posts are unchanged, so that age counts from the last good fetch; expect one small data commit a week.fetch_orcid.py: keeps its cache on any failure. Publication links are built only ashttps://doi.org/<DOI>; a work without a well-formed DOI is listed without a link. Each work keeps its ORCIDsource(not rendered), and works added to or removed from the record are printed as::notice::lines in the run summary.render_snapshot.pysplices generated HTML between the<!-- data:NAME -->markers inindex.html. It is deterministic — no timestamps, no run-dependent ordering — and a second run on unchanged data prints "index.html data blocks unchanged." It omits works typedconference-abstractandconference-posteron purpose.stage_data.pyruns only in the workflow'spublishjob, standard library only:import DIRcopies the fetched data files in after checking each one, andcheckrefuses the commit if anything but the data files and the<!-- data: -->blocks ofindex.htmlchanged.- Keep both escaping layers: the fetch scripts'
clean_text()(tags stripped, entities decoded, before writing JSON) and the renderer'sclean()/safe_url()(escaped on output; links only for http(s) URLs). The root reviewer checklist explains why. - A new script must be added to
on.push.pathsin the workflow or it never gets a test run. - After any change:
python -m py_compile scripts/*.py, then run the renderer twice.