Imported from youngfish42/FL-paper-update-tracker (
AGENTS.md). Install upstream withnpx skills add youngfish42/FL-paper-update-tracker. Copyright stays with the author.
Agent Guidance for FL-paper-update-tracker
This file is intended for automated coding agents (and human maintainers) who need to understand, modify, or extend the project. Keep it up-to-date whenever architecture, logic, or conventions change.
Project Overview
FL-paper-update-tracker is an automated bot that tracks new Federated Learning (FL) papers published in 40+ top-tier computer-science conferences and journals. It is a satellite project of Awesome-FL.
High-Level Workflow
- GitHub Actions runs the tracker once per day (cron:
0 0 * * *) and on every push tomain. src/main.pyreadsconfig.yaml→ takesdblp.keywords(e.g.[federate, FedAvg, ...]) anddblp.queries(plain venue restrictions), assembles fully URL-encoded DBLP search topics, and queries the DBLP search API.- Extracted paper metadata is filtered by year (last 3 years + next 1 year) and deduplicated by
eefield. - New papers (not yet in
cached/dblp.yaml) are collected, formatted as Markdown, and written to theGITHUB_ENVvariableMSG. scripts/convert_cache_to_md.pyregeneratesFL-Papers.mdfrom the updated cache.- If new papers exist, the action
JasonEtco/create-an-issue@v2creates a GitHub Issue using.github/issue-template.md. - Both
cached/dblp.yamlandFL-Papers.mdare committed back to the repo so that subsequent runs know what has already been reported.
Tech Stack
- Language: Python 3.8+
- Core Dependencies (see
requirements.txt):fire– CLI scaffoldingrequests– HTTP calls to DBLP APIloguru– structured loggingezkfg– lightweight config loaderpyyaml– cache read/write
- CI/CD: GitHub Actions (Ubuntu runner)
- License: Apache 2.0
Directory Structure
.
├── .github/
│ ├── workflows/watch.yml # GitHub Actions workflow definition
│ └── issue-template.md # Nunjucks template for auto-created issues
├── cached/
│ └── dblp.yaml # Persistent cache of already-reported papers
├── scripts/
│ ├── convert_cache_to_md.py # Converts cache to structured Markdown (domain-specific maps)
│ ├── fetch_abstracts.py # Backfill/refresh paper abstracts via external APIs
│ ├── fetch_dois.py # Backfill missing DOIs via DBLP / Crossref / Semantic Scholar
│ ├── dedup_cache_by_title.py # Deduplicate cache entries by title
│ └── dedup_cache_global.py # Global cross-topic deduplication for the cache
├── src/
│ ├── main.py # Entry point: assembles topics from keyword+queries, orchestrates API calls
│ └── utils.py # Helper functions: API call, parsing, formatting, dedup
├── config.yaml # keyword, plain queries (venues), and mail targets
├── FL-Papers.md # Structured Markdown output of all tracked papers
├── requirements.txt # Python dependencies
├── README.md # Human-facing documentation (EN + CN)
├── TECHNICAL.md # Deployment and configuration guide
└── AGENTS.md # This file
Key Logic & Conventions
1. Year Filtering
- Location:
src/main.py(callsfilter_items_by_yearinsrc/utils.py) - Rule: Only papers whose
yearfalls in[current_year - 3, current_year + 1]are kept. - Example: In 2026, valid years are 2023–2027.
- Agent Note: If you change this window, update both the implementation and this document.
2. Deduplication
- Location:
src/utils.py+src/main.py - Three-stage dedup:
deduplicate_items_by_ee(per-topic) — Within a single DBLP query result, papers with identicaleeare deduplicated. Rationale: DBLP sometimes returns multiple records for the same paper with minor author-name differences (e.g.,Ming Hu 0003vsMing Hu).deduplicate_items_by_title(per-topic) — Within a single query result, papers with identicaltitleare also deduplicated. Rationale: DBLP may list the same paper multiple times under differenteeURLs (e.g., preprint vs. proceedings version).- Global cross-topic dedup (in
main.py) — A paper that has already been cached under any topic is skipped when processing subsequent topics. Rationale: DBLP search API can return the same paper for multiple venue queries (e.g., a keyword match may cross venue boundaries), so a globalseen_ee/seen_titleset prevents the same paper from being stored under multiple topic keys.
- Agent Note: Do not switch back to full-dict comparison (
item not in cached_items) unless you also normalize author names.
3. Cache Format (cached/dblp.yaml)
- Top-level keys: URL-encoded DBLP search topics (e.g.,
federate%20venue%3ADAC%3A:orFedAvg%20venue%3AICML%3A:). - Each key maps to a list of paper dicts with fields:
author,title,venue,year,type,access,key,doi,ee,url,abstract,abstract_cn,related_code.abstractmay be empty for legacy entries; usescripts/fetch_abstracts.pyto backfill it.abstract_cnis the Chinese translation ofabstract, auto-generated via Qwen-MT-plus.related_codestores the first GitHub repository URL detected inabstract(or empty string if none).
- The file is overwritten after every successful run.
- Agent Note: If you add new fields to the paper dict, ensure backward compatibility; old cache entries missing the new field should be handled gracefully.
4. Message Formatting (get_msg)
- Location:
src/utils.py get_msggenerates a Markdown block for a single venue: heading with[+N]count plus an unordered list of papers in the form:- {title}. [PUB]({ee})src/main.pycollects new papers from all keyword + query combinations, then merges them by venue (name_topic, e.g.journals/ml). Papers found under the same venue by different keywords are combined into a single block, deduplicated byee/title, and counted together.get_topic_short_nameextracts the venue short name (the segment after the last/, or the whole name if no/) from a topic URL. Used for the Issue title.format_title_topicsjoins short names with,and truncates to ≤80 characters, appending等N个when truncated.
5. Issue Title Construction
- Template:
.github/issue-template.md - Title:
Paper Update [{{ env.ISSUE_TITLE_TOPICS }}] @ {{ date | date('YYYY-MM-DD') }} - The environment variable
ISSUE_TITLE_TOPICSis populated insrc/main.pyand passed through the workflowenvblock.
6. Configuration (config.yaml)
dblp:
url: https://dblp.org/search/publ/api?q={}&format=json&h=1000
keywords:
- federate
- gradient inversion
- FedAvg
- ...
queries:
- "venue:IJCAI:"
- ...
mails:
- "im.young@foxmail.com"
keywords: List of research-domain keywords. The first keyword is treated as primary; the rest are secondary.- In automatic runs (
primary_only=True), only the primary keyword scans all venues initially. Secondary keywords are then run only on venues where the primary keyword discovered new papers, reducing total API calls. - In manual runs (
primary_only=False), allkeywords × queriescombinations are scanned.
- In automatic runs (
queries: List of plain-text DBLP venue restrictions. The runner automatically URL-encodes each query and prepends the keyword before calling the API.mails: The first email address (mails[0]) is used as thecontact_emailfor the Crossref API User-Agent (mailto:...), which is recommended for polite API access. Additional addresses are reserved for future mail-notification features.- Agent Note: When adding a new venue, find its DBLP query syntax (venue code or stream ID) and append it to
dblp.queries. - Backward compatibility: Old single-string
keywordfield is still supported and is automatically wrapped into a one-element list.
Maintenance Notes
Adding a New Venue
- Find the DBLP venue code (e.g.,
venue:ICMLorstreamid:journals/pami). - Append the plain query string to
config.yamlunderdblp.queries(e.g.,venue:ICML:). The runner handles URL encoding automatically. - Update
scripts/convert_cache_to_md.pyif you want the new venue mapped to a specific category inFL-Papers.md. - Update
README.md(both EN and CN sections) to list the new venue. - Update this
AGENTS.mdif the change affects architecture or conventions.
Backfilling Abstracts for Existing Papers
- A standalone script
scripts/fetch_abstracts.pyis provided to backfillabstractfields for papers already incached/dblp.yaml. - It queries five APIs in order until a non-empty abstract is found:
- OpenReview (priority, batch-only for OpenReview-source papers) — Inside
fetch_abstract_for_papers()a prefill stage (_prefill_openreview_abstracts) scans every paper whoseeecontainsopenreview.net, extracts the forum id (forum?id=XXX/pdf?id=XXX), and asks the OpenReview v2 batch endpointapi2.openreview.net/notes?ids=A,B,C(≤100 ids per call). Any id missed by v2 falls back to per-id queries against v2 then v1 (api.openreview.net/notes?forum=XXX). Both API content schemas are handled (v1plain string vsv2{value: "..."}). Inspired by the PaperVault project. - Crossref — by DOI, with
contact_emailin the User-Agent header. - Semantic Scholar — by DOI.
- arXiv — by title via
export.arxiv.org/api/query, parsing Atom XML. arXiv enforces a minimum 3-second interval between requests. - OpenAlex (final fallback) — by DOI via
api.openalex.org/works/doi:.... OpenAlex returnsabstract_inverted_index, which is reconstructed into plain text. Amailtoparameter is appended whencontact_emailis available.
- OpenReview (priority, batch-only for OpenReview-source papers) — Inside
- All queries use rate limiting, timeout handling (10s + exponential backoff), and automatic newline cleaning.
- The cache is backed up to
cached/dblp.yaml.bakbefore each overwrite;*.bakfiles are ignored by git (see.gitignore). - Usage:
# Process current-year papers (default) python scripts/fetch_abstracts.py # Process all years python scripts/fetch_abstracts.py --year all # Retry previously failed entries (empty abstracts) python scripts/fetch_abstracts.py --retry-failed - Automatic abstract fetching:
src/main.pyalready callsfetch_abstract_for_papers()for every batch of new papers before saving the cache, so newly discovered papers get their abstracts filled automatically during the daily GitHub Actions run. - Automatic related code extraction: Immediately after a non-empty English
abstractis obtained,extract_github_links()scans the text forhttps://github.com/<user>/<repo>patterns and stores the first match inrelated_code. This happens insidefetch_abstract_for_papers()(both in the OpenReview prefill stage and the per-paper fallback loop), so new papers get their code links without extra steps. - Automatic Chinese translation: After a non-empty English
abstractis obtained,translate_abstracts_for_papers()calls Qwen-MT-plus (via the Alibaba Cloud Bailian OpenAI-compatible API) to translate the abstract into Chinese, storing it asabstract_cn. Translation is skipped if theDASHSCOPE_API_KEYenvironment variable is missing, and individual translation failures do not block the pipeline.
Backfilling DOIs for Existing Papers
- A standalone script
scripts/fetch_dois.pyis provided to backfill missingdoifields for papers already incached/dblp.yaml. - It queries APIs in the following priority order until a non-empty DOI is found:
- DBLP API (primary) — re-queries the paper by its
key(e.g.conf/dac/ChandrasekaranE22) to check whether a DOI has been assigned since the initial fetch. This is the most authoritative source. - Crossref (fallback) — searches by title via
api.crossref.org/works?query.title=.... - Semantic Scholar (final fallback) — searches by title via
api.semanticscholar.org/graph/v1/paper/search.
- DBLP API (primary) — re-queries the paper by its
- All queries use rate limiting, timeout handling (10s + exponential backoff), and title fuzzy-matching verification (
is_title_match) to avoid assigning an incorrect DOI. - The cache is backed up to
cached/dblp.yaml.bakbefore each overwrite;*.bakfiles are ignored by git (see.gitignore). - Usage:
# Process current-year papers with missing DOI (default) python scripts/fetch_dois.py # Process all years python scripts/fetch_dois.py --year all # Re-fetch DOI for all papers (even those that already have one) python scripts/fetch_dois.py --retry-all - A GitHub Actions workflow
.github/workflows/backfill-dois.ymlallows manual triggering from the repository UI.
Backfilling Related Code for Existing Papers
- A standalone script
scripts/fetch_related_code.pyis provided to backfill missingrelated_codefields for papers already incached/dblp.yaml. - It scans the
abstracttext of each paper for GitHub repository URLs (https://github.com/<user>/<repo>), cleans trailing punctuation, and stores the first match inrelated_code. - The cache is backed up to
cached/dblp.yaml.bakbefore each overwrite;*.bakfiles are ignored by git (see.gitignore). - Usage:
# Process current-year papers (default) python scripts/fetch_related_code.py # Process all years python scripts/fetch_related_code.py --year all # Retry previously failed entries (empty related_code) python scripts/fetch_related_code.py --retry-failed - Automatic related code extraction:
src/main.pyalready callsextract_github_links()insidefetch_abstract_for_papers()for every batch of new papers, so newly discovered papers get their code links filled automatically during the daily GitHub Actions run.
Switching to a Different Research Domain
The tracker is domain-agnostic. To pivot from Federated Learning to any other field (e.g., diffusion models, LLMs, reinforcement learning):
- Change the keywords in
config.yaml:dblp: keywords: - diffusion # or LLM, "reinforcement learning", etc. - Adjust the venue list under
dblp.queriesto match the venues relevant to the new domain. - Update
scripts/convert_cache_to_md.py:VENUE_MAP— map DBLP raw venue names to your preferred display names.CATEGORY_MAP— assign each display name to a category.CATEGORY_ORDERandVENUE_ORDER— control the output ordering.
- Reset the cache by deleting or renaming
cached/dblp.yamlso the next run treats all fetched papers as new. - (Optional) Update
README.mdandTECHNICAL.mdto reflect the new domain.
No changes to src/main.py or the GitHub Actions workflow are required (unless you want to adjust the primary_only behavior).
Changing Message Format
- Edit
src/utils.py→get_msg. - If you modify the Markdown structure, verify that the issue template renders correctly in GitHub.
- Keep the
aggregatedparameter behavior:True= heading only,False= heading + list. - The default list item format is
- {title}. [[PUB]({ee})]. Whenrelated_codeis present, it appends[[CODE]({related_code})]after[PUB].
Changing Issue Template
- Edit
.github/issue-template.md. - The template engine is Nunjucks (via
JasonEtco/create-an-issue@v2). Available variables:{{ date }}{{ env.MSG }}{{ env.ISSUE_TITLE_TOPICS }}
- If you introduce a new
envvariable, also update.github/workflows/watch.ymlto pass it through.
Changing Workflow Schedule
- Edit
.github/workflows/watch.yml→schedule.cron. - Default is daily at 00:00 UTC+8 (
0 0 * * *).
Cache Corruption / Manual Reset
- Simply delete
cached/dblp.yaml(or remove specific topic keys). - The next run will treat all papers as "new" and recreate the cache.
Testing Locally
cd src
python main.py run --env=dev
In dev mode the script will:
- Load the cache
- Query DBLP with all keywords and venues
- Print logs to stdout
- Not write to
GITHUB_ENV
You can inspect aggregated_msg and msg in the logs to preview the issue content.
Additional CLI flags:
--primary_only— Emulates the automatic cron/push behavior: primary keyword scans all venues first, secondary keywords only scan venues that produced new papers.--all_years— Disables the year filter and skips abstract fetching / translation (used by theFetch All Years Papersworkflow).
Common Pitfalls
- Rate Limiting: DBLP does not publish explicit rate limits. The code uses a conservative base interval (
6 + random(0.5, 2.5)seconds for search requests) plus stricter exponential backoff with jitter and Retry-After support when requests fail or are rate-limited. Do not weaken this behavior. - YAML Encoding:
cached/dblp.yamlis written withallow_unicode=True. Editing it manually with an editor that strips Unicode may corrupt author names. - Empty
eeField: Some older DBLP entries lack anee. The deduplication logic skips emptyeevalues, so those papers are always retained. This is intentional to avoid data loss. - Topic String Parsing:
get_topic_short_namerelies onsplit(":")[-2]followed bysplit("/")[-1]. If DBLP changes its topic URL format, this will break.
Contact & References
- Parent project: Awesome-FL
- Upstream inspiration: dblp-watcher
- DBLP API docs: https://dblp.org/faq/How+to+use+the+dblp+search+API.html
