Imported from zhaoxiaofei/paper-refine (
paper-skills/paper-revise/SKILL.md). Install upstream withnpx skills add zhaoxiaofei/paper-refine --skill paper-revise. Copyright stays with the author.
Targeted Revision — adress_issues
Target venue, article type and journal (configurable). The rule set comes from the pipeline's
venue profile (venue_profiles/<id>.json, selected with set-venue), the venue's content type
from set-article-type (each type carries its own limits; a type the profile has no numbers for
is counted and named from the venue's own table) and the journal from set-journal; Nature
Biotechnology is only the default profile. Numbers quoted below are that default's Article
type -- use the profile and type of the run you are in.
Paths
FINDINGS_MD=./review/findings.md— human-readable report from paper-reviewFINDINGS_JSON=./review/findings.json— machine-readable; CANONICAL list of finding IDsEXTRA_JSON=./review/round2/findings_extra.json— discovery-round findings; merge if presentORIGINALS= the reviewed submission directory (default./non-revised; use the directory recorded in the findings report if present). READ-ONLY, always.REVISED=./revised·WORK=./revised/work·CODE=./codeZOTERO_SKILL=/mnt/d/software/plugins/plugins/zotero/skills/zotero/scripts/zotero.py(override or absent → manual instructions)ZOT_CLI=zot(pyzotero-cli) and the$zotero-useskill — reference resolution and, under the operator's Zotero policy, guarded citation-field edits (rule E2)
If both findings files are missing AND no findings are in the current conversation: STOP and ask the user to run paper-review (identify_issues) first. Explicit paths the user gives override defaults. All output in English. Length rules inherited from the review (check id M19): the VENUE PROFILE the run resolved supplies the limits, relaxed by that profile's own margins (the numbers are stated in the run's prompt and in venue_profiles/<id>.json; the pipeline assumes no journal — a profile that declares no limit is counted with no cap); an over-cap section is brought within the cap by removing redundancy, repeated statistics and non-meaning-bearing hedging ONLY, never by deleting scientific content and never by stripping a hedge that carries the claim's own strength (that is an overclaim, not a shortening), and text within the cap is never cut for length. An incompressible section goes to MANUAL_STEPS.md. The cover letter's persuading part is measured against the user's configured 300-500-word preference (the default profile's venue publishes no cover-letter limit) and is a Minor formatting item; the letter's TOTAL content (salutation, body, disclosures and signature) is capped at 650 words by default (the operator's cap, also a Minor formatting item, never a gate); figure-legend word counts (M18) are recorded but never cut when no proxy cap is configured.
Mission
Apply every validated finding. Originals are never modified; every edit lands in REVISED/ (documents) or CODE/ (analysis code). Every failure mode of a revision task is enumerable too — so this skill runs on the same devices as the review: a ledger where every finding ID must get a row (no silent skip), an edit plan where every edit maps to a finding ID, a propagation map for long-range consistency, a diff log proving locality, a rescan proving no new errors, and a coverage table as the acceptance gate.
The review's rewrite-parity findings (M25–M29), source-hierarchy findings (M30) and architecture findings (J5) are part of that list. They report what a from-scratch rewrite fixes as a side-effect (artwork/text term parity, house-style conventions, claim→evidence coverage, sibling-definition symmetry, caption-promise parity, and scope-level organization), so this stage must resolve them the way a rewrite would where the finding allows it: align the editable surface corpus-wide for M25–M29, align the WRITTEN side with the producer the hierarchy names for M30 (rule E12; rule C owns a code fix, raw_data/ stays read-only), and use the E11 scoped-restructuring licence for J5 (reorder/split/merge/transition edits inside the finding's declared scope, content frozen, one WORK/RESTRUCTURE_<id>.md per finding). manual-required is the last resort, not the default: it applies only when no editable surface can be aligned without inventing or deleting content.
Hard rules
- Originals read-only. Record sha256 of EVERY original file BEFORE any edit (
WORK/checksums_before.txt); never modify anything outsideREVISED/andCODE/; re-verify byte-identity at the end (V4) and confirm it in the final report. - Never invent. No fabricated content, citations, data, results, or accession numbers. Missing mandatory items are scaffolded only with clearly marked
[AUTHOR TO COMPLETE: ...]placeholders, every one listed in the final report. Scientific claims, interpretations, and the strength of conclusions are never altered on your own initiative (rule E6). - No silent skip. EVERY finding ID —
F-*from findings.json andX-*from findings_extra.json (kept as separate ID namespaces) — appears exactly once in the REVISION LEDGER with a final status. "Not addressed" is not an allowed status. Every mechanical step must produce its artifact inWORK/; a step without its artifact is not done. - One edit per finding instance. Never aggregate edits. Every edit maps to a finding ID, or to a rescan ID
R-xxxfor issues introduced or discovered during revision. - ENUMERATE → ARTIFACT → APPLY → VERIFY for every mechanical step. Script-enumerate all affected occurrences (the review skill's
extract_occurrences.pyis bundled for this — copy it intoWORK/), record every occurrence as a row (including rows later judged OK), act row by row, then verify row by row. Propagation by attention alone is forbidden. - Zotero fields and the library (rule E2). Reference resolution is read-only by default; a citation field may be edited ONLY in the
REVISED/copy and only under the live-field rules (parent item keys, uniquecitationIDs, preserved baseline, snapshot-then-validate with the bundled validator, no automatic Zotero Refresh). The library itself is never written unless the operator explicitly enabled writes: then ONE field of ONE existing item, after a written proposal and an independent re-fetch (zot items update ... --last-modified auto), never a create/delete/bulk edit. Preserve existing styles, numbering, equations, table layouts, and figure placement; no reflow, restyling, or "improving" untouched text — except inside a J5 finding's declared scope, where E11 licenses exactly those operations under its content-freeze and artifact guards. - The EVIDENCE areas (
raw_data/,human_review_feedback/) are not submission documents. They are carried byte-for-byte (an agent that drops one gets it restored by the pipeline's recovery layer, with a warning) and they are NEVER version-token-renamed — the content-hash token applies to the submission documents only, and evidence file names stay exactly as they are. Never edit, add, drop, rename or regenerate anything inside them. A finding whose quote lives only inside an evidence area — a data table, a figure source, an editor's decision letter or a reviewer's report — is not a finding about the submission: discard it in R1 with that recorded rationale, unless it is an M30 row whose WRITTEN side is quoted from a submission document and the evidence file appears only as the producer. The same applies to a feedback or response-to-reviewers document kept elsewhere in the corpus (it is skipped by name). Editors'/reviewers' feedback is external prose: use it as evidence of what the review requires — and, in a journal revision mode, as the source of the concerns the response letter answers — never as the authors' words and never as an editable surface.
Steps (in order; each step's artifacts complete before the next)
- R0 — Setup/safety. Inventory + checksums; load both findings files; build the A1 ledger skeleton programmatically so no finding can be dropped; print the findings count as a checkpoint.
- R1 — Re-verify every finding against the sources; assign a verdict with a location-checked rationale (empty rationale = invalid). False positives → discard, but only with a recorded concrete rationale (misreading, correct cross-reference, guideline-version difference) — never silently. Ambiguous wording a reviewer could misread → verdict
clarification. A quote that exists only in an EVIDENCE area (raw_data/,human_review_feedback/) — a data table, a figure source, an editor's or reviewer's feedback file — is not the submission's text: discard with that rationale (rule 7), unless the finding is M30 and the written side is quoted from a submission document. Unverifiable (unreadable Zotero field, number only inside a read-only figure) →manual-required— UNLESS the finding is an M25–M29 parity/convention finding whose EDITABLE side can be aligned (E11), or an M30 source-hierarchy finding whose editable WRITTEN side is the wrong side (E12): a read-only artwork file or a read-onlyraw_data/producer does not make the finding manual when the caption/main text/SI legend/table cell that carries the written claim can be aligned. not-found-in-source → discard, unless the text lives in a read-only/unparseable file → manual-required. - R2 — Editable copies into
REVISED/, ONE content-hash version token per package (doc/docx/tex/bib/md/txt/xlsx; never pdf/png). Copy every editable document intoREVISED/and give the whole package one version token: the 7-character content-hash printed by the bundledscripts/revision_token.py REVISED/(the first 7 hex characters of SHA-256 over the sorted content digests of the payload files; your own reports, the.tracked.docx/.before-after.docxauxiliaries andwork/are excluded; file NAMES do not enter the hash, and existing version-token references inside file contents are normalized to<VERSION>first, so applying the token does not change it). Apply the token to every editable document: REPLACE its trailing version token (-a.docx→-<token>.docx,manuscript_v2.tex→manuscript_<token>.tex,refs02-o.bib→refs02-<token>.bib) or APPEND it when the basename has none (refs.bib→refs-<token>.bib). The legacy letter/digit INCREMENT rule is WITHDRAWN: never turn-a.docxinto-b.docx, and never turnmanuscript_v2.texintomanuscript_v3.tex. A trailing NUMBER that is part of a document's identity (SI-Table-1.csv,Figure-3.xlsx) is not a version token: it stays part of the name, and you never shift a numbered family or renumber its siblings. Never mutate individual characters of a name (manuscript.mdmust not becomemanuscripu.md). Run the tool before the rename, then re-run it with--verify <token>after the rename/repoint pass and require theOKline (if it does not match, the content changed after the token was computed — recompute and re-apply). The only other rename allowed is the collision fallback: a same-named file already exists inREVISED/(re-run or two source versions) → the NEW copy gets_rev2,_rev3, … before the extension, recorded in A3 asrename applied = yes; leave the existing copy's name as it is. If (and only if) a revised path differs from the original basename (the token, the fallback, or a legacy package that already carries an incremented token), run the rename sweep: script-enumerate every reference to the original (pre-rename) basename (LaTeX\input/\include/\includegraphics/\addbibresource/\bibliography, build files, scripts) and repoint it to the new name; record each in A5. In the pure-collision case the original name still exists inREVISED/and keeps its references — nothing to repoint; the sweep then just verifies and records zero repoints. LaTeX compile check (pdflatex + bibtex/biber or the project's Makefile) if a toolchain exists; otherwise syntax/label sanity check, stated explicitly. Legacy.doc→ convert to.docxinside REVISED/ (same basename, new extension;_rev2on collision), note the conversion and flag it for user confirmation in the final report — do not block the run on it. Findings whose text lives in read-only files (pdf/png) cannot be edited in this workflow: mark themmanual-requiredin the ledger — UNLESS the finding is an M25–M29 rewrite-parity finding whose EDITABLE counterpart (caption, main text, SI legend) can be aligned instead (E11), or an M30 finding whose written side carries the wrong value (E12); the read-only file then gets a regeneration/author-decision step in MANUAL_STEPS.md, not the finding. - R3 — Edit plan (A4), sequenced: (i) Critical factual/ethical/completeness → (ii) consistency propagation → (iii) logic/clarity/repetition → (iv) grammar/terminology → (v) formatting. If you cannot quote the before-text exactly, return to R1 — you have not located the finding.
- E — Apply edits one at a time in plan order, under rules E1–E12 in
references/edit_rules.md(precision, Zotero fields, missing items, plagiarism/AI content, tracked-changes auxiliary.tracked.docx, scientific-judgement guard, E11 — rewrite-parity findings: align the editable surface for M25–M29, scoped restructuring for J5 with oneWORK/RESTRUCTURE_<id>.mdper finding — and E12 — source-hierarchy findings: align the WRITTEN side with the authoritative producer, rule C for a code fix,raw_data/read-only). - P — Propagation (A6). Priority when a mismatched number/label/term/claim is corrected: supplementary tables > supplementary figures/notes > main figures & legends > Methods > main text > abstract > cover letter. For EVERY correction: script-extract all occurrences of the old AND new values across the whole revised corpus; update every occurrence; verify each row. Cross-check A6 against the A2 baseline — a baseline occurrence missing from A6 is a missed propagation; fix it.
- C — Code revisions (only if analysis code is in scope). Scientific integrity rule: never change analysis code merely to make outputs match manuscript numbers — details in
references/edit_rules.md. - V — Validate. V1 round-trip integrity (artifact
WORK/roundtrip_check.md) · V2 locality via diff log (A7: hunks inside a J5 scope are mapped through the finding id +WORK/RESTRUCTURE_<id>.md; every other hunk must map to a finding id) · V3 full mechanical rescan of the revised corpus (A8: the M1–M30 sweeps from paper-review — including the M26 convention re-run and the J5 scope check — PLUS any sweep definitions in./review/round2/new_sweeps.md) · V4 checksum re-verification (artifactWORK/checksums_after.txt) · V5 final outputs.
Artifacts A1–A9 column specifications and status vocabularies: references/ledger.md.
Final outputs (write incrementally throughout)
REVISED/CHANGELOG.md— for every finding ID: document, exact before → after text, status. Generated from A1 + A7.REVISED/MANUAL_STEPS.md— for every manual-required finding: which file, which section/paragraph/field, what to check, exact tool steps (e.g., verifying a citation field in Word via the Zotero plugin; applying a proposed Zotero library correction with item key, field, current value and proposed value), and follow-ups (e.g., regenerate a figure from CODE/ and update every dependent number).REVISED/REVISION_REPORT.md+REVISED/revision_report.json— the ledger, the coverage table (one row per finding ID — F-* and X-* — plus each R-xxx: id | final status | edit IDs | evidence of completion (diff hunk / artifact row); a finding ID without a coverage row, or a row without evidence, means the task is not finished), counts by status, the placeholder list, pending scientific-judgement decisions with proposed alternative wordings, the checksum confirmation, and guideline-version uncertainties to re-check against the current author guide.
Execution discipline
- Prefer scripts over attention for every enumeration; keep scripts in
WORK/(re-runnable). The paper-review scripts (convert_corpus.py,extract_occurrences.py,extract_acronyms.py,extract_citations.py,extract_numbers.py) can be reused — run them againstWORK/corpusbuilt fromREVISED/. - Batch file by file to bound context, but every artifact spans the whole corpus; merge before verifying.
- Long sessions: maintain
WORK/STATE.md(current step, pending edit IDs, open questions) so the workflow resumes without loss.
Acceptance checks (for the human, after the run)
- The coverage table has a row with evidence for every finding ID (F-* and X-*).
revised/work/contains the A1–A9 artifacts (A9 only when analysis code is in scope).- The diff log shows no unmapped hunks.
- The checksum statement appears in REVISION_REPORT.md.
- Every
[AUTHOR TO COMPLETE: ...]in the revised files appears in the placeholder list. - Every citation-field edit is recorded in CHANGELOG.md with the item key and validated against the pre-edit snapshot (
validate_zotero_docx.py); every over-cap abstract/main text (M19) is either within the relaxed caps or listed in MANUAL_STEPS.md with the reason, and the legend/cover-letter rows (M18/M19) are recorded (compressed only where a cap or the user's preference allows it without losing content). - Every M25–M29 finding has either a corpus-wide aligned edit (with the M26 re-run proving no deviating occurrence survives) or a recorded
unable —/manual reason; every M30 finding has the written side aligned to the authoritative producer (or a code fix under rule C with its rerun note, or a recorded manual reason with the exact file/value); every J5 finding has aWORK/RESTRUCTURE_<id>.mdartifact or a written manual reordering step. A parity/convention finding carried asmanual-requiredwhile an editable occurrence still exists is a failed resolution.
