Imported from zerlinpi/amazon-ads-skills (
AGENTS.md). Install upstream withnpx skills add zerlinpi/amazon-ads-skills. Copyright stays with the author.
AGENTS.md
Repository purpose
This repository contains reusable Amazon Ads Agent Skills. The canonical business logic lives under skills/<skill-name>/SKILL.md and must remain portable across Codex, Claude Code, WorkBuddy, and other Agent Skills-compatible runtimes.
First-run / installation guidance
For a fresh runtime, repository URL, or manual checkout, read docs/GETTING-STARTED.md before inventing runtime-specific install paths. The stable default is repository-workspace mode: clone the whole repository, open the repository root, keep AGENTS.md plus the canonical skills/ tree and shared references together, then progressively load only the Skill needed for the task.
A runtime manifest or cloned directory is not proof that Skill discovery succeeded. Verify a fresh setup with the bootstrap/smoke-test prompts in docs/GETTING-STARTED.md. If automatic discovery fails, explicitly load the matching skills/<name>/SKILL.md; for cross-domain work use skills/amazon-ads-optimizer/SKILL.md.
Skill discovery
- Treat every
skills/*/SKILL.mdas an independently invocable skill. - Read only the matching skill first; load shared files from
references/,playbooks/andschemas/only when needed. - For a custom host without native Skill discovery, use
scripts/resolve_skill_context.py: expose only the metadata catalog at discovery time, preload exactly one selectedSKILL.md, and treat returned resources as on-demand pointers. Do not preload all Skill bodies,docs/research/, orevals/merely because they exist. - For cross-domain requests, use
skills/amazon-ads-optimizer/SKILL.mdas the orchestrator. - For recurring weekly/Monday account reviews, start with
playbooks/weekly-review.md, then load only specialist Skills required by material findings. - When the task asks whether a previous optimization worked, route to
skills/post-change-review/SKILL.mdbefore proposing another edit on the same entity. - When prior actions may overlap a new decision, load
references/optimization-memory.mdand retrieve only the bounded relevant entity history. - When compared windows come from different APIs, MCPs, exports, warehouses or semantic/report versions, load
references/data-lineage.mdbefore making high-confidence trend or causal claims. - When a decision depends on what the active MCP/connector can actually expose for the required profile, report, metric, dimension, history, pagination, freshness or semantic identity, load
references/connector-capability.md; treatPartial,UnsupportedandUnknownas acquisition-path evidence states, never as metric zero. - Resolve decision-required connector capability IDs from
references/connector-capability-catalog.jsonviascripts/resolve_skill_capabilities.py; do not invent new capability IDs ad hoc. When a causal review uses a registereddecision_surface, pass that surface to the resolver so its versioned material-control set is attached toentity-state-readback; generic readback support must not substitute for exactcontrol_types_exposedcoverage. When a machine-readable connector capability snapshot is available, feed the resolved required IDs and merged data requirements intoscripts/evaluate_connector_capability_gate.pybefore metric interpretation;Degraded/Blockedmust cap the dependent decision to conservative classes instead of preserving high confidence. - When a bid, budget, placement or other monetary-control Skill must turn a direction/raw estimate into a concrete magnitude, load
references/action-sizing.md; do not invent a repository-global change percentage or damping constant. - When a recommendation depends on a time-varying platform capability, exact platform limit, control eligibility, or console/API availability, load
references/platform-capability-lineage.mdand reconcile source date + capability scope before using the fact as action-safe. - Historical replay fixtures under
evals/are test inputs, not reusable operating instructions; do not load them during ordinary account analysis unless explicitly running an eval. - Do not duplicate business logic into this file.
Default operating mode
Default to Suggest mode. Analysis may propose actions, but no skill or playbook in this repository directly modifies a live Amazon Ads account.
Modes:
Read-only— inspect and explain data.Suggest— generate recommended actions without applying them. Default.Shadow— simulate actions for review/backtesting.Execute— only after explicit authorization; hand validated actions to an external connector/executor.
Safety rules
- Never invent missing advertising data, source lineage, optimization history, or action-sizing precision.
- Do not recommend aggressive bid/budget changes when sample size is insufficient.
- A raw economic or directional estimate is not automatically an action-safe final magnitude; require an explicit account/caller policy, calibrated response, defensible marginal headroom, experiment design, or another auditable sizing basis for a concrete proposed value.
- Check marketplace, profile/account scope, currency, timezone, attribution window, date range, promotion context, data freshness and source comparability before high-confidence recommendations.
- Treat
extracted_atandavailable_throughas different concepts; a freshly fetched downstream table may still be incomplete for recent event dates. - Treat a time-varying platform capability as lineage-bearing evidence: if current official/API/account evidence materially conflicts or scope/supersession is unresolved, fail closed on the exact numeric rule instead of silently choosing a convenient percentage, limit or formula.
- Treat Prime Day, Best Deal, Lightning Deal, Coupon, Prime-exclusive promotions, stockouts, listing suppression, major price changes, and parent/variation-family retail changes as potential confounders.
- Before reversing or stacking another action on the same entity, check whether a recent action is still inside its validation window when history is available.
- Distinguish
proposed,applied,readback confirmed, andworked; none of these imply the next stage automatically. - If history is unavailable or partial, expose that limitation instead of treating the ledger as complete.
- Prefer
Hold,Experiment Only, orManual Reviewwhen application is unknown/drifted, evaluation is still pending, source comparability is unresolved, action magnitude lacks a defensible sizing basis, or a new action would contaminate an active experiment unless a safety guardrail triggered. - Weekly reviews must include a hold list; do not force every material entity into an action.
- Do not place credentials, refresh tokens, client secrets, profile IDs, account IDs, or customer secrets in generated files or logs.
- Any
Executeplan must include evidence, confidence, guardrails, validation window, and rollback criteria. - A
Rollback Candidateis a proposal only; this repository does not perform the rollback itself.
Shared references, playbooks and evals
- Weekly operating cadence:
playbooks/weekly-review.md - Metrics:
references/amazon-ads-metrics.md - Optimization framework:
references/optimization-framework.md - Decision boundaries:
references/decision-boundaries.md - Contextual action sizing:
references/action-sizing.md - Platform capability version/scope lineage:
references/platform-capability-lineage.md - Benchmark policy:
references/benchmark-policy.md - Optimization memory:
references/optimization-memory.md - Canonical data model:
references/data-schema.md - Source lineage and cross-source comparability:
references/data-lineage.md - Active MCP/connector capability evidence:
references/connector-capability.mdandschemas/connector-capability-snapshot.json - Canonical connector capability IDs / Skill decision profiles:
references/connector-capability-catalog.json - Action proposal schema:
schemas/optimization-action.json - Action/readback/evaluation event schema:
schemas/optimization-event.json - Derived entity-history schema:
schemas/entity-history.json - Experiment plan schema:
schemas/experiment-plan.json - Historical replay framework:
evals/README.md - Eval-case schema:
schemas/eval-case.json - Paired Skill-effectiveness result schema:
schemas/skill-effectiveness-benchmark.json - Decision-time measurement-state projector:
scripts/project_measurement_history.py - Paired with/without-Skill benchmark summarizer:
scripts/summarize_skill_effectiveness.py
Evaluation rules
When adding or changing decision logic, prefer adding a focused synthetic replay fixture for material failure modes.
- Keep deterministic contract checks separate from capability evals.
- Every fixture must remain valid under
scripts/validate_evals.py: valid JSON, allowed contract fields/enums, collision-free ID, filename/ID alignment, and an existing repository-local entrypoint. - Score decision behavior rather than exact prose.
- Use
met,not_met, orinsufficient_evidence; do not coerce missing evidence into pass/fail. - A fixture may allow multiple conservative outcomes, but must list forbidden unsafe behaviors explicitly.
- Do not weaken a fixture merely because a current model fails it; change the Skill only when the fixture represents the intended behavior.
- Prefer synthetic identifiers and values. Never commit client/account secrets or proprietary exports as fixtures.
- Eval fixtures must remain closed-world by default and must not trigger live Amazon Ads writes.
- For important with-Skill / without-Skill comparisons, keep pair identity and harness/model/tool/evidence configuration aligned.
scripts/summarize_skill_effectiveness.pyfails closed on missing/duplicate pair variants and reports measured deltas only; it does not claim statistical significance.
Read-only repository utilities
The repository includes deterministic helpers for derived memory and evaluation summaries. These are not Amazon Ads executors.
scripts/project_measurement_history.pyreads an optimization-event slice from stdin and projects only the newest traceableevidence_snapshotintolatest_measurement_state. It fails closed on requested scope mismatch, preservesunknown,retired_or_deletedandNot Comparable, and never converts unavailable history to zero.scripts/resolve_skill_capabilities.pyresolves canonical required/optional connector capability IDs plus profile-owneddata_requirementsfor a registered Skill decision profile and can merge explicit task-level constraints such as an exacthistory_window. Task requirements may only strengthen profile policy, must target a required capability, and retain profile/task provenance. Pass the merged requirements unchanged to the connector capability gate. It never calls a connector or grants write authority.scripts/evaluate_connector_capability_gate.pyreads a connector capability snapshot plus required capability IDs from stdin and returnsPass,Degraded, orBlocked. It is read-only, never calls a connector, and enforcesmissing_evidence_policy = never_zerofor Partial/Unsupported/Unknown/absent capability evidence. Optional per-capabilitydata_requirementsadditionally gate exact reporting generation and history availability: retired/mismatched generations or unavailable/retired required history block; partial/unknown required history degrades rather than becoming zero. A task-levelhistory_windowchecks exact start/end dates against the matching observed grain; no cross-grain coverage inference is allowed.scripts/evaluate_metric_aggregation_gate.pyis a read-only policy gate for direct summation. It permits a sum only when metric aggregation semantics are explicitly additive and source rows are explicitly disjoint; unknown semantics never default to additive, and non-additive de-duplicated or ratio/derived metrics are blocked from direct summation.scripts/evaluate_skill_effectiveness_comparability.pycompares two benchmark measurement identities before longitudinal interpretation. Fixture/evaluator/rubric/model/runtime/config/tool/evidence mismatches returnNot Comparable; missing identity returnsUnknown, never implicit comparability. It is read-only.scripts/summarize_skill_effectiveness.pyreads benchmark trial records that conform toschemas/skill-effectiveness-benchmark.json. It validates pairedwith_skill/without_skilltrials and reports counts/rates/deltas only from pairs where both arms were actually measured. Scorer/harness errors and insufficient-evidence arms exclude the whole pair rather than being imputed as zero quality. It does not call a model, invoke a live agent runtime, or mutate an advertiser account.
Actual model/harness execution remains an external evaluation-runner concern. Raw transcripts and provider credentials should stay outside normal Skill loading paths and outside committed fixtures unless explicitly synthetic and safe.
Agent Skills frontmatter compatibility
SKILL.md frontmatter must stay compatible with the Agent Skills specification and strict reference validators.
Allowed top-level fields are:
name— required and must match the parent directory name;description— required and should state both capability and trigger/use case;license— optional; repository-owned Skills should normally useMIT;compatibility— optional, only when environment requirements are material;metadata— optional string-to-string map for repository-specific metadata;allowed-tools— optional/experimental and should not be used to broaden execution authority.
Do not add custom top-level fields such as display_name, display_name_en, description_zh, description_en, version or author. Put repository-specific values under metadata instead.
Keep name within the Agent Skills naming constraints and description within the specification limit. Metadata compatibility is part of multi-agent portability; a Skill that works in one permissive runtime but fails strict validation is not considered portable.
Versioning and releases
VERSION is the canonical repository release version. Keep the README version label and all runtime manifests (.codex-plugin/plugin.json, .claude-plugin/plugin.json, and .workbuddy-plugin/plugin.json) synchronized with it, and add the release entry to CHANGELOG.md.
Use semantic versioning for release-worthy changes:
- PATCH — backward-compatible bug fixes, documentation corrections, research/source maintenance, or internal changes that do not materially expand a public decision/runtime contract.
- MINOR — backward-compatible new capability, materially stronger decision/safety/data/runtime behavior, new optional schema surface, or another substantial improvement users should be able to identify as a new release.
- MAJOR — breaking changes to public schemas, CLI/output contracts, canonical Skill/capability identity, runtime integration contracts, or changes that require downstream migration.
A material MINOR/MAJOR change must not merge while the repository continues to advertise the previous release indefinitely. Classify version impact in the pull request; when a bump is required, update VERSION, README, all runtime manifests, and CHANGELOG.md in the same PR. Do not add repository release versions to individual SKILL.md top-level frontmatter.
Deterministic repository validation
After changing a Skill, its supporting references, shared references, schemas, playbooks, eval fixtures, validators, projector/summarizer utilities, or runtime manifests, run:
python -m unittest discover -s tests -v
python scripts/validate_skills.py .
python scripts/validate_evals.py .
The validators are intentionally dependency-free and check repository invariants that should fail before a runtime or eval harness discovers them interactively. Unit tests additionally cover measurement-state projection, realized-ad projection, experiment-plan semantics, paired effectiveness aggregation and release/runtime compatibility.
Skill/package validation checks:
- every direct child of
skills/containsSKILL.md; - required
name/descriptionfrontmatter exists; - top-level frontmatter stays within the Agent Skills allowed field set;
- Skill
namematches its folder name; - nested
SKILL.mdfiles are rejected; - relative Markdown references from
SKILL.mdstay inside the repository and resolve to existing paths.
Eval-contract validation checks:
- every
evals/fixtures/*.jsonfile parses as JSON; - required/allowed fixture and expected-result fields are respected;
- mode and acceptable-decision values stay within the closed-world schema contract;
- fixture ID matches filename and does not collide;
- entrypoint stays inside the repository and exists;
- rubric/required-observation contracts do not positively require live mutation.
.github/workflows/validate-skills.yml runs unit tests plus both validators for relevant pushes and pull requests when GitHub Actions is enabled. Do not treat a local or CI pass as proof of Amazon Ads decision quality; deterministic structure/contract checks complement, rather than replace, capability replay behavior checks.
Contribution rules
New skills must:
- use kebab-case directory names;
- include Agent Skills-compatible YAML frontmatter with at least
nameanddescription; - keep custom metadata under the
metadatamap rather than inventing top-level fields; - keep the core
SKILL.mdconcise and progressively load shared references; - state required inputs, workflow, output contract, safety checks, and stop conditions;
- output proposals rather than performing live account mutation.
Add a playbook instead of a new Skill when the new material is primarily a recurring operating rhythm that composes existing Skills rather than a distinct decision capability.
Shared memory/history features should prefer append-first events plus derived compact summaries rather than mutable prose logs.
When a change materially affects source comparability, negatives, bid/budget reversals or sizing, retail causality, Mixed-ASIN safety, growth qualification, experiments, readback, or rollback decisions, add or update a replay fixture when practical.
