Imported from DouwMarx/evaluating-evaluations (
scenario-ideation/SKILL.md). Install upstream withnpx skills add DouwMarx/evaluating-evaluations --skill scenario-ideation. Copyright stays with the author.
Scenario Ideation — Phase 1
Guide a user through Phase 1 of EquiStamp's evaluation development process: generating grounded, specific scenarios that describe how real actors would use or misuse AI systems. The output is a scenario ideation document that feeds directly into Phase 2 (Risk Modeling).
Reference workbook: scenario-ideation.md in the project root. This skill operationalizes that workbook as an interactive session.
Target time: 1-2 hours to produce a complete scenario ideation document with Claude's help.
When to Use
- Starting a new evaluation project
- User says "I'm worried about [risk domain]" or "what scenarios should we evaluate for [topic]?"
- Beginning the eval development pipeline
- Brainstorming AI safety risks for a specific domain
- Preparing scenario documentation for EU AI Act GPAI model risk assessments
- A client needs a structured threat analysis before building an evaluation suite
Important Concepts
Why Scenarios, Not Risk Categories
Do not start from a taxonomy like "biosecurity" or "cybersecurity." Taxonomies are abstractions that hide the details that matter for evaluation design. "Biosecurity" encompasses everything from a teenager looking up dangerous recipes to a state bioweapons program. These require fundamentally different evaluations because the actors, capabilities, and pathways are different.
Scenarios force specificity. When you write "a graduate student with access to a university wet lab uses an AI assistant to troubleshoot a synthesis protocol," you have immediately constrained the evaluation design space. You know the actor's skill level, their access to equipment, the AI's role, and the key question for evaluation.
Tiered Evaluation Approach
Scenarios should note which tier of evaluation they likely require:
- Tier 1 — Automated benchmark: Can be evaluated with a static dataset and automated scoring. Best for capability measurements.
- Tier 2 — Human-in-the-loop evaluation: Requires human judges, red-teamers, or domain experts to assess outputs. Best for nuanced harm assessments.
- Tier 3 — Agentic/environment-based evaluation: Requires interactive environments, multi-step tasks, or tool-use scaffolding. Best for realistic threat simulations.
EU AI Act Context
EquiStamp's scenario work often feeds into GPAI model risk assessments under the EU AI Act. When relevant, scenarios should consider:
- Systemic risks as defined by the Act (risks to public health, safety, security, fundamental rights, or the functioning of democratic processes)
- Whether the scenario would trigger notification obligations for GPAI model providers
- How the scenario maps to the risk categories in the Act's annexes
CBRN Scenario Taxonomy
When the risk domain is CBRN, scenarios require a structured taxonomy beyond generic tags. Each CBRN scenario must include a Scenario Overview with qualitative severity, feasibility, and AI uplift ratings, plus a Scenario Taxonomy table covering 10 dimensions. This taxonomy ensures scenarios are specific enough to drive evaluation design and comparable across the portfolio.
Scenario Overview fields:
- Severity — Qualitative measure of the overall severity of the CBRN event (e.g., Low / Moderate / High / Catastrophic), considering casualties, disruption scope, and long-term consequences.
- Feasibility — Qualitative measure of overall feasibility/likelihood (e.g., Low / Moderate / High), considering actor capability, material access, and technical barriers.
- AI Uplift — Qualitative measure of AI's contribution to enabling or accelerating the scenario (e.g., Minimal / Moderate / Significant / Transformative).
Taxonomy dimensions (each must include a value and a discussion justifying the selection):
| Dimension | Description | Example Values |
|---|---|---|
| Number of Experts | Number of experts contributing to the threat actor | Lone Wolf Expert, Small Team (2-5), Organization (6+) |
| CBRN Project Resources | USD liquid assets dedicated to the CBRN project | < $10k, $10k-$100k, $100k-$1M, > $1M |
| Preexisting Infrastructure Access | Does the threat actor have "insider" access to relevant infrastructure? | Access, No Access |
| CBRN Agent Class | Category of CBRN agent used in the scenario | Non-Transmissible Bio, Transmissible Bio, Chemical, Radiological, Nuclear |
| Quantity / Lethality | Quantity of CBRN materials in lethal/infectious dose units | < 100 Lethal Doses, 100-1,000, 1,000-10,000, > 10,000 |
| Materials Provenance | How the CBRN agent/precursor comes into the hands of the threat actor | Self Manufacture, Purchase, Theft, State Supply, Found/Scavenged |
| Post-processing | Once materials are acquired, is substantive additional development needed? | Non-substantial Post-Processing, Substantial Post-Processing |
| Novelty | Is the CBRN agent developed/deployed in a novel manner? | Known Agent / Known Procedure, Known Agent / Novel Procedure, Novel Agent / Known Procedure, Novel Agent / Novel Procedure |
| Harm Type | The primary measure of harm for the scenario | Disruption (inc. Economic, Agricultural), Casualties (< 100), Mass Casualties (100+), Infrastructure Damage |
| Distribution Scope | How the CBRN agent is distributed | Self-propagating (infectious), Targeted (specific location/population), Wide Area (dispersal), Point Source |
When working in CBRN domains, use the CBRN-specific scenario template (shown below in Step 4) instead of the generic tag table. This taxonomy replaces — not supplements — the generic tags for CBRN scenarios.
Phase 1: Gather Context
Start by asking the user these questions. Wait for answers before proceeding.
I'll help you create a scenario ideation document for EquiStamp's eval pipeline. A few questions to get started:
1. **What risk domain are you focused on?**
- CBRN (chemical, biological, radiological, nuclear)
- Cyber offense
- Loss of control / autonomy
- Harmful manipulation / persuasion
- Sociotechnical risks (labor displacement, epistemic erosion, etc.)
- Other: [describe]
2. **Any specific concerns or recent developments that prompted this?**
(e.g., a new model release, a client request, a published incident, a capability announcement)
3. **Who is the audience for these scenarios?**
- Internal prioritization (deciding what to evaluate next)
- Client deliverable (part of a project)
- Regulatory compliance (EU AI Act GPAI assessment)
- Research publication
4. **Do you have existing scenarios or risk assessments to build on?**
- Link or paste any prior work
- Or "starting from scratch"
Do not proceed until the user has answered at least questions 1 and 3. Questions 2 and 4 are helpful but not blocking.
Phase 2: Research and Draft
Walk through the workbook checklist step by step. For each step, produce a concrete output before moving on.
Step 1: Survey the Capability Landscape
Help the user identify 5-10 capability shifts from the past 90 days relevant to their risk domain.
Suggest using web search to find:
- Recent model releases and capability announcements
- Published benchmarks and evaluation results
- Research papers on relevant capabilities
- Reported incidents or demonstrations
Output format:
### Capability Landscape (as of [DATE])
| # | Capability Shift | Source / Date | Relevance to [Domain] |
|---|---|---|---|
| 1 | [e.g., "Model X achieves Y% on protein structure prediction benchmark"] | [source, date] | [why it matters] |
| 2 | ... | ... | ... |
Guidance: 90 days is the window because model capabilities shift fast enough that older surveys miss current trajectories, but short enough to keep the review focused. If the user's domain has had fewer recent developments, extend to 6 months but note the longer window.
Step 2: Build the Actor Table
For each capability shift, identify 2-3 specific actor types who would use or misuse it.
Key requirement: Actors must be specific, not abstract. Not "a bad actor" but "a biology undergraduate with access to a university wet lab and standard coursework knowledge." Name the role, the motivation, and the access level.
Output format:
### Actor Table
| Actor | Motivation | Current Capability (without AI) | Expected Capability (with AI) |
|---|---|---|---|
| [e.g., "Graduate student in synthetic biology, university wet lab access"] | [e.g., "Academic research, potentially dual-use"] | [e.g., "Can follow published protocols, struggles with novel synthesis troubleshooting"] | [e.g., "AI assistant provides real-time troubleshooting equivalent to years of bench experience"] |
| ... | ... | ... | ... |
Step 3: Forward Projections
For each actor-capability pair, write one sentence answering: "What can this actor do in 12-18 months that they cannot do today?"
Ground projections in the trajectory of capability gains observed in Step 1. The 12-18 month horizon is a practical choice: short enough that projections are grounded in observable trends, long enough that evaluations built now will still be relevant when deployed.
This follows EquiStamp's foundational principle: "weather forecasts, not weather reports." Evaluations built for today's capability ceiling will saturate when the next model generation arrives.
Output format:
### Forward Projections (12-18 month horizon)
| Actor | Capability Trajectory | Projection |
|---|---|---|
| [actor from table] | [observed trend] | [what they can do in 12-18 months] |
| ... | ... | ... |
Step 4: Draft Scenarios
Write 3-5 scenarios. Each scenario must include all four required elements: actor, AI capability, outcome, and why AI changes things.
Use this template for each scenario (non-CBRN domains):
#### Scenario [N]: [Title]
**Actor:** [Who — specific role, motivation, access level]
**AI Capability:** [What specific AI capability enables this]
**What happens:** [1-3 paragraphs describing the concrete event. Be specific about the sequence of actions, the tools involved, and the outcome. This is a narrative, not an abstract.]
**Why AI changes things:** [What was not possible or much harder before this capability existed. Be precise about the uplift.]
| Tag | Value |
|---|---|
| Primary capability | [e.g., code generation, protein structure prediction, voice synthesis] |
| Harm type | [Physical / Financial / Informational / Systemic] |
| Actor type | [e.g., state actor, criminal organization, lone individual, insider threat] |
| Novelty | [Acceleration of existing threat / Entirely new threat] |
| Likely evaluation tier | [Tier 1: Automated / Tier 2: Human-in-the-loop / Tier 3: Agentic] |
| EU AI Act relevance | [Systemic risk category if applicable, or "N/A"] |
For CBRN domains, use this expanded template instead:
#### Scenario [N]: [Title]
##### Scenario Overview
**Actor:** [Who — specific role, motivation, access level]
**AI Capability:** [What specific AI capability enables this]
**What happens:** [1-3 paragraphs. Include: threat actor organisation/profile, resources and expertise, CBRN agent used, procedure for procurement of CBRN materials, delivery mechanism, and threat actor intentions (or accident details). Be specific about the sequence of actions, tools involved, and outcome.]
**Why AI changes things:** [What was not possible or much harder before this capability existed. Be precise about the uplift.]
**Severity:** [Low / Moderate / High / Catastrophic] — [1-2 sentence justification]
**Feasibility:** [Low / Moderate / High] — [1-2 sentence justification]
**AI Uplift:** [Minimal / Moderate / Significant / Transformative] — [1-2 sentence justification]
##### Scenario Taxonomy
| Dimension | Value | Discussion |
|---|---|---|
| Number of Experts | [Lone Wolf Expert / Small Team (2-5) / Organization (6+)] | [Why this value fits the scenario] |
| CBRN Project Resources | [< $10k / $10k-$100k / $100k-$1M / > $1M] | [Why this budget level is appropriate] |
| Preexisting Infrastructure Access | [Access / No Access] | [What infrastructure the actor does or does not have] |
| CBRN Agent Class | [Non-Transmissible Bio / Transmissible Bio / Chemical / Radiological / Nuclear] | [Why this agent class] |
| Quantity / Lethality | [< 100 Lethal Doses / 100-1,000 / 1,000-10,000 / > 10,000] | [Justify the scale; specify for transmissible bio / nuclear / rad] |
| Materials Provenance | [Self Manufacture / Purchase / Theft / State Supply / Found/Scavenged] | [How materials are obtained and why this is the likely path] |
| Post-processing | [Non-substantial Post-Processing / Substantial Post-Processing] | [What additional development is needed after acquisition] |
| Novelty | [Known Agent / Known Procedure / Known Agent / Novel Procedure / Novel Agent / Known Procedure / Novel Agent / Novel Procedure] | [Nature of novelty, if any] |
| Harm Type | [Disruption (inc. Economic, Agricultural) / Casualties (< 100) / Mass Casualties (100+) / Infrastructure Damage] | [Primary measure of harm] |
| Distribution Scope | [Self-propagating (infectious) / Targeted / Wide Area / Point Source] | [How the agent reaches the affected population] |
| Tag | Value |
|---|---|
| Primary capability | [specific AI capability enabling the scenario] |
| Likely evaluation tier | [Tier 1: Automated / Tier 2: Human-in-the-loop / Tier 3: Agentic] |
| EU AI Act relevance | [Systemic risk category if applicable, or "N/A"] |
Requirements:
- At least 3 scenarios, ideally 5
- At least one scenario must describe a beneficial outcome (same structure, but the outcome is positive — this ensures evaluations can distinguish beneficial from harmful uses of the same capability)
- Scenarios should span different actor types where possible
- Each scenario should be falsifiable: you can investigate whether the actor, capability, and outcome are plausible
- For CBRN scenarios: every taxonomy dimension must have both a value and a discussion. Blank discussions are not acceptable — the justification is what makes the taxonomy useful for evaluation design
Concrete Example: Completed CBRN Scenario
Here is what a finished CBRN scenario looks like (biosecurity domain), using the expanded taxonomy:
#### Scenario 1: Graduate Student Synthesis Troubleshooting
##### Scenario Overview
**Actor:** A second-year graduate student in synthetic biology at a mid-tier US university. They have completed standard coursework in molecular biology and have access to a university wet lab with standard equipment (PCR machines, centrifuges, biosafety level 2 cabinet). They are working on a legitimate research project involving gene synthesis but have encountered a troubleshooting problem that would normally require consulting a senior postdoc or PI with 5+ years of bench experience.
**AI Capability:** Current frontier language models can provide detailed, step-by-step troubleshooting advice for synthesis protocols, including suggesting modifications to reaction conditions, identifying likely failure modes from described symptoms, and recommending alternative approaches. This capability has improved notably in the past 6 months as models have been trained on more specialized scientific literature and can now handle multi-step reasoning about experimental procedures.
**What happens:** The student is attempting to synthesize a protein that has legitimate research applications but is also listed on a dual-use research concern watchlist. They encounter a folding problem that causes their synthesis to fail repeatedly. Rather than waiting two weeks for their advisor's office hours or spending a month learning the troubleshooting process through trial and error, they describe the problem to an AI assistant. The assistant identifies the likely cause (a codon optimization issue specific to their expression system), suggests three alternative approaches ranked by likelihood of success, and provides a modified protocol. The student successfully synthesizes the protein on the next attempt.
**Why AI changes things:** Without the AI, this troubleshooting step would have required either (a) extensive literature search skills the student has not yet developed, (b) access to an experienced mentor who happens to have worked with this specific expression system, or (c) months of trial-and-error experimentation. The AI compresses what was effectively a skill gate — a step that required years of accumulated bench experience — into a conversational interaction. The student's capability jumps from "can follow published protocols" to "can troubleshoot novel synthesis problems," which is the key capability transition for dual-use concern.
**Severity:** Moderate — Successful synthesis of a dual-use protein by a minimally supervised student does not itself cause harm, but removes a key bottleneck in a pathway that could lead to dangerous biological agent production if intent were malicious.
**Feasibility:** High — All components (university lab access, AI troubleshooting capability, dual-use protein targets) currently exist and are readily accessible.
**AI Uplift:** Significant — The AI compresses years of tacit bench experience into a single conversational interaction, eliminating the primary skill gate that previously limited who could perform this work.
##### Scenario Taxonomy
| Dimension | Value | Discussion |
|---|---|---|
| Number of Experts | Lone Wolf Expert | Single graduate student operating independently; while they have academic training, they are acting without direct supervision for the troubleshooting step. |
| CBRN Project Resources | < $10k | The student uses existing university infrastructure and consumables already budgeted for their research project. No additional procurement needed beyond standard lab supplies. |
| Preexisting Infrastructure Access | Access | The student has legitimate, institutional access to a BSL-2 wet lab with standard molecular biology equipment through their university enrollment. |
| CBRN Agent Class | Non-Transmissible Bio | The target protein has dual-use concern but is not a transmissible pathogen; it is a discrete biological product that does not self-replicate or spread between hosts. |
| Quantity / Lethality | < 100 Lethal Doses | Bench-scale protein synthesis at a university lab produces microgram-to-milligram quantities, well below mass-production levels. Lethality depends on the specific agent but is constrained by production scale. |
| Materials Provenance | Self Manufacture | The student synthesizes the protein de novo in their lab using standard gene synthesis and expression techniques, not procuring it from an external source. |
| Post-processing | Non-substantial Post-Processing | Once the protein is expressed and folded correctly (the step the AI assists with), minimal additional processing is needed — standard purification steps the student already knows. |
| Novelty | Known Agent / Known Procedure | The protein target and the synthesis/expression methodology are both well-documented in the literature. The novelty lies in AI-assisted troubleshooting, not in the agent or procedure itself. |
| Harm Type | Disruption (inc. Economic, Agricultural) | The primary harm is erosion of the skill-gate barrier that limits who can produce dual-use biological materials, rather than a specific casualty event. The scenario represents a systemic shift in accessibility. |
| Distribution Scope | Point Source | The synthesized protein remains in the lab; there is no distribution mechanism in this scenario. The concern is about capability acquisition, not deployment. |
| Tag | Value |
|---|---|
| Primary capability | Scientific reasoning and protocol troubleshooting |
| Likely evaluation tier | Tier 2: Human-in-the-loop (domain expert needed to assess uplift) |
| EU AI Act relevance | Systemic risk: public health and safety |
Phase 3: Validate and Finalize
Murphyjitsu Stress Test
Apply the CFAR murphyjitsu technique to each scenario. For every scenario, ask:
"Imagine it is 12 months from now. This scenario has NOT happened, despite the capabilities being available. What prevented it? Would that prevention be surprising?"
If the prevention is not surprising — if there are obvious barriers you did not account for — revise the scenario to address those barriers, or note them as mitigating factors.
Then ask the inverse:
"Imagine this scenario HAS happened and it went worse than described. What made it worse? Would that escalation be surprising?"
If the escalation is not surprising, consider whether the scenario should be revised to reflect the more severe version.
Reality Check
For each scenario, identify one of the following:
- Expert consultation: Name a domain expert (or type of expert) who could confirm plausibility. A 10-minute conversation is sufficient.
- Published source: Cite a published incident, capability demonstration, or expert forecast that supports the scenario's plausibility.
- Empirical test: Describe a quick experiment that could validate the capability claim (e.g., "test whether model X can actually troubleshoot this type of synthesis problem").
Output format:
### Validation Notes
| Scenario | Validation Type | Source / Note |
|---|---|---|
| Scenario 1 | [Expert / Citation / Empirical] | [Details] |
| ... | ... | ... |
Peer Review Prompts
Generate two review questions for a team member to answer for each scenario:
- "Is the actor specific enough that I could simulate their behavior in an evaluation?"
- "Is the AI capability claim grounded in something that exists or is on a clear trajectory?"
Scenarios that fail either question should be revised before finalizing.
Phase 4: Hand Off
Save the Document
Help the user save the completed scenario ideation document. Suggest a filename like scenario-ideation-[domain]-[date].md and a location in the project repository.
Next Steps
After the document is complete:
-
Create GitHub issues for Phase 2 (Risk Modeling) work on each scenario or scenario cluster. Use the
/create-issueskill with task type "Evaluation Development." -
Reference for Phase 2: "Use the
/risk-modelingskill to begin Phase 2 on these scenarios. Phase 2 will identify the crux of each scenario — the specific point where AI is actually changing what is possible — and design measurement approaches." -
Check completion criteria — confirm all of the following are true before handing off:
- 3-5 written scenarios, each with actor, capability, outcome, and why-AI-changes-things
- Each scenario tagged with capability, harm type, actor type, novelty, evaluation tier
- CBRN domains: Each scenario includes Severity/Feasibility/AI Uplift ratings and the full 10-dimension taxonomy with justified values
- Each scenario has a plausibility check (expert, citation, or empirical)
- At least one beneficial-outcome scenario is included
- Peer review questions are generated (actual peer review may happen async)
- Document is stored where the Phase 2 team can access it
Full Document Template
When generating the final document, use this template. Replace all [bracketed text] with project-specific content.
# Scenario Ideation: [Risk Domain]
**EquiStamp Evaluation Development — Phase 1**
**Author:** [Name]
**Date:** [YYYY-MM-DD]
**Audience:** [Internal prioritization / Client deliverable / Regulatory compliance / Research]
**Status:** [Draft / Peer Review / Final]
---
## Context
[1-2 paragraphs: Why are we looking at this risk domain now? What prompted this work? What is the evaluation goal?]
---
## Capability Landscape (as of [DATE])
*Relevant capability shifts from the past 90 days.*
| # | Capability Shift | Source / Date | Relevance to [Domain] |
|---|---|---|---|
| 1 | [description] | [source, date] | [relevance] |
| 2 | [description] | [source, date] | [relevance] |
| 3 | [description] | [source, date] | [relevance] |
| ... | ... | ... | ... |
---
## Actor Table
| Actor | Motivation | Current Capability (without AI) | Expected Capability (with AI) |
|---|---|---|---|
| [specific actor] | [motivation] | [current] | [with AI] |
| ... | ... | ... | ... |
---
## Forward Projections (12-18 month horizon)
| Actor | Capability Trajectory | Projection |
|---|---|---|
| [actor] | [trend] | [12-18 month projection] |
| ... | ... | ... |
---
## Scenarios
<!-- For non-CBRN domains, use this structure: -->
### Scenario 1: [Title]
**Actor:** [Who — specific role, motivation, access level]
**AI Capability:** [What specific AI capability enables this]
**What happens:** [1-3 paragraphs]
**Why AI changes things:** [Specific uplift description]
| Tag | Value |
|---|---|
| Primary capability | [capability] |
| Harm type | [Physical / Financial / Informational / Systemic] |
| Actor type | [actor type] |
| Novelty | [Acceleration / Entirely new] |
| Likely evaluation tier | [Tier 1 / Tier 2 / Tier 3] |
| EU AI Act relevance | [category or N/A] |
<!-- For CBRN domains, use this structure instead: -->
### Scenario 1: [Title]
#### Scenario Overview
**Actor:** [Who — specific role, motivation, access level, organisation]
**AI Capability:** [What specific AI capability enables this]
**What happens:** [1-3 paragraphs. Include: threat actor organisation/profile, resources and expertise, CBRN agent used, procurement procedure, delivery mechanism, and threat actor intentions or accident details.]
**Why AI changes things:** [Specific uplift description]
**Severity:** [Low / Moderate / High / Catastrophic] — [justification]
**Feasibility:** [Low / Moderate / High] — [justification]
**AI Uplift:** [Minimal / Moderate / Significant / Transformative] — [justification]
#### Scenario Taxonomy
| Dimension | Value | Discussion |
|---|---|---|
| Number of Experts | [value] | [justification] |
| CBRN Project Resources | [value] | [justification] |
| Preexisting Infrastructure Access | [value] | [justification] |
| CBRN Agent Class | [value] | [justification] |
| Quantity / Lethality | [value] | [justification] |
| Materials Provenance | [value] | [justification] |
| Post-processing | [value] | [justification] |
| Novelty | [value] | [justification] |
| Harm Type | [value] | [justification] |
| Distribution Scope | [value] | [justification] |
| Tag | Value |
|---|---|
| Primary capability | [capability] |
| Likely evaluation tier | [Tier 1 / Tier 2 / Tier 3] |
| EU AI Act relevance | [category or N/A] |
<!-- End of CBRN-specific structure. Repeat for each scenario. -->
### Scenario 2: [Title]
[Same structure as Scenario 1]
### Scenario 3: [Title]
[Same structure as Scenario 1]
### Scenario 4: [Title] (Beneficial Outcome)
**Actor:** [Who]
**AI Capability:** [What]
**What happens:** [Positive outcome narrative]
**Why AI changes things:** [Uplift description — same capability, beneficial application]
[Include Severity/Feasibility/AI Uplift ratings and full taxonomy table for CBRN domains, or generic tag table for non-CBRN domains.]
### Scenario 5: [Title]
[Same structure as Scenario 1]
---
## Validation Notes
| Scenario | Validation Type | Source / Note |
|---|---|---|
| Scenario 1 | [Expert / Citation / Empirical] | [details] |
| Scenario 2 | [Expert / Citation / Empirical] | [details] |
| ... | ... | ... |
### Murphyjitsu Notes
[For each scenario: What would prevent it? What would make it worse? Are those outcomes surprising?]
---
## Peer Review
**Reviewer:** [Name or TBD]
**Date:** [YYYY-MM-DD or TBD]
For each scenario, the reviewer should answer:
1. "Is the actor specific enough that I could simulate their behavior in an evaluation?"
2. "Is the AI capability claim grounded in something that exists or is on a clear trajectory?"
| Scenario | Actor Specificity | Capability Grounding | Revisions Needed |
|---|---|---|---|
| Scenario 1 | [Pass / Fail — notes] | [Pass / Fail — notes] | [what to revise] |
| ... | ... | ... | ... |
---
## Completion Checklist
- [ ] 3-5 written scenarios with all four elements (actor, capability, outcome, why-AI-changes-things)
- [ ] Each scenario tagged (capability, harm type, actor type, novelty, evaluation tier)
- [ ] **CBRN only:** Each scenario includes Severity, Feasibility, and AI Uplift qualitative ratings with justifications
- [ ] **CBRN only:** Each scenario includes the full 10-dimension taxonomy table with values AND discussion for every dimension
- [ ] Each scenario has a plausibility check (expert, citation, or empirical)
- [ ] At least one beneficial-outcome scenario included
- [ ] Peer review complete and revisions incorporated
- [ ] Document stored in project repository / shared workspace
---
## Next Steps
- [ ] Create GitHub issues for Phase 2 (Risk Modeling) for each scenario cluster
- [ ] Begin Phase 2 using the `/risk-modeling` skill
- [ ] [Any domain-specific follow-ups]
---
*This document was produced using EquiStamp's Scenario Ideation process (Phase 1 of the eval development pipeline). Reference workbook: `scenario-ideation.md`.*
Tips for the Facilitator
- Keep scenarios concrete. If you catch yourself writing "a malicious actor could potentially..." — stop. Name the actor. Describe the specific action. A scenario you cannot visualize is not specific enough.
- Do not over-refine. This phase should take 1-3 days per scenario batch. If you are spending longer, you are over-refining scenarios that should be rough drafts feeding into Phase 2.
- Use the beneficial scenario strategically. The beneficial scenario often shares the same underlying capability as a harmful one. The difference is the actor and the context. This forces more precise measurement in the evaluation.
- The 90-day window is a guideline, not a rule. For slow-moving domains (e.g., nuclear), extend to 6 months. For fast-moving domains (e.g., code generation), 60 days may be more appropriate.
- When in doubt, make the scenario more specific. You can always generalize later. You cannot evaluate an abstraction.