# Running Living Dungeon simulations locally

The local simulator is a reproducible scenario and pacing model for **The
Cinder Vault**. It is deliberately more detailed than a spreadsheet and much
smaller than a Magic rules engine.

It executes the Dungeon's chapters, Encounter library, face-up Reserve,
casting windows, battlefield and graveyard, cleanup and recycling, Room gates,
Sealed transitions, Focus rolls, Heat, Doom, combat timing, Heartbound, and
both boss phases. Each adventurer remains a disclosed statistical envelope for
development, attack, collective blocking, answers, and variance. The model
does not invent hands, mana, priority, target legality, or exact card-by-card
combat.

Run every command below from the repository root.

## Characterize opening turns

`simulate-openings.py` is the smaller companion model for hands, mana, and
coarse battlefield development through turns two to seven. It draws from the
exact selected 58/59-card libraries and their fingerprint-bound dated card
metadata. It applies one shared London mulligan policy, normal first-turn
draws, reviewed early land restrictions, simplified colored payment, starting
Treasure consumption, selected ramp and token effects, commander timing, d20
events, and each deck's first signature event.

Compare no starting Treasure with one and two using identical paired hands:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-openings.py `
  --runs 10000 --turns 5 --seed 1729 `
  --treasures 0,1,2
```

The first Treasure count is the paired baseline. Every later scenario receives
the same shuffled library, mulligan decision, opening hand, normal draws, and
named die-result streams. If an effect rolls extra dice, its kept result can
still differ as it should. This isolates the resource change instead of quietly
letting the faster scenario keep riskier hands.

Trace the same exact opening under every selected Treasure count:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-openings.py `
  --runs 1000 --turns 5 --treasures 0,2 `
  --trace-seed 1729
```

JSON and CSV output use the same `--format` and `--output` conventions as the
Dungeon pacing model. The JSON report fingerprints every selected list,
metadata snapshot, semantic annotation file, and the helper itself.

The helper is intentionally not a hand-level Magic engine. Its reviewed
casting policy prioritizes commanders, acceleration, signature-enabling
rolls, engines, and bodies. It does not select targets, model an opponent's
interaction, assign combat, calculate exact ready blockers or power, or infer
unreviewed card effects from Oracle-text patterns. Body counts characterize
table presence only. Trace a surprising seed before accepting an aggregate
explanation.

## Fast feedback

Use 1,000 runs while editing one layer. `latest` selects the current paper
baseline without rerunning historical loops.

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 1000 --seed 1729 --revision latest
```

The text report includes a 95% Wilson interval around the win rate, loss
causes, damage-source averages, maximum Reserve and permanent load, answer use,
gated-Vitality share, target status, and any paired comparison.

## Stable baseline

Use the preserved seed and 10,000 runs before recording a selected result:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 10000 --seed 1729 --revision latest `
  --format json --output .\reports\d20-baseline.json `
  --fail-on-target
```

`reports/` is ignored, so generated results remain local. The JSON report binds
the output to SHA-256 digests of the scenario and revision inputs and records
the base seed plus deterministic stride.

`--fail-on-target` exits with status 3 if any selected scenario misses one of
the configured first-playtest targets. Omit it for exploratory runs that are
expected to miss.

## Walk through exact seeds

Add one or more trace seeds to a normal run:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 1000 --revision latest `
  --trace-seed 2 --trace-seed 14
```

Text output gives a narrated expedition. JSON output includes the same
narration plus structured events containing round, Dungeon half, stage, Focus,
Party Life, Vitality, Doom, Reserve, battlefield, and event-specific details.
This is the preferred way to investigate a suspicious aggregate without
trying to infer a game from averages.

Every batch also names seeds for a median win, narrow win, median loss, and
deepest loss when those outcomes exist. Paired comparisons list example seeds
that flipped in each direction. Feed any suggested seed back through
`--trace-seed`, or ask the same batch to select and trace one directly:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 5000 --revision latest `
  --trace-representative median-win `
  --trace-representative deepest-loss
```

Available names are `median-win`, `narrow-win`, `median-loss`, and
`deepest-loss`.

## Compare tuning values with common seeds

`--sweep` evaluates every listed value against the same run seeds. This makes
the reported loss-to-win and win-to-loss flips much more useful than comparing
unrelated random batches.

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 5000 --seed 1729 --revision latest `
  --sweep dungeon_pressure_multiplier=1.15,1.22,1.29,1.36 `
  --format csv --output .\reports\d20-pressure-sweep.csv
```

Multiple sweeps form a Cartesian grid, capped at 64 scenarios:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 2000 --revision latest `
  --sweep party_attack_multiplier=0.95,1.0,1.05 `
  --sweep party_defense_multiplier=0.9,1.0,1.1
```

Useful revision parameters are:

- `party_life`
- `doom_limit`
- `party_attack_multiplier`
- `party_defense_multiplier`
- `answer_multiplier`
- `dungeon_pressure_multiplier`
- `setup_damage_per_stage`

Use `--set FIELD=VALUE` for one ad hoc override. Repeat it to change more than
one input. Overrides never edit the tracked revision file.

Use `--source PATH` or `--revisions-file PATH` to evaluate an experimental copy
without replacing the tracked baseline. Every JSON report records the exact
resolved inputs and their hashes, so an interesting result can be reproduced
or rejected before any source change is selected.

## Export every run

Aggregate reports are normally enough. When a distribution or outlier needs
independent analysis, add `--runs-output`:

```powershell
python .\sites\d20-adventurers-pod\scripts\simulate-dungeon.py `
  --runs 10000 --revision latest `
  --format json --output .\reports\d20-summary.json `
  --runs-output .\reports\d20-runs.jsonl
```

The JSON Lines file writes one self-contained record per scenario and seed,
including outcome, stage, gate rounds, Party damage sources, answer use,
Vitality contribution, Encounter draw/cast/resolve/answer/removal counts,
Focus counts, lair rolls, and table-load maxima. It intentionally omits verbose
trace events; request exact trace seeds for those.

Summary output supports `text`, `json`, and `csv`. Use `--output` to write a
file, or omit it to print to the terminal. `--list-revisions` shows the
available historical comparison ids.

## Reading the diagnostics

- **Win and boss reach** answer whether the pacing envelope reaches the desired
  destination, not whether the cards are fun.
- **Confidence intervals** expose sampling noise. A value barely over a target
  is not robust merely because its point estimate passes.
- **Outcome causes** distinguish Party-Life pressure from the Doom clock.
- **Gate distributions** show where expeditions stall and whether a stage is
  consuming too many rounds.
- **Damage sources** separate combat, direct Encounters, persistent Encounters,
  Room pressure, Heat, and setup.
- **Encounter counts** reveal dead cards, answer magnets, cleanup losses, and
  late graveyard recycling.
- **Table load** estimates the physical burden of Reserve and Dungeon
  permanents. It reports both non-Depth objects and the total after ordinary
  Depth creation, including approximate setup, lair, and boss object counts.
  It does not count adventurer boards and may compress several tokens created
  by one coarse Encounter value into a single modeled permanent.
- **Paired outcome flips** show which exact seeds changed result under a tuning
  variant. Trace a flipped seed before accepting the aggregate explanation.

## Refinement discipline

Change only one conceptual layer between selected comparisons. A sensible
order is:

1. confirm rules and timing;
2. inspect gate and loss-cause distributions;
3. tune scenario pressure or vitality;
4. inspect each adventurer's answer use and gated contribution;
5. use the opening helper for changes to starting resources or mulligans;
6. trace representative wins, losses, openings, and paired flips; and
7. validate the idea in paper play.

Never tune individual card rules from the model alone. Record target ambiguity,
combat texture, signature moments, decision quality, and administrative burden
at the table; those are intentionally outside this simulator.
