# Project Map

## Structure

| Directory | Contents |
|-----------|----------|
| `docs/methodology.md` | Assessment phases, validation levels L1-L4 |
| `docs/quality-criteria.md` | Completeness and validation standards |
| `docs/log-source-discovery.md` | Log-source taxonomy, suitability scoring, source-discovery methodology |
| `docs/scripts-index.md` | Script categorization (core pipeline vs engagement one-offs) |
| `docs/processing/` | Per-topic analysis notes |
| `findings/` | Raw data: dumps, batches, vault, breaches, keepass, moneros |
| `findings/data/` | Aggregated extracts, analysis results (.csv, .tsv, .md) |
| `artifacts/` | Deliverables: executive summary, gap analysis, remediation matrix |
| `scripts/` | Automation: parsers, checkers, exporters |
| `references/` | Authorization, standards, external docs |
| `redteam/` | Per-target engagement dossiers (L1-L4 validation, evidence). Local-only — gitignored, never committed. Contains attacker-side working files (creds TSVs, recon logs, JWTs). |

## Workflow

1. Source data in `findings/` → process with `scripts/` → results in `findings/data/`
2. Analysis in `docs/processing/` → synthesize into `artifacts/`
3. Validation per `docs/methodology.md` levels L1-L4
4. Operational rules: `.claude/rules/` (loaded per path scope)
   - `validation.md` — [VERIFIED] = command executed + output observed, no speculation
   - `coverage.md` — process every entry, no skipping
   - `operator-gate.md` — L3/L4 require explicit "go"; never auto-execute
   - `no-masking.md` — never truncate/redact secret values in any script output
     (JWT/tokens/passwords/env vars/etc.) unless the user explicitly asks.
     List-pagination and ID-shortening are OK; value-truncation is not.
     Applies to all scripts that touch findings/scripts/redteam/artifacts.

## Canonical status

`artifacts/status-matrix.csv` (phases PH1-PH5, gaps, artifacts) is the canonical
machine-readable source of truth. Keep narrative docs aligned with it via
`make sync-status`. See `docs/status-workflow.md`.

## File Naming

- `{SOURCE}-{ID}_remote_access.md` — per-source access analysis
- `{SOURCE}-{ID}.tsv` — raw credential extract
- `batch-{UUID}-results.txt` — batch processing output
- `*_creds.csv` — aggregated credential data
- `corp_search_{SOURCE}_results.md` — per-source corporate-access search output
- `{SOURCE}_api_keys.tsv` — per-source secret-pattern scan output
- `{SOURCE}_rce_candidates.tsv` — per-source RCE candidate L0 probe results
- `{SOURCE}_L1_{provider}.tsv` — per-source L1 validation results for a specific provider

## Breach Intake & Post-Intake Analysis

Local breach-archive intake is handled by `scripts/intake_local.py` (no SMB / no
network fetch). Each archive becomes a "source" with a stable ID derived from the
archive name. Idempotent: re-running on an already-processed archive is safe.

```bash
python3 scripts/intake_local.py --archive "NAME.rar"                       # password auto-guessed for known channels
python3 scripts/intake_local.py --archive "NAME.rar" --password "@PIXELCLOUD3"   # explicit password
python3 scripts/intake_local.py --extracted "findings/breaches/SOURCE_ID" --source SOURCE_ID  # archive already extracted
```

Per archive, the script runs 5 steps: test integrity → extract → aggregate_breach.py →
corp_search.py → append to findings/data/index.csv. See SETUP.md for source-ID
derivation rules and supported stealer-log formats.

Post-intake analysis scripts (operate on the resulting `findings/data/<SOURCE>.tsv`):

```bash
python3 scripts/scan_secrets.py --source SOURCE_ID                       # scan TSV for API keys/tokens
python3 scripts/scan_secrets.py --source SOURCE_ID --validate            # + L1 validate (read-only API checks)
python3 scripts/scan_secrets.py --all                                    # scan every TSV in findings/data/
python3 scripts/probe_rce.py --source SOURCE_ID                          # L0 reachability for RCE candidates
python3 scripts/probe_rce.py --source SOURCE_ID --forms                  # + extract login form fields
python3 scripts/corp_search.py --source SOURCE_ID                        # per-source corp search (avoids rename-hack)
python3 scripts/corp_search.py --source SOURCE_ID --category infra_rce   # one category
```

DIY Telegram bot collection (see docs/log-source-discovery.md §5.2, §8.1):

```bash
python3 scripts/tg_bot_harvest.py --source threatfox   # harvest + validate bot tokens (abuse.ch Auth-Key in .env as MB_AUTH_KEY)
python3 scripts/tg_bot_collect.py                      # daemon: poll getUpdates → per-victim ZIPs to findings/breaches/TGBOT-<id>/
python3 scripts/tg_bot_collect.py --once               # single poll round
```

### Known channel passwords

Auto-detected by `intake_local.py` from the archive name via the `KNOWN_PASSWORDS`
dict (extend it when new channels appear). See SETUP.md for the current set.
