# ir-assessment — Setup & Agent Onboarding

## What this is

Clean project skeleton for breach-log post-compromise scope assessment:
code + docs + operational rules only. No raw data, no artifacts, no redteam
files. Populate `findings/` via the intake pipeline, generate `artifacts/`
as analysis matures. Assessment methodology: NIST SP 800-61r3, PTES
(Post-Exploitation), SANS FOR508/FOR572, MITRE ATT&CK (T1078, T1539, T1528),
FIRST CSIRT, OWASP WSTG. See `docs/methodology.md`.

## Prerequisites (system)

- Python 3.11+
- `unrar` and `7z` (archive extraction in intake_local)

Debian/Ubuntu:
    sudo apt install unrar p7zip-full

## Install

    python3 -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

## Verify

    pytest tests/ -x                          # offline-pipeline smoke tests
    python3 scripts/intake_local.py --help    # local intake CLI

## Project layout

    scripts/
      intake_local.py           MAIN: local archive → extract → aggregate → corp_search → index
      aggregate_breach.py       CORE: per-victim dirs → TSV (3 formats auto-detected)
      corp_search.py            127 patterns, 17 categories, 3-layer search
      scan_secrets.py           60+ regex patterns + L1 API validation (optional)
      probe_rce.py              L0 reachability for RCE-grade endpoints
      kdbx_inventory.py         KeePass inventory + hash buckets
      tg_session_load.py        Telegram breach-archive access (pyrogram)
      tg_list_channel.py        List Telegram channel messages
      tg_download.py            Resumable channel document download (offset-resume)
      tg_bot_harvest.py         Bot token harvest (ThreatFox / MalwareBazaar) + Bot API validation
      tg_bot_collect.py         Passive getUpdates listener → per-victim ZIPs
      tg_group_read.py          Group-join reader: full exfil-group history via invite link
      tg_botapi_probe.py        Bot API visibility probe (methodology §8.2)
      tg_l1l2_browser.py        Playwright browser L1+L2 (phpMyAdmin, Jenkins)
      validate_remaining_tokens.py  Residual token validation (ghp_, glpat-, dckr_pat_, Azure PAT)
      validate_status_matrix.py Validate status-matrix.csv + narrative sync
      generate_status_summary.py Regen managed block in gap-analysis.md
      status_report.py          Read-only status summary
      validate_data_inventory.py    Validate findings/data/index.csv
      validate_artifact_inventory.py Validate artifacts/index.csv
      password_reset_chain.py   Password reset chain inference
      monetary_impact.py        Monetary impact assessment
      analyze_cookies.py        Cookie session analysis
      analyze_credit_cards.py   Credit card analysis
      vault_creds_export.py     Vault credentials export
      legacy_dump_assessment/   standalone KeePass/dump triage tools (stdlib)
      ulp_to_tsv.py             ULP format → TSV
      fast_download.py          LEGACY (Telethon) — use tg_download.py instead
      extract_tg_hashes.py      Extract Telegram session hashes
      imap_checker.py           IMAP L1 validation (optional, dnspython)
      sanitized_export.py       Sanitized export for sharing
    docs/                       methodology, pipeline map, threat modeling
      methodology.md            L1-L4 levels, victim pivot technique
      log-source-discovery.md   Log-source taxonomy + P0 pilot (bot collection)
      scripts-index.md          Script categorization (what to use vs one-offs)
      pipeline.md               input → script → output → downstream map
      quality-criteria.md       Completeness + tier-list standards
      threat-modeling-guide.md  EC1-EC5 exposure categories
      status-workflow.md        Canonical status source rules
      data-inventory.md         findings/data/index.csv schema
      artifact-inventory.md     artifacts/index.csv schema
      processing/               Per-topic analysis notes (empty — populate during work)
    .claude/rules/              operator-gate, validation, coverage rules
    findings/                   EMPTY — populate via intake_local
    artifacts/                  status-matrix.csv + index.csv skeletons; reports generated during work
    redteam/                    Per-target engagement dossiers (local-only, gitignored)
    CLAUDE.md                   project map

## Core pipeline

    # 1. Intake local breach archive
    python3 scripts/intake_local.py --archive "PATH/TO/6226_05.07.2026_@PIXELCLOUD3.rar"
    python3 scripts/intake_local.py --archive "NAME.rar" --password "@PIXELCLOUD3"
    python3 scripts/intake_local.py --extracted "findings/breaches/SOURCE_ID" --source SOURCE_ID

    # 2. Post-intake analysis
    python3 scripts/scan_secrets.py --source SOURCE_ID                       # secret scan
    python3 scripts/scan_secrets.py --source SOURCE_ID --validate            # + L1 API check
    python3 scripts/probe_rce.py --source SOURCE_ID                          # L0 reachability
    python3 scripts/probe_rce.py --source SOURCE_ID --forms                  # + login forms
    python3 scripts/corp_search.py --source SOURCE_ID                        # corp search

    # 3. Make helpers
    make intake ARCHIVE="NAME.rar" PASSWORD="@PW"
    make scan-all                # scan + validate all sources
    make sync-status             # regenerate + validate status

## Methodology (L1-L4)

See docs/methodology.md. Summary:
- L1 identity: confirm cred active (auto-approved)
- L2 scope: enumerate what cred accesses (auto-approved)
- L3 write: verify modify capability (**operator "go" required**)
- L4 priv: determine admin/elevated (**operator "go" required**)

## Rules (.claude/rules/)

- operator-gate.md: L3-L4 → print + explain + wait for "go". Never auto.
- validation.md: [VERIFIED] = command executed + output observed. No speculation.
- coverage.md: process every entry. No skipping.
- no-masking.md: full secret values in all script outputs.

## Data format

- findings/data/<SOURCE_ID>.tsv — URL<TAB>user<TAB>password (aggregated creds)
- findings/data/<SOURCE_ID>_api_keys.tsv — secret scan output
- findings/data/<SOURCE_ID>_rce_candidates.tsv — L0 probe output
- findings/data/corp_search_<SOURCE_ID>_results.md — corp search output
- findings/data/index.csv — dataset registry

## Source ID derivation (intake_local.py)

    "6226_05.07.2026_@PIXELCLOUD3.rar"                    → PIXELCLOUD3-6226
    "@beetraffic 2000 MIX 05-07-2026"                      → BEETRAFFIC-2000
    "AYANKOUJI PRIVATE #528 PART 4"                         → AYANKOUJI-528
    "@fatetraffic 1700 MIX 10-07-2026.rar"                 → FATETRAFFIC-1700
    "@UP_DAISYCLOUD_CHAMPIONING!_05_JULY_5530_ON_CHANNEL.rar" → UP_DAISYCLOUD-5530
    "WINGSCLOUD-ULP-JUNE-18" (ULP txt)                     → WINGSCLOUD-ULP-JUNE-18

## aggregate_breach.py — supported log formats

Auto-detected per file:
- Format A: Soft/Host/Login/Password (classic cracked stealers)
- Format B: SOFT/URL/USER/PASS
- Format C: Browser/URL/Username/Password with --- separators (new stealers)
- TSV: URL<TAB>USER<TAB>PASS (ready format)

Unicode-obfuscated stealer-channel watermarks (kir3inf variants) auto-filtered.

## Known channel passwords (auto-guessed)

Extend KNOWN_PASSWORDS dict in scripts/intake_local.py when new channels appear.
The template ships with the current operational set already populated.
