# Data Inventory

## Purpose

This document explains the structure and update rules for
[findings/data/index.csv](findings/data/index.csv:1).

`index.csv` is the machine-readable registry of important datasets and derived
outputs. It complements [docs/pipeline.md](docs/pipeline.md:1):

- `pipeline.md` explains flows: `input -> script -> output -> artifact`
- `findings/data/index.csv` lists actual tracked outputs and their current role

## Schema

Columns:

- `dataset_id`
  Stable identifier for the dataset or output.

- `path`
  Repository-relative path.

- `kind`
  One of: `raw`, `intermediate`, `derived`, `artifact`, `reference`.

- `source`
  High-level origin, for example `vault`, `dumps`, `batches`, `telegram_logs`,
  `keepass`, `multi-source`, `osint`.

- `format`
  File format such as `csv`, `tsv`, `md`, `json`, `txt`, `ods`.

- `producer`
  Script or process that creates the dataset. Use `manual_analysis`,
  `manual_extract`, or `manual_export` when no stable script exists yet.

- `consumers`
  Pipe-delimited list of downstream documents or analysis steps.

- `status`
  Use:
  - `active` for current tracked outputs
  - `partial` when the output exists but is incomplete or blocked
  - `stale` when it should not be trusted as current
  - `deprecated` when retained only for history

- `notes`
  Short operational description.

## Update Rules

Add or update an entry when:
1. a new high-value file appears in `findings/data/`;
2. a producer script changes its output contract;
3. a dataset becomes stale, partial, or deprecated;
4. a downstream artifact starts depending on the dataset.

## Commands

```bash
make validate-data-index
make data-index-report
```

Both targets currently run [scripts/validate_data_inventory.py](scripts/validate_data_inventory.py:1).

- `make validate-data-index` validates schema, paths, producers, consumers, and allowed status values
- `make data-index-report` is the same command but can be used as a read-only reporting entrypoint in local workflows

## Current Scope

`index.csv` is intentionally selective. It should track:
- high-value derived outputs;
- validation outputs;
- datasets referenced by `artifacts/` and `docs/processing/`;
- files that materially affect reproducibility.

It does not need to enumerate every scratch file, lock file, or ad hoc export.
