Document legacy CSV migration

This commit is contained in:
2026-07-17 21:49:41 +02:00
parent ebd38d145e
commit 1e47a66ff8

151
README.md
View File

@ -528,6 +528,157 @@ Do not recreate the database this way after it contains manually maintained meta
OCR data must also be recreated if the database is deleted. OCR data must also be recreated if the database is deleted.
## Legacy CSV migration
The historical index is stored in:
```text
migration/persons.utf8bom.csv
```
This file is version-controlled and can be used to restore manually maintained metadata that is not available from the WCX site or thumbnail OCR.
The migration source contains:
* stable movie IDs
* manually corrected names and aliases
* nationality, age, shoot location, and shoot date
* duration
* descriptions and ratings
* site URLs, thumbnail URLs, and publication dates
The movie ID corresponds to the filename portion of `WEBURL` without the `.html` extension.
Example:
```text
WEBURL: https://www.woodmancastingx.com/casting-x/ysana_268.html
ID: ysana_268
```
### Migration script
The migration is performed by:
```text
scripts/migrate_csv.py
```
The script:
* reads UTF-8 CSV files with BOM support
* validates required headers
* validates ISO dates
* converts durations to seconds
* converts historical age notation to whole years
* trims surrounding whitespace
* verifies that `ID` matches the ID derived from `WEBURL`
* updates existing database rows
* inserts CSV rows that are missing from the site-derived database
* skips rows without an ID
* runs database writes in a transaction
Historical age values may include additional month or week notation:
```text
21.2 → 21
18,1 → 18
19,1w → 19
18 (same day) → 18
? → NULL
```
Only whole years are stored in the database.
### Dry-run validation
Always validate the migration source before applying it:
```bash
/storage/disk1/WCX/scripts/migrate_csv.py \
--database /storage/disk1/WCX/database/wcx-test.db \
--dry-run
```
A limited validation run can be performed with:
```bash
/storage/disk1/WCX/scripts/migrate_csv.py \
--database /storage/disk1/WCX/database/wcx-test.db \
--limit 20 \
--dry-run
```
The report includes:
* rows read
* valid and invalid rows
* duplicate IDs
* rows already present in the database
* rows missing from the database
* ID and `WEBURL` mismatches
* warnings and errors
Rows without an ID are skipped and reported as warnings.
### Test migration
Create a test database from the current database:
```bash
cp /storage/disk1/WCX/database/wcx.db \
/storage/disk1/WCX/database/wcx-test.db
```
Apply the migration to the test database:
```bash
/storage/disk1/WCX/scripts/migrate_csv.py \
--database /storage/disk1/WCX/database/wcx-test.db \
--apply
```
Inspect the test database before applying the migration to the main database.
### Applying the migration
Create a backup first:
```bash
cp /storage/disk1/WCX/database/wcx.db \
/storage/disk1/WCX/database/wcx-before-csv-migration.db
```
Apply the migration:
```bash
/storage/disk1/WCX/scripts/migrate_csv.py \
--database /storage/disk1/WCX/database/wcx.db \
--apply
```
Existing rows are updated from the CSV metadata. A CSV row whose ID is not present in the site-derived database is inserted. This allows locally retained movies to remain in the index even if their publication has been removed from the site.
Migrated rows with complete OCR-related metadata receive:
```text
ocr_status = completed
```
Rows with incomplete OCR-related metadata receive:
```text
ocr_status = manual_review
```
New site publications that are not present in the historical CSV retain:
```text
ocr_status = pending
```
The migration is intended primarily for initial population or database restoration. Repeated execution is possible, but it will overwrite the corresponding movie metadata with values from the version-controlled CSV source.
## Git and generated data ## Git and generated data
The Git repository contains source code, schema, and documentation. The Git repository contains source code, schema, and documentation.