Document legacy CSV migration
This commit is contained in:
151
README.md
151
README.md
@ -528,6 +528,157 @@ Do not recreate the database this way after it contains manually maintained meta
|
||||
|
||||
OCR data must also be recreated if the database is deleted.
|
||||
|
||||
## Legacy CSV migration
|
||||
|
||||
The historical index is stored in:
|
||||
|
||||
```text
|
||||
migration/persons.utf8bom.csv
|
||||
```
|
||||
|
||||
This file is version-controlled and can be used to restore manually maintained metadata that is not available from the WCX site or thumbnail OCR.
|
||||
|
||||
The migration source contains:
|
||||
|
||||
* stable movie IDs
|
||||
* manually corrected names and aliases
|
||||
* nationality, age, shoot location, and shoot date
|
||||
* duration
|
||||
* descriptions and ratings
|
||||
* site URLs, thumbnail URLs, and publication dates
|
||||
|
||||
The movie ID corresponds to the filename portion of `WEBURL` without the `.html` extension.
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
WEBURL: https://www.woodmancastingx.com/casting-x/ysana_268.html
|
||||
ID: ysana_268
|
||||
```
|
||||
|
||||
### Migration script
|
||||
|
||||
The migration is performed by:
|
||||
|
||||
```text
|
||||
scripts/migrate_csv.py
|
||||
```
|
||||
|
||||
The script:
|
||||
|
||||
* reads UTF-8 CSV files with BOM support
|
||||
* validates required headers
|
||||
* validates ISO dates
|
||||
* converts durations to seconds
|
||||
* converts historical age notation to whole years
|
||||
* trims surrounding whitespace
|
||||
* verifies that `ID` matches the ID derived from `WEBURL`
|
||||
* updates existing database rows
|
||||
* inserts CSV rows that are missing from the site-derived database
|
||||
* skips rows without an ID
|
||||
* runs database writes in a transaction
|
||||
|
||||
Historical age values may include additional month or week notation:
|
||||
|
||||
```text
|
||||
21.2 → 21
|
||||
18,1 → 18
|
||||
19,1w → 19
|
||||
18 (same day) → 18
|
||||
? → NULL
|
||||
```
|
||||
|
||||
Only whole years are stored in the database.
|
||||
|
||||
### Dry-run validation
|
||||
|
||||
Always validate the migration source before applying it:
|
||||
|
||||
```bash
|
||||
/storage/disk1/WCX/scripts/migrate_csv.py \
|
||||
--database /storage/disk1/WCX/database/wcx-test.db \
|
||||
--dry-run
|
||||
```
|
||||
|
||||
A limited validation run can be performed with:
|
||||
|
||||
```bash
|
||||
/storage/disk1/WCX/scripts/migrate_csv.py \
|
||||
--database /storage/disk1/WCX/database/wcx-test.db \
|
||||
--limit 20 \
|
||||
--dry-run
|
||||
```
|
||||
|
||||
The report includes:
|
||||
|
||||
* rows read
|
||||
* valid and invalid rows
|
||||
* duplicate IDs
|
||||
* rows already present in the database
|
||||
* rows missing from the database
|
||||
* ID and `WEBURL` mismatches
|
||||
* warnings and errors
|
||||
|
||||
Rows without an ID are skipped and reported as warnings.
|
||||
|
||||
### Test migration
|
||||
|
||||
Create a test database from the current database:
|
||||
|
||||
```bash
|
||||
cp /storage/disk1/WCX/database/wcx.db \
|
||||
/storage/disk1/WCX/database/wcx-test.db
|
||||
```
|
||||
|
||||
Apply the migration to the test database:
|
||||
|
||||
```bash
|
||||
/storage/disk1/WCX/scripts/migrate_csv.py \
|
||||
--database /storage/disk1/WCX/database/wcx-test.db \
|
||||
--apply
|
||||
```
|
||||
|
||||
Inspect the test database before applying the migration to the main database.
|
||||
|
||||
### Applying the migration
|
||||
|
||||
Create a backup first:
|
||||
|
||||
```bash
|
||||
cp /storage/disk1/WCX/database/wcx.db \
|
||||
/storage/disk1/WCX/database/wcx-before-csv-migration.db
|
||||
```
|
||||
|
||||
Apply the migration:
|
||||
|
||||
```bash
|
||||
/storage/disk1/WCX/scripts/migrate_csv.py \
|
||||
--database /storage/disk1/WCX/database/wcx.db \
|
||||
--apply
|
||||
```
|
||||
|
||||
Existing rows are updated from the CSV metadata. A CSV row whose ID is not present in the site-derived database is inserted. This allows locally retained movies to remain in the index even if their publication has been removed from the site.
|
||||
|
||||
Migrated rows with complete OCR-related metadata receive:
|
||||
|
||||
```text
|
||||
ocr_status = completed
|
||||
```
|
||||
|
||||
Rows with incomplete OCR-related metadata receive:
|
||||
|
||||
```text
|
||||
ocr_status = manual_review
|
||||
```
|
||||
|
||||
New site publications that are not present in the historical CSV retain:
|
||||
|
||||
```text
|
||||
ocr_status = pending
|
||||
```
|
||||
|
||||
The migration is intended primarily for initial population or database restoration. Repeated execution is possible, but it will overwrite the corresponding movie metadata with values from the version-controlled CSV source.
|
||||
|
||||
## Git and generated data
|
||||
|
||||
The Git repository contains source code, schema, and documentation.
|
||||
|
||||
Reference in New Issue
Block a user