Document OCR processing workflow

This commit is contained in:
2026-07-17 18:03:20 +02:00
parent a55cf8da15
commit f34f74697b

497
README.md
View File

@ -1,31 +1,40 @@
# WCX Index
Lokalt index för WCX-publiceringar.
Local SQLite-based index for WCX publications.
Projektet bygger ett SQLite-baserat index från JSON-data som genereras av det befintliga scraper-scriptet.
The project imports metadata produced by an existing site scraper, stores it in a SQLite database, and enriches new records with metadata extracted from thumbnail images using Google Cloud Vision OCR.
## Nuvarande flöde
## Current processing flow
```text
WCX-site
WCX site
befintligt scraper-script
existing scraper
import/wcx_site_index.json
scripts/import_site.py
database/wcx.db
scripts/check_ocr.py
├── scripts/ocr.sh
└── scripts/parse_ocr.py
OCR metadata stored in SQLite
```
Importscriptet:
The current implementation supports:
* lägger till nya filmer
* uppdaterar befintliga filmer när sitens `update`-datum är nyare
* lämnar oförändrade filmer orörda
* lagrar speltid som antal sekunder
* inserting new movies from the scraper JSON
* updating existing movies based on the site's `update` field
* storing movie duration as seconds
* running OCR for an individual movie
* parsing and validating OCR output
* comparing the OCR name with the database name
* storing OCR metadata, raw OCR text, status, and errors
## Katalogstruktur
## Directory structure
```text
/storage/disk1/WCX/
@ -33,79 +42,117 @@ Importscriptet:
│ └── wcx.db
├── import/
│ └── wcx_site_index.json
── scripts/
├── schema.sql
── import_site.py
── scripts/
├── schema.sql
── import_site.py
│ ├── ocr.sh
│ ├── parse_ocr.py
│ └── check_ocr.py
├── .gitignore
└── README.md
```
## Förutsättningar
## Requirements
Projektet använder:
The project currently uses:
* Python 3
* SQLite 3
* jq, valfritt för inspektion av JSON
* Bash
* curl
* jq
* base64
* Google Cloud Vision API access
Ubuntu/Debian:
Install the local command-line dependencies on Ubuntu or Debian:
```bash
sudo apt update
sudo apt install sqlite3 jq
sudo apt install sqlite3 jq curl
```
Python-modulerna `json`, `sqlite3` och `pathlib` ingår i Pythons standardbibliotek. Inga externa Python-paket behövs.
The Python scripts only use modules from the Python standard library.
## Skapa databasen
## Google Vision API key
Databasen skapas från `schema.sql`.
The OCR script expects the Google Vision API key in the environment:
```bash
export GOOGLE_VISION_API_KEY="your-api-key"
```
The key must be available in the environment whenever `ocr.sh` or `check_ocr.py` is executed.
Do not commit the API key to Git.
## Creating the database
Create the database directory:
```bash
mkdir -p /storage/disk1/WCX/database
```
Create the database from the schema:
```bash
sqlite3 /storage/disk1/WCX/database/wcx.db \
< /storage/disk1/WCX/scripts/schema.sql
```
Kontrollera tabellen:
Inspect the resulting table:
```bash
sqlite3 /storage/disk1/WCX/database/wcx.db ".schema movie"
```
## Databasmodell
## Database model
Tabellen `movie` innehåller:
The `movie` table currently contains the following columns:
| Kolumn | Beskrivning |
| ------------------ | ------------------------------------------ |
| `id` | Stabilt och unikt ID från siten |
| `name` | Filmens titel |
| `aka` | Alternativt namn |
| `nationality` | Nationalitet |
| `age` | Ålder |
| `shoot_location` | Inspelningsplats |
| `shoot_date` | Inspelningsdatum |
| `duration_seconds` | Speltid i sekunder |
| `description` | Manuell beskrivning |
| `rating` | Manuell rating |
| `web_url` | Länk till detaljsidan |
| `thumbnail` | Länk till thumbnail |
| `published` | Ursprungligt publiceringsdatum |
| `updated` | Senaste uppdateringsdatum från siten |
| `created_at` | När posten skapades i databasen |
| `modified_at` | När posten senast uppdaterades i databasen |
| Column | Description |
| ------------------ | -------------------------------------------------- |
| `id` | Stable and unique ID from the site |
| `name` | Movie title |
| `aka` | Alternative name |
| `nationality` | Nationality extracted by OCR or edited manually |
| `age` | Manually maintained age |
| `shoot_location` | Shoot location extracted by OCR or edited manually |
| `shoot_date` | Shoot date stored as `YYYY-MM-DD` |
| `duration_seconds` | Movie duration stored in seconds |
| `description` | Manually maintained description |
| `rating` | Manually maintained rating |
| `web_url` | URL to the movie detail page |
| `thumbnail` | URL to the thumbnail image |
| `published` | Original publication date |
| `updated` | Latest site update date |
| `created_at` | Timestamp when the database record was created |
| `modified_at` | Timestamp when the record was last modified |
| `ocr_status` | Current OCR processing status |
| `ocr_raw_text` | Unmodified OCR text returned by Google Vision |
| `ocr_error` | OCR or validation error |
| `ocr_processed_at` | Timestamp of the latest OCR attempt |
## JSON-indata
Possible OCR statuses are currently:
Importscriptet läser:
```text
pending
completed
failed
manual_review
```
## Site JSON input
The site importer reads:
```text
/storage/disk1/WCX/import/wcx_site_index.json
```
Filen ska innehålla en JSON-array.
The file must contain a JSON array.
Exempel:
Example entry:
```json
{
@ -118,7 +165,7 @@ Exempel:
}
```
Vid en uppdaterad publicering kan även följande fält finnas:
An updated publication may also contain:
```json
{
@ -126,44 +173,44 @@ Vid en uppdaterad publicering kan även följande fält finnas:
}
```
Importscriptet accepterar både `update` och `updated`.
The importer accepts both `update` and `updated`.
## Fältmappning
## Site field mapping
| JSON-fält | Databaskolumn |
| ------------------------ | ------------------ |
| JSON field | Database column |
| --------------------- | ------------------ |
| `id` | `id` |
| `titel` | `name` |
| `details` | `web_url` |
| `duration` | `duration_seconds` |
| `thumb` | `thumbnail` |
| `published` | `published` |
| `update` eller `updated` | `updated` |
| `update` or `updated` | `updated` |
Speltiden konverteras till sekunder.
Durations are converted to seconds.
Exempel:
Examples:
```text
37:35 → 2255
1:37:00 → 5820
```
## Kör importen
## Importing site data
Gör scriptet körbart om det inte redan är gjort:
Make the importer executable:
```bash
chmod +x /storage/disk1/WCX/scripts/import_site.py
```
Kör importen:
Run the import:
```bash
/storage/disk1/WCX/scripts/import_site.py
```
Exempel på resultat:
Example output:
```text
Poster i JSON: 1250
@ -172,24 +219,30 @@ Uppdaterade: 1
Oförändrade: 1246
```
## Importregler
## Site import rules
### Ny film
### New movie
Om filmens ID inte finns i databasen skapas en ny post.
If the movie ID does not exist in the database, a new row is inserted.
### Befintlig film utan `update`
New rows receive the default OCR status:
Om filmen redan finns och site-posten saknar `update` görs ingen ändring.
```text
pending
```
### Befintlig film med `update`
### Existing movie without a site update
Posten uppdateras när:
If the movie already exists and the site entry has no `update` date, no database changes are made.
* databasen saknar ett `updated`-värde, eller
* sitens `update`-datum är nyare än databasens `updated`-värde
### Existing movie with a site update
Vid en uppdatering skrivs följande sitefält över:
An existing row is updated when:
* the database has no `updated` value and the site does, or
* the site's update date is newer than the database value
When an update is triggered, all fields supplied by the site are replaced:
* `name`
* `duration_seconds`
@ -198,20 +251,20 @@ Vid en uppdatering skrivs följande sitefält över:
* `published`
* `updated`
ID och manuella metadatafält påverkas inte.
The movie ID and manually maintained metadata fields are left unchanged.
Datumen lagras som `YYYY-MM-DD`, vilket gör att de kan jämföras kronologiskt som text.
Dates use the ISO format `YYYY-MM-DD`, allowing chronological comparison as text.
## Inspektera databasen
## Inspecting the database
Visa antal filmer:
Show the total number of movies:
```bash
sqlite3 /storage/disk1/WCX/database/wcx.db \
"SELECT COUNT(*) FROM movie;"
```
Visa några poster:
Show some movies:
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
@ -220,7 +273,7 @@ sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
LIMIT 10;"
```
Visa en specifik film:
Show a specific movie:
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
@ -229,29 +282,238 @@ sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
WHERE id = 'susana-melo_6707';"
```
## Testa uppdateringslogiken
## OCR processing
Ett enkelt test är att medvetet ändra en post:
OCR processing is currently performed for one movie at a time.
The complete flow is:
```text
movie ID
check_ocr.py
thumbnail URL loaded from SQLite
ocr.sh
Google Cloud Vision raw text
parse_ocr.py
name validation
SQLite update
```
## Raw OCR script
Run OCR directly against an image URL:
```bash
sqlite3 /storage/disk1/WCX/database/wcx.db \
"UPDATE movie
SET duration_seconds = 1,
updated = NULL
/storage/disk1/WCX/scripts/ocr.sh \
"https://example.com/thumbnail.jpg"
```
Example output:
```text
SUSANA MELO
PORTUGUESE
WOODMAN CASTINGX.COM
Budapest (Hungary), March 9, 2014
```
The raw OCR output is not considered a stable data format and must be parsed and validated before metadata is stored.
## Expected OCR structure
The parser currently expects exactly four non-empty lines:
```text
Line 1: Performer name
Line 2: Nationality
Line 3: Branding text, ignored
Line 4: Shoot location and shoot date
```
Example final line:
```text
Budapest (Hungary), March 9, 2014
```
The line is split on the first comma:
```text
shoot_location = Budapest (Hungary)
shoot_date = March 9, 2014
```
The date is normalized before storage:
```text
March 9, 2014 → 2014-03-09
```
## Testing the OCR parser
Raw OCR output can be piped directly into the parser:
```bash
/storage/disk1/WCX/scripts/ocr.sh \
"https://example.com/thumbnail.jpg" \
| /storage/disk1/WCX/scripts/parse_ocr.py
```
Example result:
```text
ocr_name=SUSANA MELO
nationality=PORTUGUESE
shoot_location=Budapest (Hungary)
shoot_date=2014-03-09
```
## Processing OCR for a movie
Run the complete OCR workflow using a movie ID:
```bash
/storage/disk1/WCX/scripts/check_ocr.py susana-melo_6707
```
Example output:
```text
id=susana-melo_6707
database_name=Susana Melo
ocr_name=SUSANA MELO
name_match=yes
nationality=PORTUGUESE
shoot_location=Budapest (Hungary)
shoot_date=2014-03-09
ocr_status=completed
```
The name comparison is case-insensitive and ignores repeated whitespace.
For example:
```text
Susana Melo
SUSANA MELO
```
are considered equal.
## Successful OCR processing
When OCR execution and parsing succeed and the OCR name matches the database name, the following fields are updated:
```text
nationality
shoot_location
shoot_date
ocr_raw_text
ocr_status = completed
ocr_error = NULL
ocr_processed_at
modified_at
```
## OCR failure handling
If the Google Vision request or `ocr.sh` execution fails:
```text
ocr_status = failed
```
The error is stored in:
```text
ocr_error
```
## Manual review handling
The movie is marked for manual review when:
* the OCR output cannot be parsed
* the expected fields are missing
* the OCR name does not match the database name
* the OCR output has an unexpected structure
In this case:
```text
ocr_status = manual_review
```
The raw OCR text is retained in `ocr_raw_text`, but the parsed metadata fields are not updated.
This prevents uncertain OCR output from silently replacing valid metadata.
## Inspecting OCR results
Show the structured OCR result:
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
"SELECT
id,
name,
nationality,
shoot_location,
shoot_date,
ocr_status,
ocr_error,
ocr_processed_at
FROM movie
WHERE id = 'susana-melo_6707';"
```
Kör sedan importen igen:
Show the raw OCR text:
```bash
/storage/disk1/WCX/scripts/import_site.py
sqlite3 /storage/disk1/WCX/database/wcx.db \
"SELECT ocr_raw_text
FROM movie
WHERE id = 'susana-melo_6707';"
```
Posten ska då uppdateras från JSON-filen och `duration_seconds` ska återställas till sitens värde.
List movies waiting for OCR:
## Återskapa databasen
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
"SELECT id, name, ocr_status
FROM movie
WHERE ocr_status = 'pending'
ORDER BY published DESC;"
```
Under utveckling kan databasen enkelt återskapas eftersom all automatisk site-data kan importeras igen.
List movies requiring manual review:
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
"SELECT id, name, ocr_error
FROM movie
WHERE ocr_status = 'manual_review';"
```
List failed OCR attempts:
```bash
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
"SELECT id, name, ocr_error
FROM movie
WHERE ocr_status = 'failed';"
```
## Recreating the database
During development, the database can be recreated from the schema and site JSON:
```bash
rm /storage/disk1/WCX/database/wcx.db
@ -262,31 +524,70 @@ sqlite3 /storage/disk1/WCX/database/wcx.db \
/storage/disk1/WCX/scripts/import_site.py
```
Radera inte databasen på detta sätt när den senare innehåller manuellt registrerad metadata utan att först ta en backup.
Do not recreate the database this way after it contains manually maintained metadata unless a backup has been created first.
## Versionshantering
OCR data must also be recreated if the database is deleted.
Källkod och dokumentation ska versionshanteras i Git.
## Git and generated data
Databasen och importerade datafiler bör normalt inte checkas in.
The Git repository contains source code, schema, and documentation.
Föreslagen `.gitignore`:
The SQLite database and scraper output are runtime data and are not version-controlled.
Current `.gitignore` rules:
```gitignore
# SQLite database files
database/*.db
database/*.db-shm
database/*.db-wal
# Imported/generated data
import/*.json
# Python cache
__pycache__/
*.pyc
```
## Nästa steg
Files that should be version-controlled include:
Planerade kommande delar:
```text
README.md
.gitignore
scripts/schema.sql
scripts/import_site.py
scripts/ocr.sh
scripts/parse_ocr.py
scripts/check_ocr.py
```
* OCR-behandling av thumbnails för nya filmer
* manuell redigering av metadata
* export av komplett index
* integration med Emby
* schemalagd körning
Files that should not be version-controlled include:
```text
database/wcx.db
database/wcx.db-shm
database/wcx.db-wal
import/wcx_site_index.json
```
## Current limitations
The current OCR parser assumes a four-line OCR result.
Google Cloud Vision does not guarantee this exact structure. Unexpected output is therefore stored for manual review rather than automatically accepted.
OCR processing is currently started manually for one movie ID at a time.
## Planned next steps
Potential next steps include:
* automatically process movies with `ocr_status = pending`
* add retry support for failed OCR requests
* provide commands for manually approving or correcting OCR results
* add database backup handling
* export the complete index to JSON or CSV
* add Emby-compatible metadata export
* schedule the scraper and import process
* create an orchestration script for the complete automated workflow