Document OCR processing workflow
This commit is contained in:
497
README.md
497
README.md
@ -1,31 +1,40 @@
|
|||||||
# WCX Index
|
# WCX Index
|
||||||
|
|
||||||
Lokalt index för WCX-publiceringar.
|
Local SQLite-based index for WCX publications.
|
||||||
|
|
||||||
Projektet bygger ett SQLite-baserat index från JSON-data som genereras av det befintliga scraper-scriptet.
|
The project imports metadata produced by an existing site scraper, stores it in a SQLite database, and enriches new records with metadata extracted from thumbnail images using Google Cloud Vision OCR.
|
||||||
|
|
||||||
## Nuvarande flöde
|
## Current processing flow
|
||||||
|
|
||||||
```text
|
```text
|
||||||
WCX-site
|
WCX site
|
||||||
↓
|
↓
|
||||||
befintligt scraper-script
|
existing scraper
|
||||||
↓
|
↓
|
||||||
import/wcx_site_index.json
|
import/wcx_site_index.json
|
||||||
↓
|
↓
|
||||||
scripts/import_site.py
|
scripts/import_site.py
|
||||||
↓
|
↓
|
||||||
database/wcx.db
|
database/wcx.db
|
||||||
|
↓
|
||||||
|
scripts/check_ocr.py
|
||||||
|
├── scripts/ocr.sh
|
||||||
|
└── scripts/parse_ocr.py
|
||||||
|
↓
|
||||||
|
OCR metadata stored in SQLite
|
||||||
```
|
```
|
||||||
|
|
||||||
Importscriptet:
|
The current implementation supports:
|
||||||
|
|
||||||
* lägger till nya filmer
|
* inserting new movies from the scraper JSON
|
||||||
* uppdaterar befintliga filmer när sitens `update`-datum är nyare
|
* updating existing movies based on the site's `update` field
|
||||||
* lämnar oförändrade filmer orörda
|
* storing movie duration as seconds
|
||||||
* lagrar speltid som antal sekunder
|
* running OCR for an individual movie
|
||||||
|
* parsing and validating OCR output
|
||||||
|
* comparing the OCR name with the database name
|
||||||
|
* storing OCR metadata, raw OCR text, status, and errors
|
||||||
|
|
||||||
## Katalogstruktur
|
## Directory structure
|
||||||
|
|
||||||
```text
|
```text
|
||||||
/storage/disk1/WCX/
|
/storage/disk1/WCX/
|
||||||
@ -33,79 +42,117 @@ Importscriptet:
|
|||||||
│ └── wcx.db
|
│ └── wcx.db
|
||||||
├── import/
|
├── import/
|
||||||
│ └── wcx_site_index.json
|
│ └── wcx_site_index.json
|
||||||
└── scripts/
|
├── scripts/
|
||||||
├── schema.sql
|
│ ├── schema.sql
|
||||||
└── import_site.py
|
│ ├── import_site.py
|
||||||
|
│ ├── ocr.sh
|
||||||
|
│ ├── parse_ocr.py
|
||||||
|
│ └── check_ocr.py
|
||||||
|
├── .gitignore
|
||||||
|
└── README.md
|
||||||
```
|
```
|
||||||
|
|
||||||
## Förutsättningar
|
## Requirements
|
||||||
|
|
||||||
Projektet använder:
|
The project currently uses:
|
||||||
|
|
||||||
* Python 3
|
* Python 3
|
||||||
* SQLite 3
|
* SQLite 3
|
||||||
* jq, valfritt för inspektion av JSON
|
* Bash
|
||||||
|
* curl
|
||||||
|
* jq
|
||||||
|
* base64
|
||||||
|
* Google Cloud Vision API access
|
||||||
|
|
||||||
På Ubuntu/Debian:
|
Install the local command-line dependencies on Ubuntu or Debian:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sudo apt update
|
sudo apt update
|
||||||
sudo apt install sqlite3 jq
|
sudo apt install sqlite3 jq curl
|
||||||
```
|
```
|
||||||
|
|
||||||
Python-modulerna `json`, `sqlite3` och `pathlib` ingår i Pythons standardbibliotek. Inga externa Python-paket behövs.
|
The Python scripts only use modules from the Python standard library.
|
||||||
|
|
||||||
## Skapa databasen
|
## Google Vision API key
|
||||||
|
|
||||||
Databasen skapas från `schema.sql`.
|
The OCR script expects the Google Vision API key in the environment:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export GOOGLE_VISION_API_KEY="your-api-key"
|
||||||
|
```
|
||||||
|
|
||||||
|
The key must be available in the environment whenever `ocr.sh` or `check_ocr.py` is executed.
|
||||||
|
|
||||||
|
Do not commit the API key to Git.
|
||||||
|
|
||||||
|
## Creating the database
|
||||||
|
|
||||||
|
Create the database directory:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
mkdir -p /storage/disk1/WCX/database
|
mkdir -p /storage/disk1/WCX/database
|
||||||
|
```
|
||||||
|
|
||||||
|
Create the database from the schema:
|
||||||
|
|
||||||
|
```bash
|
||||||
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
||||||
< /storage/disk1/WCX/scripts/schema.sql
|
< /storage/disk1/WCX/scripts/schema.sql
|
||||||
```
|
```
|
||||||
|
|
||||||
Kontrollera tabellen:
|
Inspect the resulting table:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sqlite3 /storage/disk1/WCX/database/wcx.db ".schema movie"
|
sqlite3 /storage/disk1/WCX/database/wcx.db ".schema movie"
|
||||||
```
|
```
|
||||||
|
|
||||||
## Databasmodell
|
## Database model
|
||||||
|
|
||||||
Tabellen `movie` innehåller:
|
The `movie` table currently contains the following columns:
|
||||||
|
|
||||||
| Kolumn | Beskrivning |
|
| Column | Description |
|
||||||
| ------------------ | ------------------------------------------ |
|
| ------------------ | -------------------------------------------------- |
|
||||||
| `id` | Stabilt och unikt ID från siten |
|
| `id` | Stable and unique ID from the site |
|
||||||
| `name` | Filmens titel |
|
| `name` | Movie title |
|
||||||
| `aka` | Alternativt namn |
|
| `aka` | Alternative name |
|
||||||
| `nationality` | Nationalitet |
|
| `nationality` | Nationality extracted by OCR or edited manually |
|
||||||
| `age` | Ålder |
|
| `age` | Manually maintained age |
|
||||||
| `shoot_location` | Inspelningsplats |
|
| `shoot_location` | Shoot location extracted by OCR or edited manually |
|
||||||
| `shoot_date` | Inspelningsdatum |
|
| `shoot_date` | Shoot date stored as `YYYY-MM-DD` |
|
||||||
| `duration_seconds` | Speltid i sekunder |
|
| `duration_seconds` | Movie duration stored in seconds |
|
||||||
| `description` | Manuell beskrivning |
|
| `description` | Manually maintained description |
|
||||||
| `rating` | Manuell rating |
|
| `rating` | Manually maintained rating |
|
||||||
| `web_url` | Länk till detaljsidan |
|
| `web_url` | URL to the movie detail page |
|
||||||
| `thumbnail` | Länk till thumbnail |
|
| `thumbnail` | URL to the thumbnail image |
|
||||||
| `published` | Ursprungligt publiceringsdatum |
|
| `published` | Original publication date |
|
||||||
| `updated` | Senaste uppdateringsdatum från siten |
|
| `updated` | Latest site update date |
|
||||||
| `created_at` | När posten skapades i databasen |
|
| `created_at` | Timestamp when the database record was created |
|
||||||
| `modified_at` | När posten senast uppdaterades i databasen |
|
| `modified_at` | Timestamp when the record was last modified |
|
||||||
|
| `ocr_status` | Current OCR processing status |
|
||||||
|
| `ocr_raw_text` | Unmodified OCR text returned by Google Vision |
|
||||||
|
| `ocr_error` | OCR or validation error |
|
||||||
|
| `ocr_processed_at` | Timestamp of the latest OCR attempt |
|
||||||
|
|
||||||
## JSON-indata
|
Possible OCR statuses are currently:
|
||||||
|
|
||||||
Importscriptet läser:
|
```text
|
||||||
|
pending
|
||||||
|
completed
|
||||||
|
failed
|
||||||
|
manual_review
|
||||||
|
```
|
||||||
|
|
||||||
|
## Site JSON input
|
||||||
|
|
||||||
|
The site importer reads:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
/storage/disk1/WCX/import/wcx_site_index.json
|
/storage/disk1/WCX/import/wcx_site_index.json
|
||||||
```
|
```
|
||||||
|
|
||||||
Filen ska innehålla en JSON-array.
|
The file must contain a JSON array.
|
||||||
|
|
||||||
Exempel:
|
Example entry:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@ -118,7 +165,7 @@ Exempel:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Vid en uppdaterad publicering kan även följande fält finnas:
|
An updated publication may also contain:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
{
|
{
|
||||||
@ -126,44 +173,44 @@ Vid en uppdaterad publicering kan även följande fält finnas:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
Importscriptet accepterar både `update` och `updated`.
|
The importer accepts both `update` and `updated`.
|
||||||
|
|
||||||
## Fältmappning
|
## Site field mapping
|
||||||
|
|
||||||
| JSON-fält | Databaskolumn |
|
| JSON field | Database column |
|
||||||
| ------------------------ | ------------------ |
|
| --------------------- | ------------------ |
|
||||||
| `id` | `id` |
|
| `id` | `id` |
|
||||||
| `titel` | `name` |
|
| `titel` | `name` |
|
||||||
| `details` | `web_url` |
|
| `details` | `web_url` |
|
||||||
| `duration` | `duration_seconds` |
|
| `duration` | `duration_seconds` |
|
||||||
| `thumb` | `thumbnail` |
|
| `thumb` | `thumbnail` |
|
||||||
| `published` | `published` |
|
| `published` | `published` |
|
||||||
| `update` eller `updated` | `updated` |
|
| `update` or `updated` | `updated` |
|
||||||
|
|
||||||
Speltiden konverteras till sekunder.
|
Durations are converted to seconds.
|
||||||
|
|
||||||
Exempel:
|
Examples:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
37:35 → 2255
|
37:35 → 2255
|
||||||
1:37:00 → 5820
|
1:37:00 → 5820
|
||||||
```
|
```
|
||||||
|
|
||||||
## Kör importen
|
## Importing site data
|
||||||
|
|
||||||
Gör scriptet körbart om det inte redan är gjort:
|
Make the importer executable:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
chmod +x /storage/disk1/WCX/scripts/import_site.py
|
chmod +x /storage/disk1/WCX/scripts/import_site.py
|
||||||
```
|
```
|
||||||
|
|
||||||
Kör importen:
|
Run the import:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
/storage/disk1/WCX/scripts/import_site.py
|
/storage/disk1/WCX/scripts/import_site.py
|
||||||
```
|
```
|
||||||
|
|
||||||
Exempel på resultat:
|
Example output:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Poster i JSON: 1250
|
Poster i JSON: 1250
|
||||||
@ -172,24 +219,30 @@ Uppdaterade: 1
|
|||||||
Oförändrade: 1246
|
Oförändrade: 1246
|
||||||
```
|
```
|
||||||
|
|
||||||
## Importregler
|
## Site import rules
|
||||||
|
|
||||||
### Ny film
|
### New movie
|
||||||
|
|
||||||
Om filmens ID inte finns i databasen skapas en ny post.
|
If the movie ID does not exist in the database, a new row is inserted.
|
||||||
|
|
||||||
### Befintlig film utan `update`
|
New rows receive the default OCR status:
|
||||||
|
|
||||||
Om filmen redan finns och site-posten saknar `update` görs ingen ändring.
|
```text
|
||||||
|
pending
|
||||||
|
```
|
||||||
|
|
||||||
### Befintlig film med `update`
|
### Existing movie without a site update
|
||||||
|
|
||||||
Posten uppdateras när:
|
If the movie already exists and the site entry has no `update` date, no database changes are made.
|
||||||
|
|
||||||
* databasen saknar ett `updated`-värde, eller
|
### Existing movie with a site update
|
||||||
* sitens `update`-datum är nyare än databasens `updated`-värde
|
|
||||||
|
|
||||||
Vid en uppdatering skrivs följande sitefält över:
|
An existing row is updated when:
|
||||||
|
|
||||||
|
* the database has no `updated` value and the site does, or
|
||||||
|
* the site's update date is newer than the database value
|
||||||
|
|
||||||
|
When an update is triggered, all fields supplied by the site are replaced:
|
||||||
|
|
||||||
* `name`
|
* `name`
|
||||||
* `duration_seconds`
|
* `duration_seconds`
|
||||||
@ -198,20 +251,20 @@ Vid en uppdatering skrivs följande sitefält över:
|
|||||||
* `published`
|
* `published`
|
||||||
* `updated`
|
* `updated`
|
||||||
|
|
||||||
ID och manuella metadatafält påverkas inte.
|
The movie ID and manually maintained metadata fields are left unchanged.
|
||||||
|
|
||||||
Datumen lagras som `YYYY-MM-DD`, vilket gör att de kan jämföras kronologiskt som text.
|
Dates use the ISO format `YYYY-MM-DD`, allowing chronological comparison as text.
|
||||||
|
|
||||||
## Inspektera databasen
|
## Inspecting the database
|
||||||
|
|
||||||
Visa antal filmer:
|
Show the total number of movies:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
||||||
"SELECT COUNT(*) FROM movie;"
|
"SELECT COUNT(*) FROM movie;"
|
||||||
```
|
```
|
||||||
|
|
||||||
Visa några poster:
|
Show some movies:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
@ -220,7 +273,7 @@ sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
|||||||
LIMIT 10;"
|
LIMIT 10;"
|
||||||
```
|
```
|
||||||
|
|
||||||
Visa en specifik film:
|
Show a specific movie:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
@ -229,29 +282,238 @@ sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
|||||||
WHERE id = 'susana-melo_6707';"
|
WHERE id = 'susana-melo_6707';"
|
||||||
```
|
```
|
||||||
|
|
||||||
## Testa uppdateringslogiken
|
## OCR processing
|
||||||
|
|
||||||
Ett enkelt test är att medvetet ändra en post:
|
OCR processing is currently performed for one movie at a time.
|
||||||
|
|
||||||
|
The complete flow is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
movie ID
|
||||||
|
↓
|
||||||
|
check_ocr.py
|
||||||
|
↓
|
||||||
|
thumbnail URL loaded from SQLite
|
||||||
|
↓
|
||||||
|
ocr.sh
|
||||||
|
↓
|
||||||
|
Google Cloud Vision raw text
|
||||||
|
↓
|
||||||
|
parse_ocr.py
|
||||||
|
↓
|
||||||
|
name validation
|
||||||
|
↓
|
||||||
|
SQLite update
|
||||||
|
```
|
||||||
|
|
||||||
|
## Raw OCR script
|
||||||
|
|
||||||
|
Run OCR directly against an image URL:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
/storage/disk1/WCX/scripts/ocr.sh \
|
||||||
"UPDATE movie
|
"https://example.com/thumbnail.jpg"
|
||||||
SET duration_seconds = 1,
|
```
|
||||||
updated = NULL
|
|
||||||
|
Example output:
|
||||||
|
|
||||||
|
```text
|
||||||
|
SUSANA MELO
|
||||||
|
PORTUGUESE
|
||||||
|
WOODMAN CASTINGX.COM
|
||||||
|
Budapest (Hungary), March 9, 2014
|
||||||
|
```
|
||||||
|
|
||||||
|
The raw OCR output is not considered a stable data format and must be parsed and validated before metadata is stored.
|
||||||
|
|
||||||
|
## Expected OCR structure
|
||||||
|
|
||||||
|
The parser currently expects exactly four non-empty lines:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Line 1: Performer name
|
||||||
|
Line 2: Nationality
|
||||||
|
Line 3: Branding text, ignored
|
||||||
|
Line 4: Shoot location and shoot date
|
||||||
|
```
|
||||||
|
|
||||||
|
Example final line:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Budapest (Hungary), March 9, 2014
|
||||||
|
```
|
||||||
|
|
||||||
|
The line is split on the first comma:
|
||||||
|
|
||||||
|
```text
|
||||||
|
shoot_location = Budapest (Hungary)
|
||||||
|
shoot_date = March 9, 2014
|
||||||
|
```
|
||||||
|
|
||||||
|
The date is normalized before storage:
|
||||||
|
|
||||||
|
```text
|
||||||
|
March 9, 2014 → 2014-03-09
|
||||||
|
```
|
||||||
|
|
||||||
|
## Testing the OCR parser
|
||||||
|
|
||||||
|
Raw OCR output can be piped directly into the parser:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/storage/disk1/WCX/scripts/ocr.sh \
|
||||||
|
"https://example.com/thumbnail.jpg" \
|
||||||
|
| /storage/disk1/WCX/scripts/parse_ocr.py
|
||||||
|
```
|
||||||
|
|
||||||
|
Example result:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ocr_name=SUSANA MELO
|
||||||
|
nationality=PORTUGUESE
|
||||||
|
shoot_location=Budapest (Hungary)
|
||||||
|
shoot_date=2014-03-09
|
||||||
|
```
|
||||||
|
|
||||||
|
## Processing OCR for a movie
|
||||||
|
|
||||||
|
Run the complete OCR workflow using a movie ID:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
/storage/disk1/WCX/scripts/check_ocr.py susana-melo_6707
|
||||||
|
```
|
||||||
|
|
||||||
|
Example output:
|
||||||
|
|
||||||
|
```text
|
||||||
|
id=susana-melo_6707
|
||||||
|
database_name=Susana Melo
|
||||||
|
ocr_name=SUSANA MELO
|
||||||
|
name_match=yes
|
||||||
|
nationality=PORTUGUESE
|
||||||
|
shoot_location=Budapest (Hungary)
|
||||||
|
shoot_date=2014-03-09
|
||||||
|
ocr_status=completed
|
||||||
|
```
|
||||||
|
|
||||||
|
The name comparison is case-insensitive and ignores repeated whitespace.
|
||||||
|
|
||||||
|
For example:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Susana Melo
|
||||||
|
SUSANA MELO
|
||||||
|
```
|
||||||
|
|
||||||
|
are considered equal.
|
||||||
|
|
||||||
|
## Successful OCR processing
|
||||||
|
|
||||||
|
When OCR execution and parsing succeed and the OCR name matches the database name, the following fields are updated:
|
||||||
|
|
||||||
|
```text
|
||||||
|
nationality
|
||||||
|
shoot_location
|
||||||
|
shoot_date
|
||||||
|
ocr_raw_text
|
||||||
|
ocr_status = completed
|
||||||
|
ocr_error = NULL
|
||||||
|
ocr_processed_at
|
||||||
|
modified_at
|
||||||
|
```
|
||||||
|
|
||||||
|
## OCR failure handling
|
||||||
|
|
||||||
|
If the Google Vision request or `ocr.sh` execution fails:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ocr_status = failed
|
||||||
|
```
|
||||||
|
|
||||||
|
The error is stored in:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ocr_error
|
||||||
|
```
|
||||||
|
|
||||||
|
## Manual review handling
|
||||||
|
|
||||||
|
The movie is marked for manual review when:
|
||||||
|
|
||||||
|
* the OCR output cannot be parsed
|
||||||
|
* the expected fields are missing
|
||||||
|
* the OCR name does not match the database name
|
||||||
|
* the OCR output has an unexpected structure
|
||||||
|
|
||||||
|
In this case:
|
||||||
|
|
||||||
|
```text
|
||||||
|
ocr_status = manual_review
|
||||||
|
```
|
||||||
|
|
||||||
|
The raw OCR text is retained in `ocr_raw_text`, but the parsed metadata fields are not updated.
|
||||||
|
|
||||||
|
This prevents uncertain OCR output from silently replacing valid metadata.
|
||||||
|
|
||||||
|
## Inspecting OCR results
|
||||||
|
|
||||||
|
Show the structured OCR result:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
|
"SELECT
|
||||||
|
id,
|
||||||
|
name,
|
||||||
|
nationality,
|
||||||
|
shoot_location,
|
||||||
|
shoot_date,
|
||||||
|
ocr_status,
|
||||||
|
ocr_error,
|
||||||
|
ocr_processed_at
|
||||||
|
FROM movie
|
||||||
WHERE id = 'susana-melo_6707';"
|
WHERE id = 'susana-melo_6707';"
|
||||||
```
|
```
|
||||||
|
|
||||||
Kör sedan importen igen:
|
Show the raw OCR text:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
/storage/disk1/WCX/scripts/import_site.py
|
sqlite3 /storage/disk1/WCX/database/wcx.db \
|
||||||
|
"SELECT ocr_raw_text
|
||||||
|
FROM movie
|
||||||
|
WHERE id = 'susana-melo_6707';"
|
||||||
```
|
```
|
||||||
|
|
||||||
Posten ska då uppdateras från JSON-filen och `duration_seconds` ska återställas till sitens värde.
|
List movies waiting for OCR:
|
||||||
|
|
||||||
## Återskapa databasen
|
```bash
|
||||||
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
|
"SELECT id, name, ocr_status
|
||||||
|
FROM movie
|
||||||
|
WHERE ocr_status = 'pending'
|
||||||
|
ORDER BY published DESC;"
|
||||||
|
```
|
||||||
|
|
||||||
Under utveckling kan databasen enkelt återskapas eftersom all automatisk site-data kan importeras igen.
|
List movies requiring manual review:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
|
"SELECT id, name, ocr_error
|
||||||
|
FROM movie
|
||||||
|
WHERE ocr_status = 'manual_review';"
|
||||||
|
```
|
||||||
|
|
||||||
|
List failed OCR attempts:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sqlite3 -header -column /storage/disk1/WCX/database/wcx.db \
|
||||||
|
"SELECT id, name, ocr_error
|
||||||
|
FROM movie
|
||||||
|
WHERE ocr_status = 'failed';"
|
||||||
|
```
|
||||||
|
|
||||||
|
## Recreating the database
|
||||||
|
|
||||||
|
During development, the database can be recreated from the schema and site JSON:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
rm /storage/disk1/WCX/database/wcx.db
|
rm /storage/disk1/WCX/database/wcx.db
|
||||||
@ -262,31 +524,70 @@ sqlite3 /storage/disk1/WCX/database/wcx.db \
|
|||||||
/storage/disk1/WCX/scripts/import_site.py
|
/storage/disk1/WCX/scripts/import_site.py
|
||||||
```
|
```
|
||||||
|
|
||||||
Radera inte databasen på detta sätt när den senare innehåller manuellt registrerad metadata utan att först ta en backup.
|
Do not recreate the database this way after it contains manually maintained metadata unless a backup has been created first.
|
||||||
|
|
||||||
## Versionshantering
|
OCR data must also be recreated if the database is deleted.
|
||||||
|
|
||||||
Källkod och dokumentation ska versionshanteras i Git.
|
## Git and generated data
|
||||||
|
|
||||||
Databasen och importerade datafiler bör normalt inte checkas in.
|
The Git repository contains source code, schema, and documentation.
|
||||||
|
|
||||||
Föreslagen `.gitignore`:
|
The SQLite database and scraper output are runtime data and are not version-controlled.
|
||||||
|
|
||||||
|
Current `.gitignore` rules:
|
||||||
|
|
||||||
```gitignore
|
```gitignore
|
||||||
|
# SQLite database files
|
||||||
database/*.db
|
database/*.db
|
||||||
database/*.db-shm
|
database/*.db-shm
|
||||||
database/*.db-wal
|
database/*.db-wal
|
||||||
|
|
||||||
|
# Imported/generated data
|
||||||
import/*.json
|
import/*.json
|
||||||
|
|
||||||
|
# Python cache
|
||||||
__pycache__/
|
__pycache__/
|
||||||
*.pyc
|
*.pyc
|
||||||
```
|
```
|
||||||
|
|
||||||
## Nästa steg
|
Files that should be version-controlled include:
|
||||||
|
|
||||||
Planerade kommande delar:
|
```text
|
||||||
|
README.md
|
||||||
|
.gitignore
|
||||||
|
scripts/schema.sql
|
||||||
|
scripts/import_site.py
|
||||||
|
scripts/ocr.sh
|
||||||
|
scripts/parse_ocr.py
|
||||||
|
scripts/check_ocr.py
|
||||||
|
```
|
||||||
|
|
||||||
* OCR-behandling av thumbnails för nya filmer
|
Files that should not be version-controlled include:
|
||||||
* manuell redigering av metadata
|
|
||||||
* export av komplett index
|
```text
|
||||||
* integration med Emby
|
database/wcx.db
|
||||||
* schemalagd körning
|
database/wcx.db-shm
|
||||||
|
database/wcx.db-wal
|
||||||
|
import/wcx_site_index.json
|
||||||
|
```
|
||||||
|
|
||||||
|
## Current limitations
|
||||||
|
|
||||||
|
The current OCR parser assumes a four-line OCR result.
|
||||||
|
|
||||||
|
Google Cloud Vision does not guarantee this exact structure. Unexpected output is therefore stored for manual review rather than automatically accepted.
|
||||||
|
|
||||||
|
OCR processing is currently started manually for one movie ID at a time.
|
||||||
|
|
||||||
|
## Planned next steps
|
||||||
|
|
||||||
|
Potential next steps include:
|
||||||
|
|
||||||
|
* automatically process movies with `ocr_status = pending`
|
||||||
|
* add retry support for failed OCR requests
|
||||||
|
* provide commands for manually approving or correcting OCR results
|
||||||
|
* add database backup handling
|
||||||
|
* export the complete index to JSON or CSV
|
||||||
|
* add Emby-compatible metadata export
|
||||||
|
* schedule the scraper and import process
|
||||||
|
* create an orchestration script for the complete automated workflow
|
||||||
|
|||||||
Reference in New Issue
Block a user