diff --git a/.gitignore b/.gitignore index 33e736e5..d6a817a1 100644 --- a/.gitignore +++ b/.gitignore @@ -58,3 +58,6 @@ skills-lock.json #local backups *.glossa-backup + +# Istruzioni personali per agenti, solo locali +CLAUDE.local.md diff --git a/CLAUDE.md b/CLAUDE.md index cc7d43c3..1b306832 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,103 +1,130 @@ -# Glossa — Istruzioni per lo sviluppo +# Glossa — guidance for AI coding agents -## Stato del progetto +Glossa is a desktop application for scholars working on historical texts: +discover and collect sources (Library), transcribe them with optional OCR/HTR +assistance (Transcriptions), translate them through configurable LLM/DeepL +pipelines with glossaries and phrase memory (Translations), and export the +results. It is a private beta: the 2.x version numbers come from release-automation +tests and do not indicate completeness. Product order of work: Library → +Transcriptions → Translations → Export (see `docs-dev/ROADMAP_2_0.md`). -Beta privata in sviluppo, senza una base di utenti esterni. La numerazione 2.x deriva da prove di rilascio automatico e non indica completezza. Obiettivo: completare Biblioteca → Trascrizioni → Traduzioni → Export, preservando la modalità documento/editoriale. Ordine in docs-dev/ROADMAP_2_0.md. Scriptoria resta riferimento tecnico per fonti, deposito, lavori, trascrizione ed export. UI sandbox tocca solo regressioni bloccanti. +The interface, in-app guide and developer documentation are written in +Italian; public documentation is published in Italian and English. ## Stack -- **Frontend**: React 19, TypeScript, Tailwind CSS v4, Zustand, Vite -- **Backend**: Rust (Tauri v2), SQLite via SQLx, reqwest -- **Test**: Vitest + Testing Library (Frontend), tokio-test + wiremock (Backend) - -## Principi Fondamentali - -- **Semplicità**: Codice minimo. No feature speculative future. -- **UI (vincolo utente)**: Comandi visivi solo `IconButton` neutri, icona + tooltip hover. No pill o pulsanti testuali/colorati. Verde solo per tab, selettori e stati attivi, salvo richiesta esplicita utente. -- **Leggibilità**: Nomi descrittivi. Commenti solo per logiche non ovvie, vincoli nascosti, workaround. -- **Immutabilità**: No mutare oggetti esistenti (preferisci `let` a `let mut`, usa spread operator e metodi funzionali in JS). -- **File**: Max 400-800 righe. Organizza per dominio/feature, non per tipo file. -- **TypeScript**: Tipi espliciti (mai `any`), costanti nominate. Gestione esplicita `null`/`undefined`. Validazione rigorosa input esterni. -- **Rust**: No `.unwrap()` in produzione. Uso sistematico `?` e `thiserror`. Formattazione/linting rigorosi (`cargo fmt`, zero warning `clippy`). Evita `clone()` inutili. -- **Architettura**: Handler backend snelli (logica in moduli dominio). Frontend con hook custom; Zustand solo per stato globale reale. - -## Invarianti della Pipeline - -- **Prefix Caching (CRITICO)**: Ordine blocchi system prompt (`static → blob → stage-instructions`) **mai cambia**. Inversione spezza cache provider, moltiplica costi. - -## Documentazione e Stato - -- **Lettura selettiva**: Parti da `docs-dev/README.md` e apri solo i documenti - pertinenti al task; non leggere tutta `docs-dev` per default. -- **Architettura**: Aggiorna `docs-dev/ARCHITECTURE.md` per modifiche flussi, comandi Tauri, schemi DB, store Zustand. -- **UI**: Consulta `docs-dev/UI_DESIGN_SYSTEM.md` prima di ogni modifica visiva. -- **Avanzamento**: Leggi `STATO_SESSIONE_2.0.md` inizio sessione, aggiorna obbligatorio fine task/feature. - -### Regola di documentazione (OBBLIGATORIA, non negoziabile) - -Ogni funzionalità nuova, rimossa o cambiata nel comportamento visibile va documentata **in tre posti nello stesso task**, prima di considerarlo finito: - -1. **Guida in-app** — `src/components/help/HelpGuide.tsx` più le stringhe `help.*` in `src/i18n/it.json` **e** `src/i18n/en.json`. Se serve una sezione nuova, aggiungila all'elenco di navigazione, al selettore di rendering e al tipo `HelpSection` in `src/stores/uiStore.ts`. -2. **Documentazione pubblica VitePress** — `docs/` (IT) **e** `docs/en/` (EN), pubblicata su GitHub Pages. Pagina nuova ⇒ voce in entrambe le barre laterali di `docs/.vitepress/config.ts`. Verifica con `npx vitepress build docs`, che fallisce sui collegamenti morti. -3. **Documentazione di sviluppo** — il documento pertinente secondo la tabella in `docs-dev/README.md`: architettura per flussi, comandi e schema; design system per regole visive; roadmap per il lavoro che resta. - -Nessuna delle tre è opzionale né rimandabile a un task successivo: una funzione non documentata è una funzione che nessuno sa usare e che verrà riprogettata da capo fra un mese. Vale anche per le correzioni che cambiano cosa l'utente vede, non solo per le funzioni nuove. - -Le tre superfici hanno destinatari diversi e non si copiano fra loro: la guida in-app spiega cosa fare mentre l'utente è nell'applicazione; la documentazione pubblica spiega il percorso completo e i limiti attuali; `docs-dev` registra invarianti e decisioni tecniche. -Descrivi sempre il comportamento presente e i limiti veri, mai la cronologia dello sviluppo. - -## Comunicazione con l'utente (CRITICO) - -Niki non scrive codice, non riconosce nomi tecnici. Spiegazioni utente: - -- **Mai** citare nomi file, funzioni, variabili, hook, componenti -- **Sempre** descrivere comportamenti visibili: cosa utente vede, clicca, ottiene -- **Giusto**: "la finestra della Libreria ora mostra il nome del workspace nel titolo" -- **Sbagliato**: "LibraryPanel usa panelTitle derivato da activeWorkspace?.name" - -## Git e Test - -- **Git**: Aggiorna sempre `main` prima creare branch (`git checkout main && git pull origin main && git checkout -b nome-branch`). -- **Test**: Approccio TDD. Copertura minima 80%. Nomi test descrittivi su comportamento atteso. Mai sopprimere errori in silenzio. Su task lunghi fare i test solo alla fine. - ---- - -## Strumenti e Ottimizzazione Token - -### Repomix (Esplorazione Iniziale) - -Prima di analizzare porzioni codebase estese o poco conosciute, usa **repomix** (`skill repomix-commands:pack-local`). - -- **Scopo**: Vista compatta e indicizzata intero progetto in un'unica operazione, azzera catene esplorative costose filesystem, risparmia token. -- **Misura**: usa include mirati al dominio da modificare; non generare pack completi quando bastano pochi file noti. - -### RTK (Rust Token Killer) - Filtro Output CLI - -Per prevenire esaurimento finestra contesto, **ogni comando terminale deve iniziare con `rtk`**. - -`rtk` intercetta output, filtra verbosità, restituisce formati iper-compatti, risparmia 60-90% token. - -- **Uso corretto**: `rtk cargo test`, `rtk grep pattern`, `rtk read file.ts` -- **Catene**: Anche con `&&`, applica ogni step: `rtk git add . && rtk git commit -m "msg" && rtk git push` - -### Economia di tempo e token - -- **Comunicazione**: usa il skill `caveman` nelle attività operative, salvo casi in cui la chiarezza o la sicurezza richiedano prosa normale. -- **Verifica proporzionata**: esegui soltanto test direttamente pertinenti ai file o contratti modificati — `npx vitest run `, `cargo test `. Suite complete **una sola volta, prima del commit**, e solo se il commit tocca più aree; mai fra una modifica e l'altra, mai per confermare qualcosa che il compilatore ha già detto. Vale anche per `clippy` e `tsc`: si lanciano quando servono, non a ogni passo. -- **Build**: non eseguire build dell'app, build Tauri, build della documentazione, E2E o installazioni di dipendenze salvo richiesta esplicita dell'utente o necessità indispensabile per diagnosticare un errore. -- **Esplorazione**: preferisci `rtk rg`, letture mirate e repomix compresso; evita scansioni o output completi non necessari al task. - -### MCP Tools: code-review-graph - -⚠️ **REGOLA DI INGAGGIO (OTTIMIZZAZIONE TOKEN):** -Strumenti grafo consumano molti token per esecuzione, aumentano latenza. Uso NON default. - -1. **Usa strumenti MCP (es. `query_graph`, `get_impact_radius`) SOLO se:** - - Utente chiede analisi architetturale o report impatto cross-file. - - Devi mappare dipendenze complesse per refactoring strutturale profondo. - - Stai esplorando parte completamente sconosciuta e interconnessa progetto. - -2. **Usa comandi CLI standard (`rtk grep`, `rtk read`, `rtk ls`) o repomix come DEFAULT per:** - - Fix locali, aggiunta componenti isolati o logica circoscritta. - - Interventi dentro file già noti. - - Lettura firme funzioni o ispezione file configurazione. +- **Frontend:** React 19, TypeScript, Tailwind CSS v4, Zustand, Radix UI, Vite. +- **Backend:** Rust, Tauri v2. SQLite through `@tauri-apps/plugin-sql` from the + frontend (`src/services/dbService.ts`) and through `rusqlite` + `sqlite-vec` in + the backend for the text corpus and embeddings (`src-tauri/src/vector/`). +- **Tests:** Vitest + Testing Library (frontend), `cargo test` with tokio-test and + wiremock (backend), Playwright with a Tauri mock (`e2e/`). +- **Docs:** VitePress (`docs/` Italian, `docs/en/` English). + +## Repository layout + +| Path | Content | +|---|---| +| `src/components/` | UI by domain (`library`, `transcription`, `translation`, `pipeline`, `document`, `settings`, …); shared primitives in `src/components/ui/` | +| `src/stores/`, `src/hooks/`, `src/services/` | Zustand stores (global state only), domain hooks, data and Tauri-command services | +| `src/i18n/it.json`, `src/i18n/en.json` | All UI strings, including the in-app guide (`help.*`) | +| `src/languages/` | Bundled ISO 639-3 and Glottolog language lists; regenerate with `npx tsx scripts/update-languages.ts` (in the app: Settings → Languages) | +| `src-tauri/src/` | Backend modules by domain (`llm`, `deepl`, `vector`, `federation`, `iiif`, `ocr`, `jobs`, …) | +| `src-tauri/migrations/` | SQLx migrations, fingerprinted in `src-tauri/migrations.lock` | +| `docs-dev/` | Developer documentation; start from `docs-dev/README.md` | + +## Commands + +```bash +npm run tauri:dev # run the desktop app in development +npm run lint:all # typecheck + ESLint +npm test # Vitest +npx vitest run # targeted frontend tests +cd src-tauri && cargo fmt && cargo clippy --all-targets -- -D warnings +cd src-tauri && cargo test # backend tests (includes the migration lock test) +npx vitepress build docs # public docs; fails on dead links +``` + +Run the checks relevant to what you changed; run full suites once, before +committing work that touches several areas. Do not run app builds, Tauri +builds, E2E or dependency installs unless asked or needed to diagnose a failure. + +## Engineering principles + +- **Simplicity:** minimal code, no speculative features or abstractions. +- **TypeScript:** explicit types, never `any`; explicit `null`/`undefined` + handling; validate external input (API responses, files, bundled data). +- **Rust:** no `.unwrap()` in production code; propagate errors with `?` and + `thiserror`; zero `clippy` warnings; avoid needless `clone()`. +- **Immutability:** prefer `let` over `let mut`; spread and functional methods in + TypeScript. +- **Size and structure:** files of 400–800 lines at most, organised by domain. + Backend handlers stay thin; logic lives in domain modules. Frontend logic in + custom hooks. +- **Comments:** only for non-obvious logic, hidden constraints and workarounds. +- **Testing:** descriptive test names stating the expected behaviour; never + silence errors. Target 80% coverage on new logic. + +## Invariants + +- **Prompt cache order (critical):** the system prompt blocks are always ordered + `static → blob → stage instructions`. Changing the order breaks provider prompt + caching and multiplies costs. The composed prompts are covered by a + byte-equivalence test (`src-tauri/src/llm/legacy_prompts_test.rs`). +- **Migrations:** an applied migration is never edited. Add a new migration and + its line in `migrations.lock`. Use migrations only for real schema changes, + never for one-off data fixes. Do not recreate tables (`DROP TABLE`) inside a + migration: SQLx runs it in a transaction where `PRAGMA foreign_keys=OFF` has no + effect, so cascades delete data. Pre-release consolidation of migrations is + done only on explicit request by the maintainer. +- **Text corpus:** text revisions are immutable; a correction (text or language) + creates a new revision. Embeddings always record provider, model, dimensions + and input profile; similarity search only compares compatible measures. +- **Languages:** a work's languages belong to the work (ISO 639-3 code, optional + Glottolog variety, free note), not to its pipelines. DeepL keeps its own + language pair in its phase options. + +## UI rules + +Read `docs-dev/UI_DESIGN_SYSTEM.md` before any visual change. + +- Commands are neutral icon-only `IconButton`s with a tooltip. No pills, no + coloured or text buttons outside dialogs. +- Green (accent) only for tabs, selectors and active states. +- Explanations go in hover hints, not permanent paragraphs. +- Reuse the shared primitives in `src/components/ui/` (`Dialog`, `SettingRow`, + `SearchPicker`, `ClickPopover`, `ChoiceDots`, …); no local variants. +- Every UI string exists in both `it.json` and `en.json`. + +## Documentation rule (mandatory) + +Every feature that is added, removed or changes visible behaviour is documented +in the same task, in three places: + +1. **In-app guide:** `src/components/help/HelpGuide.tsx` and the `help.*` strings + in both `src/i18n/it.json` and `src/i18n/en.json`. A new section also needs its + navigation entry, renderer case and the `HelpSection` type in + `src/stores/uiStore.ts`. +2. **Public docs:** `docs/` (Italian) and `docs/en/` (English). A new page needs an + entry in both sidebars of `docs/.vitepress/config.ts`. +3. **Developer docs:** the relevant file per `docs-dev/README.md` — architecture + for flows, commands and schema; design system for visual rules; roadmap for + remaining work. + +The three surfaces have different readers and are not copies of each other. +Always describe present behaviour and real limits, never development history. + +## Git + +- Never work on `main`. Branch from an up-to-date `main` + (`git checkout main && git pull origin main && git checkout -b `), + unless the maintainer names a different base. +- Conventional commits: `feat`, `fix`, `refactor`, `docs`, `test`, `chore`, + `perf`, `ci`, with an optional scope. +- Base and target branch of a pull request are chosen by the maintainer; + opening a PR never implies merging it. + +## References + +Scriptoria remains the technical reference for sources, storage, jobs, +transcription and export (#186, #446). diff --git a/docs-dev/ARCHITECTURE.md b/docs-dev/ARCHITECTURE.md index f9173b8c..eefe15a0 100644 --- a/docs-dev/ARCHITECTURE.md +++ b/docs-dev/ARCHITECTURE.md @@ -14,11 +14,72 @@ nuova, idempotente sui database che l'hanno già superata. una migrazione dichiarata è cambiata o sparita, sia quando ne compare una non dichiarata. Aggiornare il lucchetto è legittimo solo per aggiungere una riga. +## Contesto e anteprime della pipeline (#489) + +`PipelineConfig.workBrief` / Rust `work_brief` è l’unico contesto comune degli +LLM. Migrazione 0004 aggiunge la colonna; 0005 elimina Persona e override +lingua globali, senza conversioni semantiche o percorsi legacy. Le lingue non +stanno più nella pipeline: sono dell'opera (vedi «Lingue dell'opera» sotto). +Lettura, salvataggio e duplicazione includono la descrizione. Backup usa righe +e colonne dello schema corrente. Nome visibile «Contesto di traduzione» (campo `workBrief`, colonna `work_brief`): obbligatorio, mai vuoto — `DEFAULT_WORK_BRIEF` (inglese → italiano, come la vecchia coppia predefinita) nei default dello store e alla lettura di una riga vuota; l’editor non conferma un testo vuoto e il ripristino torna al predefinito. Template nel contesto `brief`; +contesti obsoleti non vengono riclassificati silenziosamente: la lettura +(`getPromptTemplates`) esclude le righe con contesto sconosciuto, le registra +nel log e ne restituisce i nomi in `skipped`, che lo store mostra in un avviso; +una riga sola non svuota più l’intero elenco. I tre template `persona` del +database di sviluppo sono stati riclassificati a mano in `brief` (correzione +dati una tantum, nessuna migrazione). + +Regole per fase: le frasi della memoria si aggiungono solo a traduzione e Refine (`receivesMemory` in `engine.ts`, stessa regola nell’anteprima); il glossario sta una volta nelle regole del blocco statico, senza promemoria nelle istruzioni; intestazione del contesto `Translation context:`. La prova di equivalenza confronta con una copia della composizione precedente aggiornata con gli stessi cambi di testo voluti. Testi di sistema (`src-tauri/src/llm/prompt_texts.rs`): ogni testo del programma attorno ai prompt dell’utente ha id, predefinito e segnaposto obbligatori (`SYSTEM_TEXTS`); `render(config, id, values)` usa la sostituzione della pipeline se completa, altrimenti il predefinito; sostituzione dei segnaposto in un solo passaggio. Comando `prompt_system_texts` per l’editor. Migrazione 0006 aggiunge `pipelines.prompt_composition` (JSON `{texts: {id: testo}, disabled: [id pezzo]}`, NULL = predefiniti) ↔ `PipelineConfig.promptComposition`; si copia nella duplicazione insieme a `coherence_prompt` (prima perso). I separatori fra pezzi restano nel codice, così i predefiniti producono gli stessi byte di prima. Categoria template `system`. Pezzi spenti: `promptComposition.disabled` con chiavi `fase:pezzo` (`translation`, `refine`, `format`, `audit`, `coherence`); `on()`/`when_on()` in `prompts.rs` li omettono; spenti i frammenti vicini, si omette anche l’id del frammento. Elenco dei facoltativi in `SWITCHABLE_PARTS` (`promptParts.ts`). Composizione a pezzi (`src-tauri/src/llm/composition.rs`): `compose_stage_prompts`, +`compose_judge_prompts`, `compose_coherence_prompts` producono blocchi di `PromptPart` +con id stabile; `into_structured()` li concatena (la richiesta vera), `preview_parts()` +li restituisce all’anteprima nei comandi `preview_*_prompt` (campo `parts`). Anteprima e +invio coincidono per costruzione; prova di equivalenza byte per byte con la copia della +composizione precedente (`legacy_prompts_test.rs`). L’anteprima delle opzioni +(`PromptPreviewTab` + catalogo `promptParts.ts`) mostra per fase tutti i pezzi possibili +in ordine, con tipo (fisso/tuo/dati/automatico), luogo di modifica e motivo di assenza. +L’impronta di ripresa (`pipelineFingerprint`) include modelli, prompt +delle fasi LLM attive, prompt dell’audit, descrizione e parametri DeepL. +La linguetta Anteprima dello Studio è stata rimossa: il modo «Frammento aperto» di `PromptPreviewTab` usa `useChunkPromptPreview` sul frammento selezionato (o il primo). L’hook offre anche `preview-audit` e +`preview-coherence`: stessi input dell’esecuzione manuale (traduzione attuale; per la +coerenza blocco dei frammenti vicini tradotti) tramite `preview_judge_prompt` e +`preview_coherence_prompt`; voci spente senza traduzione. + +Traduzione/refine, audit e coerenza ricevono ruolo neutro, descrizione opzionale +e istruzioni proprie; nessuna coppia implicita anche con descrizione vuota. +Format resta isolato. Ordine system cacheabile immutato: static → blob → +istruzioni della fase. Lingua report dalla UI, fallback English. Fingerprint +comprende descrizione normalizzata e opzioni DeepL. La rifinitura `brief` +preserva il contesto senza aggiungere ordini specifici delle fasi. + +DeepL usa solo `providerOptions.deepl.sourceLang/targetLang`; sorgente vuota +significa rilevamento automatico, destinazione obbligatoria. Input Tauri unico +`{text, deeplConfig}`. `build_translate_request` valida e compone il corpo sia +per HTTP sia per `preview_deepl_stage`, senza chiavi nella preview. Codici +dalle liste API; nessuna conversione euristica dei nomi lingua. Cambio coppia +scollega il glossario, cambio target azzera formality. Glossario richiede +sorgente esplicita e target. + +`preview_stage_prompt`, `preview_judge_prompt`, `preview_coherence_prompt` +riusano i costruttori dell’esecuzione. Opzioni: richiesta completa iniziale, +costruzione per blocchi selezionabile per fasi LLM, messaggi audit/coerenza e +corpo DeepL. Segnaposto espliciti per testo, blob, memoria e risultato +precedente; non sono richieste storiche. Risposte superate ignorate. +`PromptMessage` unifica carta tenue/verde, espansione e copia integrale nelle +opzioni e nel frammento. Log mantengono le richieste effettive. +Anteprima diretta audit/coerenza nel singolo frammento ancora da completare. + +Tutti i quattro ruoli restano nella configurazione; la modalità determina +`enabled`. Cambio modalità conserva prompt/modello/opzioni/profilo custom. +Sotto-tab inattive disabilitate. Editor pipeline con bozza locale e conferma; +applicazione modelli e rifinitura non salvano implicitamente. Coppia DeepL +sempre visibile in Generale, abilitata soltanto in DeepL. + ## Ricerca federata e Dashboard Le due viste vivono nella Dashboard (`DashboardArea`): `overview` e ricerca (`view: 'search'`). Si scelgono dalla barra a sinistra, -come voci sotto la Dashboard (`WorkspaceRailNext`), non da una fila di linguette +come voci sotto la Dashboard (`WorkspaceRailNext`, che in fondo porta anche il +menu generale `ShellNavFooter`, tolto dalla testata), non da una fila di linguette dentro la pagina; solo la vista corrente monta, così una ricerca nascosta non continua a leggere. Contratto di navigazione: la variante `dashboard` di `AppLocation` porta `view` e `searchId`; la variante `library` porta solo @@ -347,6 +408,10 @@ dentro. baseline si fissa e vale di nuovo la regola sopra — ogni cambiamento riceve un file di migrazione nuovo, la baseline non si tocca più. +Per il ramo Studio/corpus (PR #488), istruzione esplicita dell'utente: +conservare 0001 applicata e usare 0003 incrementale. L'eventuale consolidamento +prima del merge viene eseguito dall'utente, non dall'agente. + ## Modello di prodotto Biblioteca, Trascrizioni, Traduzioni e Analisi sono cataloghi globali. Un @@ -381,8 +446,8 @@ approvate; traduzioni attive la cui origine è una copia dell'opera o una sua trascrizione. Vista (`libraryView`: elenco, copertine, tabella) e raggruppamento -(`libraryGrouping`, `utils/libraryGrouping.ts`: secolo, autore, biblioteca, -raccolta) sono preferenze persistite in `uiStore`. Il raggruppamento lavora +(`libraryGrouping`, `utils/libraryGrouping.ts` sopra `utils/catalogGrouping.ts`: +secolo, autore, biblioteca, raccolta) sono preferenze persistite in `uiStore`. Il raggruppamento lavora sull'elenco già filtrato e ordinato; un'opera in più raccolte compare in ogni gruppo, e la scelta per intervallo segue l'ordine visibile, gruppi compresi. @@ -466,13 +531,24 @@ modo immutabile. Stato confinato a un componente resta locale. chiesta e scarta le risposte più lente: aprendo A e poi B, la risposta di A sostituiva il dettaglio di B e l'attesa non finiva più. -**Impostazioni.** `uiStore.settingsTab` ha una sola linguetta per la Biblioteca -(`library`), che al suo interno si divide in tre sotto-linguette — ritmi di -rete, biblioteche, immagini — tenute in stato locale. Le vecchie `download` e -`libraries` non esistono più; la linguetta disattivata «in arrivo» resta solo per -le Trascrizioni. La bozza di un profilo di rete vive nella finestra e non nella -scheda, perché la scheda si smonta cambiando linguetta, e il profilo in modifica -si ritrova dalla bozza al rientro. +**Tema.** `useIsDarkTheme` è l'unico calcolo di «tema scuro» (scelta +dell'utente o sistema, seguito anche mentre l'app è aperta): lo usano +`ThemeSync` (classe `dark` su ``), l'accento e le evidenziazioni +applicati a runtime. + +**Impostazioni.** Valgono per tutta l'app; quelle di un workspace stanno in +`WorkspaceSettingsModal`. `uiStore.settingsTab` (non persistito) ha otto +linguette, nella fila comune `TabStrip`: `appearance`, `library`, +`transcriptions`, `translations`, `models`, `languages`, `data`, `jobs`. Le +schede con argomenti diversi usano `SettingsSubTabs` (sotto-linguette in stato +locale): Aspetto (interfaccia, documento, evidenziazioni), Biblioteca (ritmi di +rete, biblioteche, immagini), Modelli (un provider per linguetta, provider +personalizzato, prezzi), Dati (cartelle e deposito, cache di rete, backup). Ogni +scheda legge da sé i propri store; le scelte esclusive sono `SettingChoiceRow` +(`ChoiceDots` con il nome della scelta) o `Select`. La bozza di un profilo di +rete vive nella finestra e non nella scheda, perché la scheda si smonta +cambiando linguetta, e il profilo in modifica si ritrova dalla bozza al rientro. +Il controllo pre-avvio con problemi apre direttamente Modelli. ## Trascrizioni @@ -486,6 +562,15 @@ deduplicate per impronta del contenuto (`content_hash`), un segmento senza di stato propria. Il testo si salva dopo 30 secondi senza modifiche, e subito lasciando la pagina. +Lo stato del salvataggio sta nella barra di stato, come per le traduzioni: +`useSegmentEditor` lo pubblica in `transcriptionStore.textSave` (stato, ora +dell'ultimo salvataggio riuscito, messaggio d'errore; `null` a Studio chiuso) e +`AppStatusBar` lo mostra con lo stesso `SaveIndicator` del progetto («da +salvare» = `dirty`). Il salvataggio manuale (dischetto nella testata del foglio, +rosso dopo un errore; Ctrl/⌘+S ascoltato sulla +finestra mentre lo Studio è montato, anche dentro il foglio) passa dalla stessa +coda di `save` e scrive una revisione senza nome; è spento quando il testo è +già salvato o il foglio è in sola lettura. Ogni caricamento è legato all'indice di pagina che lo ha richiesto: una risposta tardiva non può sostituire testo e storico della pagina ora aperta. Le versioni consolidate sono le stesse revisioni con `consolidated_name` valorizzato: nessuna @@ -520,6 +605,24 @@ digitalizzazione) resta su un solo blocco di testo, in posizione 0, `source_page_id` sempre `NULL` — lo stesso codice, solo che la pagina non cambia mai. +**Catalogo** (`TranscriptionsCatalogArea`): stesso modello della Biblioteca e +stessi pezzi (`ShelfItem`, `CatalogSearchField`, `CatalogViewSwitch`, +`CompletionBar`, `CommandBar`, `ui/catalogStyles.ts`, +`utils/catalogGrouping.ts`). `listTranscriptionCatalog` +(`services/transcriptionCatalogService.ts`) legge in una query tutte le +trascrizioni attive e archiviate di tutti i workspace, con pagine scritte +(ultima revisione con testo non vuoto), pagine verificate e ultima revisione; +l'opera arriva da `listLibraryCatalog` tramite `source_versions.source_id`, +quindi con le correzioni manuali e il totale di pagine della Biblioteca. +Scaffali, filtri rapidi (workspace, biblioteca, secolo), ordine e +raggruppamento vivono in `utils/transcriptionCatalogFilters.ts`; vista e +raggruppamento sono preferenze persistite (`transcriptionsView`, +`transcriptionsGrouping` in `uiStore`), il filtro workspace segue +l'indirizzo come in Biblioteca. «Verificata» vuol dire tutte le pagine +dell'opera verificate, o tutte quelle scritte se il totale non si conosce. +Rinomina (`renameDocument`) e archiviazione (`setDocumentStatus` con +`archived`) sono comandi di riga; il cestino resta `trashed`. + **Studio di trascrizione** (`TranscriptionsCatalogArea` + `TranscriptionStudio`, #388): stessa convenzione della scheda opera in Biblioteca, non quella dello Studio di traduzione — `AppLocation` porta `{ area: 'transcriptions', @@ -531,16 +634,32 @@ copia principale e dell'eventuale copia alternativa vive in `useTranscriptionSources`; storico e metadati vivono in `TranscriptionInspector`, separati dallo stato di salvataggio del testo. +**Struttura dello Studio.** `TranscriptionStudio` compone soltanto: lo stato +del testo per pagina (segmento, revisioni, bozza, catena dei salvataggi, +debounce, cambio pagina, salvataggio all'uscita) vive in `useSegmentEditor`; +verifica, ripristino, nomi, eliminazione e pulizia dello storico in +`useRevisionActions`; l'OCR in `useStudioOcr`; aggancio fra visore e testo, +richiesta di salto e cambio di copia in `useViewerSync`; larghezze e collasso +della colonna in `useInspectorLayout` (`INSPECTOR_WIDTH`, condivise con la +scheda opera); Ctrl/⌘+S in `useSaveShortcut`. Le parti visive sono +`StudioPageHeader`, `StudioViewerPane`, `StudioTextHeader`. I comandi della +copia (cambio immagini/PDF, sgancio) arrivano alla barra del visore come +elenco `ViewerCommand`, non come elementi già disegnati: sotto i 560 px la +barra (`useNarrowWidth`, una sola misura per barra e menu) li sposta nel menu +con i tre puntini insieme a «solo file locali» e «apri la pagina», così un +comando non è mai in due posti o in nessuno. `PagePendingOverlay`, mentre +copre la pagina, mette `inert` e `aria-busy` sugli elementi accanto nel suo +contenitore e li toglie quando sparisce. + **Intestazione**, quando il documento è legato a un'opera: stessa riga della scheda opera in Biblioteca (icona, titolo e autore dell'opera, uscita verso la biblioteca) — non il titolo scelto per la trascrizione, che identifica il documento nel catalogo e nel breadcrumb ma non qui, per non mostrare due titoli nella stessa schermata. Letta una volta per opera (`getLibrarySourceDetail` + `listIIIFProviders`, tenuti in `bookInfo`), non a -ogni cambio pagina. Il menu a tre puntini è **volutamente più povero** di -quello della scheda opera: solo "Rimuovi trascrizione", perché scaricare, -verificare, archiviare sono azioni sull'opera, non sul suo studio di -trascrizione — vivono già nella scheda opera. Un documento senza opera +ogni cambio pagina. L'unico comando della riga è il cestino della trascrizione: +scaricare, verificare, archiviare sono azioni sull'opera, non sul suo studio +di trascrizione — vivono già nella scheda opera. Un documento senza opera collegata mostra il proprio titolo, come prima. **Visore a sinistra** (#221, parte zoom/pan e cambio fonte — filtri visuali e @@ -832,6 +951,254 @@ lavoro sul testo di riferimento è stato rimosso perché nessuno lo calcolava. ## Pipeline di traduzione +**Catalogo delle Traduzioni** (`TranslationsArea`): stesso modello e stessi +pezzi del catalogo delle Trascrizioni (scaffali, ricerca, filtri rapidi, tre +viste, `CommandBar`, `CompletionBar`, `RenameField` comune in `ui/`). +`listTranslationCatalog` (`services/translationCatalogService.ts`) legge in una +query tutti i progetti di tutti i workspace con lingue, `updated_at` e i +conteggi dei frammenti della **prima pipeline** (quella che `openProject` +apre): totale, tradotti (`chunk_status = 'completed'`), verificati +(`translation_locked = 1`). Scaffali (Tutte, Recenti, Da iniziare, In corso, +Verificate), filtri rapidi (workspace, coppia di lingue), ordine e +raggruppamento vivono in `utils/translationCatalogFilters.ts`; vista e +raggruppamento sono preferenze persistite (`translationsView`, +`translationsGrouping` in `uiStore`), il filtro workspace segue l'indirizzo. +Comandi di riga: rinomina (`renameProject`) ed elimina (`removeProject`, che +cancella davvero: i progetti non hanno archivio). Nessun legame con opera o +trascrizione: `translation_origins` resta non scritta fino alla strada «da una +trascrizione». + +Creazione «da zero» (`CreateProjectDialog`): nome, workspace e file +facoltativo. Il file si legge alla scelta con `importTextFile` (errori mappati +da `importErrorMessageKey`, mostrati nella finestra; nulla si crea). Dopo +`createAndOpen` il file va in `uiStore.pendingImportFile` (non persistito): +l'editor montato lo consuma con `startImport`, la stessa via del comando di +import, e apre `ImportPreviewDialog`. Chiudendo l'anteprima il progetto resta +vuoto. Il libro di origine si sceglie con `SearchPicker` (titolo, copia sotto), +mai con un `Select` che si allarga al titolo più lungo. + +**Lingue dell'opera.** Le lingue sono dell'opera (`projects`), valgono per +tutte le sue pipeline; `pipelines` non ha più colonne di lingua (migrazione +0007, che converte i vecchi nomi inglesi in codici e aggiunge +`*_language_variety` e `*_language_note`). Ogni lato è un `LanguageChoice` +(`code` ISO 639-3 o null, `variety` Glottocode o null, `note`); `WorkLanguages` +vive in `pipelineStore.workLanguages`, caricato da `getProjectSource` e salvato +da `saveProjectSource`, `createProject` e `projectStore.updateWorkLanguages` +(`saveWorkLanguages`). Colonne vuote = non indicata. Elenco incluso in +`src/languages/data` (ISO 639-3 dal registro SIL, varietà Glottolog di livello +«dialect» collegate alla loro lingua ISO, nomi italiani da CLDR), rigenerato da +`npx tsx scripts/update-languages.ts`. Costruzione e unione stanno in +`languages/build.ts`, condiviso con «Aggiorna elenco lingue» (Impostazioni → +Lingue): il backend scarica le due fonti (`languages_fetch_sources`, fuori dai +limiti CORS), il frontend costruisce l'elenco, i codici spariti restano con il +segno «ritirato» (nominati, non più proposti) e il risultato si salva in +`/languages/` (`languages_save`, scrittura via file temporaneo). +`loadLanguageCatalog` preferisce l'elenco salvato (`languages_read_saved`) e +ricade su quello incluso; caricato a richiesta come testo grezzo e validato +(`languages/catalog.ts`, `useLanguageCatalog`); `catalog.info` dice origine, +date, conteggi e ritirati. Interfaccia unica +`WorkLanguagesFields` in `ImportPreviewDialog` (colonna sinistra) e nella +finestra di `WorkLanguagesControl` (salva solo con Conferma; partenza vuota +proposta dalla lingua del libro con `matchLanguage`). Dopo ogni Conferma +`countProjectPhraseRelabels` e, se l'utente accetta, `relabelProjectPhrases`. +Estrazione delle coppie: `describeLanguageForModel` (nome inglese, varietà, +nota). Catalogo e Memorie mostrano i nomi con `useLanguageLabel`. Fatti di +provenienza: codici dell'opera. + +**Studio di traduzione** (`components/translation/TranslationStudio`): si apre +quando `projectStore.currentProjectId` è valorizzato, **dentro** +`WorkspaceShellNext` come ogni altra area — la barra principale +(`WorkspaceRailNext`) resta. Con un progetto aperto la barra marca Traduzioni +come area attiva; ogni voce chiude il progetto (`closeProject`) prima di +`navigate`, e tutte si spengono mentre `chunksStore.isProcessing`: chiudere il +progetto sotto la pipeline svuoterebbe i frammenti a lavoro in corso. Il +ritorno al catalogo (`leaveTranslation` in `App`) fa lo stesso e porta a +`{ area: 'translations' }`. Ogni uscita (ritorno, voci della barra, percorso +nell'`Header`, cambio di workspace, creazione di un workspace, +`openProjectInWorkspace`) passa da `projectStore.leaveProject`: aspetta un +salvataggio già in volo, confronta l'istantanea corrente con `trackedSnapshot` +e salva se diverse; se il salvataggio fallisce restituisce `false` e la +traduzione resta aperta con l'errore. `closeProject` resta la chiusura secca +(dopo l'eliminazione, o dove non c'è niente da salvare). Limite: la chiusura +della finestra non salva. + +Salvataggio: autosalvataggio dell'intero progetto (`useProjectAutosave`, 1,2 s, +fermo durante la pipeline). Lo stato lo mostra `SaveIndicator` nella barra di +stato (suggerimento: ora dell'ultimo salvataggio e, dopo un errore, +`lastSaveError`; la barra è una regione `aria-live`, quindi si annuncia solo +l'errore). Il comando è `ProjectSaveButton` nella testata del foglio della +traduzione, o dell'originale con `paneFocus === 'source'`: spento senza +modifiche e durante la pipeline (motivo nel suggerimento), `danger` con +«Riprova» dopo un errore. Ctrl/⌘+S (`useKeyboardShortcuts`) a progetto aperto +vale anche dentro i campi e non mostra l'avviso di riuscita; senza progetto +salva solo le risorse linguistiche, fuori dai campi, come prima. + +Memoria di frasi: `vec_save_locked_phrases` **aggiunge** e +basta (niente più cancellazione delle coppie del frammento); le lingue delle +revisioni le legge dal progetto (`text_languages::project_languages`, `und` se +non indicata), non dal chiamante. Una coppia si +toglie con `vec_delete_phrase_memory`. `vec_list_phrase_memory(workspaceId?, +chunkId?)`: senza workspace tutte le frasi. `vec_search_phrase_memory` ha +`allWorkspaces` e `sourceLanguage`/`targetLanguage` (lo Studio passa solo la +lingua di arrivo dell'opera; nessun filtro se non indicata), e restituisce la +provenienza (workspace di casa = quello della traduzione o dell'importazione, +`NULL` = senza workspace; `project_id`, `chunk_id`), mostrata da +`PhraseProvenance` con `usePhraseProvenanceLookup` (due letture in tutto). +`workspaces.memory_search_all_workspaces` (migrazione 0002) è il campo del +globo nei Riferimenti (`ReferencesTab`, `updateActiveWorkspace`), non più nelle +impostazioni workspace. `ReferencesTab` ordina con `orderByCircle` +(`utils/memoryCircles.ts`): documento corrente, poi workspace, poi altrove, +dentro ogni cerchio per somiglianza; ogni riga (`ReferenceMatchRow`) porta +l'etichetta del cerchio e le lingue delle revisioni (`vec_search_phrase_memory` +restituisce `source_language`/`target_language`). La soglia ha passi di 0,01 +con − e +. +Cambiare workspace, modello di misura, lingua di arrivo dell'opera o ambito +invalida i riferimenti selezionati anche a ricerca automatica spenta; con +ricerca automatica attiva ne avvia una nuova. + +Risorse linguistiche: `LibraryPanel` usa `TabStrip`; Modelli aggiunge ricerca, +filtro OCR e `PromptTemplateForm` per creazione/modifica in posto tramite +`updateTemplate`. Un duplicato nome/ambito/flusso mantiene il modulo aperto +senza sovrascrivere un modello diverso. I modelli restano globali. La modifica +si apre nella propria voce; ricerca e filtri restano disabilitati finché la +bozza è aperta. Anteprima breve e lettura integrale non alterano il prompt. +`DictionaryEntryEditor` legge le coppie come testo; il più inserisce in cima +con focus esplicito. Modificare una voce aggiorna ancora la bozza nello store: +la spunta chiude l'editor, solo il dischetto persiste il dizionario. + +Memorie legge `listPhraseMemoryEntries(null)` per «tutti» e «senza workspace», +senza dipendere dall’esistenza di workspace; filtro e ricerca restringono +anche l’esportazione. La lista affianca i testi; i dettagli di una sola voce +mostrano provenienza compatta, misure e tag. Apertura dei dettagli è stato locale, +non una modifica ai dati. `measuringId` distingue il calcolo embedding dalle +altre operazioni impegnate: accende la rotellina del solo comando interessato, +dopo la conferma, e si azzera anche in caso di errore. Durante correzione/testo o tag sono bloccati ricerca, +filtri e chiusura dei dettagli, così la bozza non viene smontata. Correggere la sola resa conserva le misure; correggere l’originale crea una revisione e ricalcola atomicamente tutti i modelli già presenti. +La ricerca richiede misure con modello, dimensioni e profilo compatibili. +Non esistono embedding senza modello né riferimenti esatti inseriti fuori +dalla graduatoria semantica. Le coppie salvate restano nella Memoria del frammento. + +Dizionari: il filtro globale usa i **collegamenti** (`listGlossaries(ws)`), +con provenienza dal legame `is_origin`; il filtro non modifica l’ambito di +lettura/scrittura delle voci, che resta canonico (`null`) nelle risorse generali. +Nel workspace ospite la scrittura usa `saveGlossaryEntriesAsOverrides`; il +termine originale resta in sola lettura e nuove voci entrano nel canonico. +La vista aperta dichiara risorsa, destinatari delle modifiche e ambito delle +nuove voci. `entriesWorkspaceMap` conserva l’ambito della bozza per dizionario: +letture con ambito diverso ricaricano, le risposte superate si ignorano e +`saveAllDirty` conserva le correzioni locali. Chiudere dopo un salvataggio +fallito mantiene la finestra; scartare elimina le bozze dalla cache. + +Frammento in lavorazione (`status === 'processing'`): il testo della traduzione +(editor o confronto) è coperto da `PagePendingOverlay` `tone="running"` +(«Traduzione in corso…»), che lo rende inerte; prima restava scrivibile e si +scontrava col risultato della pipeline. La colonna delle fasi è fuori dal velo. +I token di Ollama arrivano solo in `stageResults[stage].content` +(`appendChunkStageContent`); il foglio sull'ultima fase mostra +`translationDisplayText`, scritto a fine fase, quindi nella vista normale non +c'è testo che arriva man mano. + +I comandi delle fasi (una per fase, confronto, coppie del confronto) sono una +colonna verticale nel margine destro della pagina (`DocumentPage.sideRail`, +`IconButton` xs con suggerimento a sinistra, ferma mentre il testo scorre): la +testata non cresce e resta alta come quella dell'originale. Il margine destro è +più largo su entrambe le pagine (`pr-11`), con o senza colonna. + +Verifica: `translationLocked` resta il dato (e `translation_locked` la colonna), +ma in interfaccia è «verificata» — `CircleCheck` accanto al titolo del foglio +della traduzione, `success` quando acceso, spento con il motivo a frammento in +lavorazione o senza testo. Accanto, non al posto, l'etichetta ocra +«Sorgente modificata» (`translationStale`); il segno del pallino è +`editorial-warning`. `toggleChunkTranslationLock`, quando verifica, azzera +`translationStale`. Nessun lucchetto sull'originale: lo stato modificabile lo +dice la matita. I comandi spenti di fogli, colonna delle fasi, riga in cima, +riquadro di esecuzione e barra principale (`ShellNavItem.disabledReason`) +portano il motivo nel formato «Comando — motivo» +(`transcription.commandBlocked`). **Limite accettato** (Niki, 2 ottobre): +`translationStale` vive solo in memoria (non è in `translation_chunks`), quindi +riaprendo il progetto il «da aggiornare» si perde; non si aggiunge una colonna. + +Storico del frammento (Revisione → Storico, `TranslationHistoryList`): legge +`translation_revisions` del frammento (`listTranslationRevisions`, con +`translations.approved_revision_id` per il segno «verificata»). Autori: `model` +(passata della pipeline e riscrittura dopo l'audit, non distinte: lo schema non +lo dice) e `human` (verifica con testo diverso, salvataggio manuale, ripristino). +Il salvataggio manuale è `projectStore.saveVersionNow`: `saveCurrentProject`, +poi `recordManualRevision` per ogni frammento di `unversionedChunks` (testo non +vuoto, non verificato, diverso dall'ultima versione). L'ultima versione per +frammento sta in `translationHistoryStore.latestText`, caricata all'apertura +della pipeline (`useLatestRevisionTexts`, una query) e aggiornata da +`insertRevision` stesso (`noteRevision`, che fa anche rileggere lo storico +aperto). Il dischetto è acceso se c'è da salvare **o** da versionare. +Ripristino = `updateChunkDraft` + versione manuale: nessuna riga si modifica o +si cancella (registro immutabile, voluto). Niente nomi né puntine: servirebbe +una colonna `consolidated_name`, rimandata. + +Configurazione della pipeline (`document/ConfigDrawer`, `Dialog` aperto da ⚙ in +`PipelineSwitch` o Ctrl/⌘+,): eyebrow «Configura pipeline», titolo = nome della +pipeline (la rinomina resta in `PipelineSwitch`), `TabStrip` (`idPrefix` +`pconfig`) nella fila della finestra con il nome della linguetta accanto. +Linguette (`ConfigSection`): `settings` Generale (`SettingsTabPanel`: modalità +su `ChoiceDots`, coppia DeepL sempre visibile e attiva solo in DeepL, Descrizione comune con bozza/conferma), `translation` Fasi +(`TranslationTabPanel` → `StageCard` per fase + memoria di contesto in fondo), +`audit` Controllo qualità (`AuditTabPanel`), `memory` Memoria (`MemoryTabPanel`, +spenta con motivo in modalità DeepL; se era aperta si torna a Generale), +`glossary` (`GlossaryTabPanel`: assegnazione, termini, dischetto acceso solo con +modifiche, caricamento DeepL), `preview` (`PromptPreviewTab`, fasi su +`TabStrip`). `PipelineConfig` monta solo il corpo della linguetta aperta e il +velo comune `PagePendingOverlay` (`components/common`, lo stesso delle +Trascrizioni) durante la pipeline, che rende inerti i comandi coperti. Fase e +giudizio condividono `ModelSection` (fornitore, modello, lucchetto se +esistono traduzioni, ricarica Ollama, ragionamento e temperatura, opzioni +Ollama in `ProviderRuntimeEditor`, cache Anthropic); le regole di taratura +sono funzioni pure in `pipeline/modelTuning.ts`. Ogni prompt pipeline (fasi, descrizione, +giudizio, coerenza) usa `PipelinePromptEditor` con `disabledReason` e +`refineDisabledReason`. I modelli di prompt salvati si applicano e si salvano da +`PromptTemplateMenus` (anche nell'OCR); si eliminano solo dalle risorse +linguistiche (`PromptTemplatesTab`). Nessuna spiegazione fissa: stanno negli `hint` di +`PanelSection`, `SettingRow`, `ToggleRow`. Footer: `IconButton` danger +«Azzera tutte le traduzioni» (`resetAllChunks`, conferma), spento con motivo. + +Composizione: `TranslationStudioHeader` (`PageHeader` area traduzioni: nome con +`RenameField`, accessorio `titleAccessory`: separatore /, `PipelineSwitch` con menu scelta/creazione/rinomina/eliminazione e ⚙, tipo Semplice/Editoriale/DeepL come sola icona con `Hint`; a destra `WorkLanguagesControl` (coppia dell'opera e icona che apre la finestra; la coppia DeepL resta nelle opzioni della fase), `CommandRule`, poi importa, esporta, +risorse linguistiche del workspace, elimina), al centro `DocumentView` invariato salvo la fila +`ChunkStrip` («nn/nn», poi una finestra di 7 `ChunkDot` con il frammento +aperto fisso al centro: la fila intera trasla di `SLOT_PX` per posto, i +pallini fuori finestra restano montati per lo scorrimento ma con `tabIndex` +-1 e `aria-hidden`; frecce ±1 e ±7, rotella con ascoltatore nativo non +passivo), `StageStatusRow` (spie delle fasi del frammento aperto, aprono +`StageTraceDialog`) e la lente che apre `SearchTab` sotto la fila (regione, +non più linguetta; Esc dal campo la chiude), a destra `TranslationInspector`: `InspectorShell` con `beforeTabs` +per l'esecuzione (`PipelineSidebarRunSection` + `ChunkCostPanel`: una riga +stima/speso che con un clic apre a sinistra un `ClickPopover` con due +`CostTable`; il consumo viene da `summarizeChunkUsage`, che per ogni fase dà +`calls` = chiamate completate, non esecuzioni; il registro non segna a quale +esecuzione appartiene una chiamata) e cinque +linguette (Glossario, Memoria, Anteprima, Revisione, Documento) su un solo +stato, `uiStore.studioTab` (`TranslationStudioTab` = +linguette del frammento ∪ linguette del documento). `studioGroupViews` conserva +in memoria la scelta di ogni gruppo; `setStudioTab` registra anche la vista +lasciata aperta dai percorsi di navigazione esterni. Il primo ingresso in +Revisione sceglie Audit/Note in base alle segnalazioni, i successivi riaprono +la scelta dell'utente; una vista indisponibile usa il ripiego del suo gruppo. +Cambio e creazione pipeline sono bloccati durante `chunksStore.isProcessing`, +nel menu e nelle operazioni di `projectStore`; il cambio ricontrolla il blocco +anche dopo le letture asincrone, prima di sostituire configurazione e frammenti. +Revisione (`ReviewTab`) è +una linguetta della colonna ma quattro valori di `studioTab` — `audit`, `notes`, +`sourceNotes`, `history` — mostrati come sottolinguette (`TabStrip`, linguette a icona con +nome e conteggio nell'etichetta, Audit spento con il motivo); `setStudioTab('notes')` da altri punti apre quindi Revisione +sulle note. Le note del testo (`SourceNotesList`) sono le note a piè di +pagina importate, in sola lettura. Memoria (`MemoryGroupTab`, colonna +`phraseMemory`) segue lo stesso schema con `references` e `memory`, Documento +(`DocumentGroupTab`) con `index`, `stats`, `coherence`; tutte e tre usano `SubTabsPanel` (fila `TabStrip` ferma, nome della sottolinguetta +accanto, un solo corpo che scorre). Le schede con barra fissa +sopra un elenco (Riferimenti, Memoria, Glossario, Indice) scorrono da +sé (`bodyScrolls` falso), le altre scorrono nella colonna. Aperta/chiusa e +larghezza restano `showInsightPanel` e `projectFlyoutWidth`; colonna chiusa = +solo traduci/stop. Le note aperte da altri punti (menu contestuale, +segnalazione dell'audit) usano `setShowInsightPanel(true)` + `setStudioTab('notes')`. + Il motore frontend coordina: 1. controllo di provider e modelli; @@ -1717,3 +2084,82 @@ aprire lo stesso archivio. I backup precedenti non sono supportati. - E2E: smoke test Chromium sul primo avvio e sul flusso progetto. - CI: TypeScript/ESLint, test frontend, E2E, `cargo check`, `cargo fmt`, `cargo clippy -D warnings`, test Rust e audit dipendenze. + + +### Corpus testuale e misure multiple (base attuale) + +`text_units` identifica un testo con riferimenti opzionali a opera/versione, parent e +provenienza JSON congelata. `text_unit_revisions` contiene revisioni immutabili per +ruolo (`source`, `translation`, `normalized`) e lingua, impronta e numero progressivo. +Un trigger impedisce UPDATE delle revisioni. La base non richiede una traduzione; +l’interfaccia attuale crea le unità dal percorso Memoria. Pagine/sezioni e analisi +restano nella roadmap, non vengono simulate associando indici di chunk a pagine. + +`text_embeddings` usa chiave `(revision_id, provider, model, dimensions, profile)`; +ogni record richiede metadati completi e un BLOB della misura corretta. Il profilo +attuale è `source-verbatim-v1`, OpenAI small 1536 / large 3072. Il comando di rete +controlla indici, completezza, dimensioni e valori; non si scartano coppie approvate +per risposte incomplete. Cache legacy `source_phrase_embeddings` rimossa: nessun +writer o consumer attivo. Vettori utente inclusi nei backup, artefatti cache no. + +`phrase_memory` punta alle revisioni correnti source/translation della stessa unità; +`phrase_memory_entries` espone i campi per UI e ricerca e deriva il workspace dal +progetto vivo. Uno spostamento non modifica le evidenze storiche; senza progetto +la coppia è senza workspace. Snapshot conserva libro/versione, traduzione/chunk, +workspace originale, impronte e selezione testuale. Offset solo quando univoco. + +`vector/text_units` gestisce salvataggio, revisioni, misure e tag; `memory_search` +filtra modello/provider/dimensioni/profilo/lingue/scope prima della distanza cosine +tramite CTE MATERIALIZED. `memory_commands` contiene i wrapper Tauri. Non inserisce +coppie salvate come match a distanza zero; il frontend conserva queste nella Memoria. +`vec_add_phrase_embedding` verifica la revisione attesa, `vec_set_phrase_tags` +modifica solo tag dell’unità. `vec_update_phrase_memory(input)` confronta entrambe +le revisioni prima della scrittura atomica. La rigenerazione del workspace aggiunge +il modello selezionato e conserva tutti gli altri, con controllo dello snapshot. + +Lingua delle revisioni: `language` (ISO 639-3, `und` = non indicata) più +`language_variety` (Glottocode) e `language_note` (migrazione 0008). Una lingua +corretta è una revisione nuova con lo stesso testo: `text_languages::relabel_project` +crea le revisioni per le frasi dell'opera con lingue diverse da quelle +dell'opera, ricopia le misure dell'originale, sposta i puntatori di +`phrase_memory` e registra `text.language.changed`, tutto in una transazione; +`count_relabels` le conta (una sola query con le due revisioni in join). Comandi +`vec_count_project_phrase_relabels`, `vec_relabel_project_phrases`. Elenco e +ricerca restituiscono anche `source_language_variety`/`target_language_variety` +(join sulle revisioni, la vista `phrase_memory_entries` resta invariata); la UI +le mostra con `LanguagePairLabel` («Italiano (Old Italian) → Inglese») in cima +allo Studio, nelle voci di Risorse linguistiche → Memorie e in testa a ogni +riferimento. + +Tag manuali in `text_unit_tags`, non proposte automatiche né vocabolario controllato. +Fatti `text.revision.created`, `text.embedding.saved`, `text.tags.changed` nel +registro esistente con chiave distinta per ogni azione. Backup aggiunge le quattro +tabelle prima delle coppie, differisce i parent autocorrelati, ripristina con INSERT +rigoroso e include revisioni/vettori/tag. `0003_text_corpus.sql` aggiorna lo schema +senza modificare la baseline già applicata e conserva testi, metadati e provenienza. +Trasferisce solo vettori con modello registrato e dimensione compatibile; nessun +modello viene inferito per quelli senza identità. Nessun reset del database. +Il consolidamento prima del merge è riservato all'utente. + +`createProject` crea progetto, pipeline e origine di libro opzionale nella stessa +transazione. La finestra di creazione propone versioni della Biblioteca, senza +deduzioni dal file. Il collegamento non cambia né importa una trascrizione. + +Costi Studio: stima e consumi usano il Popover comune; eliminati posizionamento, +portal e timer privati. Calcoli di stima e contatori del frammento invariati. +Opzioni vista su SettingRow/IconButton; importazione nella schermata vuota con +IconButton neutro. + + +### Protezioni Studio e accessibilità + +`leaveProject` rifiuta l’uscita quando il progetto aperto è in elaborazione e +ricontrolla dopo le attese di salvataggio: nessuna pulizia dei chunk mentre +la pipeline lavora. Feedback di salvataggio/storico tradotti, eccezioni nel log. +`reportUiError` applica la stessa regola ai nuovi flussi delle risorse. +Date del catalogo e memorie normalizzate UTC con `timestampOf` già comune. +`ChoiceDots` assegna il tab stop alla scelta corrente solo se disponibile, +altrimenti alla prima abilitata. `TabButton` usa aria-disabled e blocca il click; +le linguette indisponibili sono focusabili per la spiegazione e saltate dalle frecce. + +Lo Studio conserva le viste dei gruppi Memoria/Revisione/Documento in `studioGroupViews`; cambio e creazione pipeline sono bloccati durante elaborazione anche nello store. `CatalogSearchField` supporta `disabled` e motivo accessibile; il catalogo memorie disabilita la ricerca durante modifica. Le stringhe di provenienza lunghe usano `StatBlock` per andare a capo. diff --git a/docs-dev/PRODUCT_ARCHITECTURE_2_0.md b/docs-dev/PRODUCT_ARCHITECTURE_2_0.md index 1aef59bd..52ed3ee0 100644 --- a/docs-dev/PRODUCT_ARCHITECTURE_2_0.md +++ b/docs-dev/PRODUCT_ARCHITECTURE_2_0.md @@ -390,6 +390,31 @@ Glossa prepara dataset versionati ed esportabili; l'addestramento resta esterno nella 2.0. Modelli o adapter prodotti fuori dall'app possono essere registrati, valutati e riusati tramite provider locali come Ollama. +### Base del corpus testuale — decisioni, implementazione da completare + +Unità testuali di lunghezza variabile (frasi, passaggi, pagine, sezioni), con +identità, revisioni e provenienza indipendenti dai modelli. Conservare il testo +archiviato e la selezione nella fonte, anche attraverso più pagine. Originale, +normalizzazione e traduzioni sono rappresentazioni distinguibili. Correzioni +della fonte non riscrivono silenziosamente gli estratti già utilizzati. + +Memoria traduttiva, raccolte di studio ed evidenze documentali condividono la +base; una «tecnica» è un esempio di classificazione, non il tipo obbligatorio +del corpus. Tag manuali riutilizzabili separati dai metadati della fonte; +annotazioni automatiche identificano input/versione, modello e conferma umana. +Vocabolari controllati, gerarchie e classificazione automatica sono estensioni. + +Una unità può avere più embedding. Ogni misura dichiara input e revisione, +ruolo del testo, provider/modello, dimensione e preparazione. Modello obbligatorio, +senza eccezioni per dati precedenti. Il workspace mantiene un modello attivo; +aggiungere una misura conserva le altre. Ricerca solo fra misure compatibili, +senza confronto fra modelli diversi. La ricerca globale esplicita della memoria +non sposta gli oggetti né estende implicitamente l'ambito di ogni corpus. + +Issue collegate: #227, #391, #382, #380, #377, #209, #223 e #381; i passi +successivi sono nella roadmap. La base deve integrare revisioni e provenienza già +esistenti, non creare un registro parallelo scollegato. + ## 10. Scriptoria come riferimento principale Ogni issue implementativa 2.0 deve indicare quali moduli Scriptoria sono stati diff --git a/docs-dev/README.md b/docs-dev/README.md index a7ffb71b..d46cff60 100644 --- a/docs-dev/README.md +++ b/docs-dev/README.md @@ -24,3 +24,6 @@ ed export. #186 e #446 tracciano adozione e adattamento dei pattern. Piani e specifiche implementative completati non restano come documentazione: invarianti in architettura, regole visive nel design system, lavoro residuo nella roadmap. + +Piano aperto: [Ricerca federata e Dashboard](./PIANO_RICERCA_FEDERATA_DASHBOARD.md), +riferimento per il lavoro sulla Dashboard; si toglie a lavoro concluso. diff --git a/docs-dev/RESOCONTO_BIBLIOTECHE_API_E_PRIORITA_2026-09-12.md b/docs-dev/RESOCONTO_BIBLIOTECHE_API_E_PRIORITA_2026-09-12.md deleted file mode 100644 index 82bb6930..00000000 --- a/docs-dev/RESOCONTO_BIBLIOTECHE_API_E_PRIORITA_2026-09-12.md +++ /dev/null @@ -1,171 +0,0 @@ -# Biblioteche: API, alternative e priorità di integrazione - -Data: 12 settembre 2026. - -## Ambito della verifica - -Questo resoconto prende per affidabili gli esiti delle prove forniti dal maintainer. Verifica invece, attraverso documentazione ufficiale, le conclusioni sulle interfacce disponibili e propone nuove integrazioni. - -Non è un nuovo collaudo delle connessioni: un servizio documentato può essere temporaneamente irraggiungibile o bloccato dalla rete usata. Le priorità sono raccomandazioni, non garanzie di funzionamento. - -## Situazione riportata dalle prove - -- Ricerca funzionante: Internet Archive, Vaticana, Gallica, e-codices, Institut de France, Bodleian, Estense. Per Estense restano problemi con le copertine. -- Apertura per segnatura o indirizzo: Cambridge e Heidelberg. La ricerca libera Cambridge è stata disabilitata dopo blocchi anti-bot osservati. -- Nessuna risposta utilizzabile: Library of Congress e Harvard. -- I messaggi di errore della ricerca sono stati migliorati per distinguere rifiuto automatico, limiti e indisponibilità del servizio. - -## Conclusioni da correggere o qualificare - -### Library of Congress - -**Esiste un'API pubblica di ricerca documentata.** Supporta ricerca, filtri, paginazione e dettagli delle opere. Il blocco osservato può essere reale, ma non dimostra l'assenza di un'interfaccia pubblica. - -Prima di rinunciare all'integrazione, confrontare il percorso usato da Glossa con quello ufficiale. La ricerca generale include anche contenuti diversi dalle collezioni digitali: restringere opportunamente l'ambito. La presenza di un risultato non garantisce una digitalizzazione IIIF utilizzabile. - -Fonte: [Library of Congress — API endpoints](https://www.loc.gov/apis/json-and-yaml/requests/endpoints/). - -### Harvard - -LibraryCloud offre un'API documentata con ricerca per campi e faccette. Un blocco dell'indirizzo di provenienza è una possibile spiegazione del messaggio «troppe richieste», non una causa dimostrata da quel messaggio da solo. - -Registrare separatamente errore osservato e causa ipotizzata; verificare accesso al catalogo e disponibilità della digitalizzazione come capacità distinte. - -Fonti: [Harvard — APIs and datasets](https://library.harvard.edu/services-tools/harvard-library-apis-datasets), [LibraryCloud — documentazione API](https://harvardwiki.atlassian.net/wiki/spaces/LibraryStaffDoc/pages/43287734/LibraryCloud%2BAPIs). - -### Heidelberg - -Pubblica interfacce aperte, comprese IIIF e raccolta dei metadati. Questo non dimostra che esista una ricerca libera pronta da integrare, ma lascia alternative alla lettura delle pagine del sito. - -Occorre verificare quale interfaccia copra effettivamente i fondi digitalizzati interessanti per Glossa. L'esistenza di interfacce per altri servizi dell'università non garantisce copertura del catalogo storico. - -Fonti: [Heidelberg — interfacce aperte](https://www.ub.uni-heidelberg.de/helios/kataloge/datenschnittstellen.html), [heiOPENsearch — ambito e ricerca](https://www.ub.uni-heidelberg.de/Englisch/service/heiopensearch/hilfe.html). - -### Cambridge - -Mantenere per ora l'apertura diretta funzionante. In questa verifica non è emersa una ricerca pubblica documentata che risolva il problema segnalato. Non è una prova che nessuna alternativa esista. - -L'accesso IIIF e la ricerca del catalogo sono capacità separate. - -Fonte sul supporto IIIF: [Cambridge University Library — progetto IIIF](https://www.lib.cam.ac.uk/research/digital-humanities/case-studies/dynamic-digital-library-iiif-scoping-project). - -## Nuove integrazioni consigliate - -| Priorità fra le nuove integrazioni | Biblioteca o servizio | Motivo | Limite da chiarire | -| --- | --- | --- | --- | -| 1 | Wellcome Collection | API di ricerca documentata, senza registrazione, con date, lingue, tipologie, disponibilità e accesso IIIF | Distinguere opere catalogate e materiali digitalizzati effettivamente leggibili | -| 2 | e-rara | Pertinente agli stampati antichi; documenta metadati, PDF e IIIF | La documentazione consultata descrive raccolta metadati, non una ricerca libera pronta | -| 3 | e-manuscripta | Complementare a e-codices per materiali manoscritti; offre metadati e IIIF | Verificare separatamente la ricerca automatizzabile | -| 4 | Bayerische Staatsbibliothek / MDZ | Amplia il materiale storico; documenta manifesti, immagini e metadati | Accesso IIIF documentato; percorso di ricerca da valutare | -| 5 | Europeana | Amplia la scoperta trasversale fra istituzioni | Misurare copertura e qualità dei collegamenti alle scansioni | - -L'ordine di realizzazione può differire dalla pertinenza dei contenuti: Europeana può precedere e-rara/e-manuscripta se l'obiettivo immediato è ampliare la ricerca attraverso un servizio documentato. - -### Wellcome Collection - -Prima aggiunta consigliata per rapporto fra utilità e prevedibilità tecnica. Ricerca e accesso alle digitalizzazioni sono servizi esplicitamente destinati agli sviluppatori. La ricerca è accessibile senza autenticazione o registrazione. - -L'API documenta filtri per date di produzione, lingua, tipo, disponibilità e altri attributi. Controllare i parametri: quelli non riconosciuti possono essere ignorati, restituendo risultati non filtrati. - -Fonti: [Catalogue API e filtri](https://developers.wellcomecollection.org/api/catalogue), [accesso senza autenticazione ed esempi](https://github.com/wellcomecollection/catalogue-api), [panoramica API e IIIF](https://api.wellcomecollection.org/). - -### e-rara ed e-manuscripta - -Aggiungerle inizialmente tramite collegamento può già essere utile. L'interfaccia OAI-PMH per raccogliere metadati non equivale a un motore interrogabile in tempo reale per titolo e autore: potrebbe richiedere un indice locale, cioè un lavoro distinto. - -Non presentare il supporto alla raccolta dei metadati come prova di una ricerca libera già integrabile. - -Fonti: [e-rara — interfacce e dati](https://www.e-rara.ch/wiki/apiinfo), [e-manuscripta — funzionalità e interoperabilità](https://www.e-manuscripta.ch/wiki/aboutEmanuscripta). - -### Bayerische Staatsbibliothek / MDZ - -Accesso a manifesti e pagine ben documentato. Immagini e OCR hanno limiti da rispettare. La documentazione espone anche raccolte IIIF e un deposito di metadati OAI-PMH. - -Non promettere la ricerca libera prima di averne verificato il canale specifico. - -Fonte: [MDZ — interfacce ufficiali](https://www.digitale-sammlungen.de/de/schnittstellen). - -## Europeana: utile, ma non una soluzione unica - -Europeana è una buona aggiunta come aggregatore. Nella presentazione dei risultati distinguere: - -- **Trovato tramite Europeana:** servizio che ha restituito il risultato. -- **Conservato presso una biblioteca:** istituzione responsabile del materiale. -- **Immagini servite da un sito:** origine effettiva delle pagine da leggere o scaricare. - -Europeana può risolvere la scoperta di un libro senza risolvere un blocco sullo scaricamento delle sue immagini. La presenza di un'istituzione non garantisce che tutte le sue collezioni siano indicizzate. La copertura specifica di Bodleian, Heidelberg, Estense o altri fondi va misurata con esempi, non assunta. - -Non scartare automaticamente ogni risultato senza IIIF. Proposta: filtro **«Solo fonti leggibili in Glossa»**. Gli altri risultati possono restare schede bibliografiche con collegamento esterno, senza offrire lettura o download impossibili. Questa è una proposta di prodotto, non una capacità già implementata. - -Le API Europeana distinguono ricerca, record e accesso IIIF. Anche un manifesto disponibile va verificato per capire quali immagini e quale copertura dell'opera rappresenti. - -Fonti: [Europeana — panoramica API](https://api.europeana.eu/en), [Search API](https://europeana.atlassian.net/wiki/spaces/EF/pages/2385739812). - -### Chiave di accesso - -Europeana distingue chiavi personali per sperimentazione e chiavi di progetto per servizi operativi. Per la sperimentazione della beta si può partire dal percorso personale; non promettere genericamente che una chiave di progetto sia «immediata». - -Fonte: [Registrazione e gestione delle chiavi Europeana](https://www.europeana.eu/en/how-to-register-for-and-manage-an-api-key). - -## Ulteriore riferimento: Biblissima+ - -Particolarmente pertinente per manoscritti e stampe antiche. Documenta raccolte che collegano diverse digitalizzazioni dello stesso manoscritto o diversi esemplari di un'edizione. - -È interessante sia come possibile fonte sia come riferimento per presentare le relazioni bibliografiche in Glossa. Non considerarlo un sostituto pronto della ricerca federata: verificare il servizio specifico e i relativi contratti prima di pianificare l'integrazione. - -Fonti: [Biblissima+ — API Presentation](https://doc.biblissima.fr/api/api-presentation/), [modalità di condivisione dei dati](https://doc.biblissima.fr/vademecum-biblissima/). - -## Ordine operativo suggerito - -1. Verificare Library of Congress attraverso il percorso ufficiale, distinguendo API esistente e blocco osservato. -2. Aggiungere Wellcome come nuova integrazione singola. -3. Valutare Europeana come aggregatore, misurando su un campione quanti risultati conducono a fonti leggibili. -4. Sviluppare e-rara, e-manuscripta e MDZ secondo i materiali realmente usati, iniziando dall'apertura diretta dove necessario. -5. Studiare Biblissima+ per copertura specialistica e relazioni fra copie ed edizioni. - -Per ogni biblioteca verificare separatamente tre passaggi: **trovo l'opera, apro il libro corretto, leggo le pagine**. Una ricerca con molte schede e poche fonti utilizzabili non è necessariamente un'integrazione riuscita. - -## Aggiornamento 12 settembre 2026 (dopo l'implementazione) - -Le raccomandazioni di questo resoconto sono state seguite. Stato reale al -termine del lavoro, verificato interrogando i servizi: - -- **Wellcome Collection** aggiunta, come suggeriva la priorità 1. La sua - interfaccia risponde senza registrazione, e il filtro - `items.locations.locationType=iiif-presentation` toglie alla fonte i libri - mai digitalizzati: su una ricerca di prova erano quattro su cinque. -- **Europeana** aggiunta, con chiave nel portachiavi del sistema. Si tengono - solo i risultati che dichiarano una riproduzione; l'istituzione che conserva - l'originale è mostrata a parte da chi ha risposto alla ricerca. -- **e-rara, e-manuscripta, Bayerische Staatsbibliothek** aggiunte per apertura - diretta. Come il resoconto prevedeva, nessuna delle tre offre una ricerca - interrogabile: le prime due rispondono con un controllo anti-robot, Monaco - pubblica manifesti e raccolta dei metadati. -- **Library of Congress**: l'API esiste ed è documentata, come qui si diceva. - Dalla rete di sviluppo ogni percorso risponde con un controllo anti-robot, - compreso il JSON di un singolo elemento noto: è un blocco dell'indirizzo di - provenienza, non l'assenza di un'interfaccia. -- **Harvard**: l'interfaccia esiste e i parametri usati sono quelli - documentati. Risponde «troppe richieste» a ogni tentativo, da due reti - diverse e senza indicare quando riprovare: la ricerca è stata disattivata e - resta il riconoscimento del gettone. -- **Cambridge**: confermato che non emerge una ricerca pubblica utilizzabile. - Il filtro del sito blocca dopo poche richieste, quindi la ricerca libera è - stata disattivata e resta l'apertura per segnatura. -- **Biblissima+** non è stato affrontato: resta da studiare, come qui si - raccomandava, prima di pianificarne l'integrazione. - -Il filtro «Solo fonti leggibili in Glossa» non è stato costruito: per ora si -scartano alla fonte i risultati senza riproduzione, che è la stessa promessa -mantenuta con meno interfaccia. Resta una proposta valida per quando si vorrà -mostrare anche le schede bibliografiche. - -## Punti da sottoporre alla prossima analisi - -- Quale endpoint usa oggi ciascun provider e quale capacità documentata copre? -- Quali filtri sono realmente applicati al catalogo remoto? -- Quali errori sono stati osservati e quali cause sono soltanto ipotesi? -- Il risultato identifica una scheda, un'edizione, un esemplare o una digitalizzazione? -- Manifesto e immagini sono raggiungibili indipendentemente dalla ricerca? -- Un aggregatore introduce doppioni rispetto ai provider diretti? -- Quanto costa mantenere l'integrazione rispetto al valore dei fondi effettivamente usati? diff --git a/docs-dev/ROADMAP_2_0.md b/docs-dev/ROADMAP_2_0.md index 519c403d..ec4f7084 100644 --- a/docs-dev/ROADMAP_2_0.md +++ b/docs-dev/ROADMAP_2_0.md @@ -1,6 +1,15 @@ # Roadmap verso il completamento della beta -Aggiornata: 29 settembre 2026. +Aggiornata: 7 ottobre 2026. + +## Prossimi passi concordati (7 ottobre 2026) + +- **Dashboard**, insieme ai colori delle superfici (#485 D). +- Unione di #488 su `main` con l'accorpamento delle migrazioni 0004-0008; poi + chiusura di #489. +- Da chiudere dopo verifica: #481, #467, #469. +- Futuro, senza data: sblocco del prompt delle fasi quando esistono frammenti + tradotti. ## Ricerca e Dashboard: consegna corrente e consolidamento @@ -140,7 +149,8 @@ Issue: #183, #187, #397, #459, #462, #413, #471, #472; shell generale #210. - Verificare i percorsi già presenti e correggere le regressioni prima di aggiungere nuove varianti. Revisione UI/UX generale ancora da fare. -- Rendere visibili log generali, salvataggio e stato dei lavori (#413), +- Rendere visibili log generali, salvataggio (fatto per traduzioni e + trascrizioni, mancano le fonti) e stato dei lavori (#413), riusando la coda e i pannelli esistenti. - Riallineare capacità dichiarate e reali dei provider (#397). La ricerca aggregata (#395) viene dopo la verifica dei singoli provider. I risultati @@ -221,6 +231,43 @@ e il testo delle pagine vicine come contesto di continuità nel prompt. L'immagine inviata (ottimizzata a misura scelta o copia locale così com'è) si sceglie in Impostazioni → Trascrizioni e per sessione nella scheda OCR. +**Stato al 1° ottobre 2026 (#485 H e L).** Lo Studio sta nella finestra a +ogni larghezza; salvataggio manuale di una versione (comando e Ctrl+S). La +pagina iniziale delle Trascrizioni è un catalogo sul modello della Biblioteca, +con gli stessi pezzi condivisi (scaffali, ricerca, filtri rapidi, tre viste, +comandi di riga, barretta di completamento). Anche la pagina iniziale delle +Traduzioni segue lo stesso modello (#485 N, primo passo: scaffali, filtri +workspace e lingue, rinomina | elimina, creazione «da zero» con il file). +**Restano**: strada «da una trascrizione» con copia fissata, legame con opera +e trascrizione d'origine e, solo allora, i comandi apri l'opera / apri la +trascrizione, filtri per biblioteca e secolo, raggruppamento per biblioteca; +Studio di traduzione (#485 N): fatta la disposizione (T1: barra principale +sempre in vista, riga d'intestazione, una colonna Strumenti a destra) e il +salvataggio (T2: stato nella barra di stato anche per le trascrizioni, +dischetto e Ctrl/⌘+S nei fogli, salvataggio prima di uscire) e la verifica (T3: spunta, motivi dei comandi +spenti; il «da aggiornare» resta solo in memoria per scelta, si perde +riaprendo) e lo storico (T4: sottolinguetta in Revisione, versioni dal +dischetto, ripristino; **in futuro**: nomi/puntine sulle versioni, con una +colonna nuova) e la configurazione della pipeline (T5: finestra a sei linguette +comuni, sezione Modello e editor dei prompt comuni, niente spiegazioni fisse, +«Azzera tutte le traduzioni» a icona, velo comune) e il velo oro sul frammento +in traduzione (T6) e Memoria/risorse linguistiche (T7: scheda del frammento, +modelli modificabili con filtro OCR, dizionari con ambito visibile, memorie +con provenienza/modello e filtri, ricerca fra workspace). T7 implementato: +resta la prova dal vivo. Decisione del 3 ottobre: nessuna retrocompatibilità +per embedding senza modello; più misure per unità testuale, modello obbligatorio +e ricerca compatibile. Base corpus implementata: unità/revisioni/misure/tag, +provenienza e libro esplicito; selezione di pagine/sezioni e analisi restano future. +Schema aggiornato da `0003_text_corpus.sql`, baseline applicata conservata; +consolidamento prima del merge riservato all'utente. +Costi T8 su pannelli comuni; conteggio blocchi confermato dall'utente. +T9: revisione finale e nove correzioni della review implementate; i due Studio +sono inclusi nello stesso ramo della PR #488. Restano CI e prova dal vivo; +la UI/UX non è approvata dall'utente e richiede una nuova revisione visuale. +quali riepiloghi unire +nella colonna si decide dopo. Scelta multipla nei cataloghi, non chiesta per +ora. + Uscita: aprire una fonte reale, trascrivere più pagine — a mano o assistite da OCR/HTR —, correggere, riaprire e ritrovare testo, revisioni e riferimenti alla fonte. @@ -233,6 +280,10 @@ Issue: #208, #221, #222, #209, #223, #189, #224; risorse contestuali #227. - Creare il progetto di traduzione dal testo approvato senza perdere provenienza. - Collegare fonte, trascrizioni e traduzioni dalla scheda dell'opera. - Integrare corpus e suggerimenti contestuali con ambito workspace chiaro. +- Generalizzare #227 a unità testuali di lunghezza variabile, anche passaggi + attraverso più pagine; tecniche storiche come esempio di classificazione. + Tenere distinti corpus testuale ed evidenze visive #209/#223, collegandone + le provenienze. Tag manuali riutilizzabili e versioni del testo nella base. Uscita: image workbench e corpus di frammenti pronti; passaggio trascrizione → traduzione con provenienza conservata, storico ricostruibile. @@ -267,6 +318,14 @@ L'addestramento resta esterno a Glossa. Il loro perimetro minimo per la prima beta completa va deciso sulla base dei casi reali: non dichiararli rimossi né prometterli tutti nel prossimo tag. +La base del corpus testuale è consegnata (unità, revisioni, misure con +modello/dimensioni/profilo, tag, provenienza; vedi architettura). #391 va +riallineata alla conservazione delle misure; #382 riusa questa base per +originali e traduzioni. Restano passi successivi, da valutare su testi +medievali reali: selettore di pagine/sezioni, ricerca ibrida lessicale + +semantica, tag automatici con conferma, vocabolari multilingui (SKOS), +confronti fra testimoni e analisi di opere intere. + ## Riferimenti e lavori trasversali Scriptoria resta il primo riferimento tecnico per fonti, deposito, lavori, diff --git a/docs-dev/UI_DESIGN_SYSTEM.md b/docs-dev/UI_DESIGN_SYSTEM.md index 87e260c8..46576660 100644 --- a/docs-dev/UI_DESIGN_SYSTEM.md +++ b/docs-dev/UI_DESIGN_SYSTEM.md @@ -15,7 +15,7 @@ primitive condivise resta la sorgente di verità per API e varianti disponibili. - Ocra per cautele; giallo solo per attività in corso. - Comandi visivi neutri, icon-only, con tooltip. - Nessuna variante locale quando esiste una primitiva condivisa. -- Testo leggibile almeno `text-xs`; `text-[11px]` solo per caption uppercase. +- Testo leggibile almeno `text-xs`; `text-caption` solo per caption uppercase. - Focus, tastiera e ruoli ARIA fanno parte del componente. ## Palette @@ -40,11 +40,38 @@ Usare ruoli semantici per le superfici: |---|---| | `bg-surface-elevated` | dialog, menu, popover, header sticky | | `bg-surface-panel` | sidebar, colonne e pannelli | +| `bg-surface-resource` | unica carta tenue per ciascuna voce delle risorse linguistiche | | `bg-surface-hover/50` | hover su righe cliccabili | +Ogni area con un catalogo ha un **inchiostro** e una **carta** propri +(`area-library`, `area-transcriptions` seppia, `area-translations` indaco, e le +rispettive `-paper`), con le classi in `AREA_INK_CLASSNAME` e +`AREA_PAPER_CLASSNAME`. L'inchiostro va solo su segni piccoli: il filetto sotto +il titolo (`AreaHeading`), l'icona dell'area a riposo nella barra di sinistra, +il segnaposto di una copertina mancante. La carta è il fondo dell'elenco e dei +titoletti di gruppo fermi in cima; la colonna degli scaffali resta +`surface-panel`. Mai su comandi, selezione o stati: lì restano verde, rosso, +ocra e oro. Contrasto sul fondo ≥ 7:1 in entrambi i temi. + Campi, select e textarea usano `editorial-textbox` pieno. Vietati colori Tailwind grezzi e valori esadecimali nei componenti. +**Tema scuro.** Ogni colore è un token di `index.css` con il suo valore in +`html.dark`; la variante `dark:` segue la scelta fatta nell'app, non il +sistema. Il testo sopra un fondo pieno d'inchiostro o d'accento usa +`text-on-ink` / `text-on-accent`, mai `text-white`: nel tema scuro l'inchiostro +diventa chiaro e il bianco sopra non si legge. I colori che servono anche nel +codice (accento modificabile, sfondo per il controllo del contrasto, +evidenziazioni) hanno un test che li confronta con il foglio di stile. + +**Bordi.** Tre forze: `editorial-border` pieno per i bordi strutturali, `rule` +per filetti e separatori di elenco, `rule-faint` per quelli appena accennati +(valgono con `border-`, `divide-`, `bg-`). Niente opacità scritte a mano. + +**Ombre.** `shadow-modal`, `shadow-tooltip`, `shadow-warm-sm`, `shadow-warm-md`, +`shadow-page-card`, `shadow-inset-highlight(-strong)`, `shadow-chunk-current-ring`: +token con la variante scura, mai `shadow-[var(...)]`. + ## Tipografia - `font-display`: titoli di vista e valori editoriali. @@ -53,8 +80,10 @@ Tailwind grezzi e valori esadecimali nei componenti. - Scala: `text-xs` 13 px, `text-sm` 15 px, `text-base` 16 px, `text-lg` 18 px, `text-xl` 22 px, `text-2xl` 26 px. - Titolo di vista: `font-display italic`, responsive solo quando serve. -- Titolo sezione: uppercase, `text-[11px]`, `tracking-[0.16em]`. -- Label statistica: uppercase, `text-[11px]`, `tracking-[0.1em]`. +- Titolo sezione: uppercase, `text-caption` (11 px), `tracking-section`. +- Label statistica, intestazione di tabella: lo stile `caption-label` + (`text-caption`, maiuscolo, `tracking-caption`, muted). Mai `text-[..px]` né + `tracking-[..]` scritti a mano. - Valore statistico: `font-display text-sm italic`. - Metrica focale singola: `font-display text-lg italic`. @@ -117,6 +146,14 @@ interattivi. - Il chiamante controlla `open` e `onOpenChange`. - Il trigger usa `ariaPressed={open}` e non ribalta manualmente lo stato. - Overlay sopra le finestre: `z-[210]`. +- `SearchPicker`: un `IconButton` che apre un elenco lungo con ricerca + (`CatalogSearchField`) e gruppi sotto piccole didascalie maiuscole; voci + `PopoverItem` con il codice o la seconda informazione in `description`; una + scelta chiude. Si usa al posto di un `Select` quando le voci sono molte o + lunghe (lingue, libri della Biblioteca): un `Select` nativo si allarga alla + voce più lunga e allarga la finestra. Il valore scelto sta accanto, nella + riga, troncato con il testo intero nel suggerimento; la «x» lo toglie. +- `CommandRule`: filetto verticale fra gruppi di comandi in una riga. ### PopoverItem e LinkChip @@ -130,6 +167,8 @@ interattivi. - `LinkChip`: etichetta di un legame già stabilito che, cliccata, lo scioglie. Il motivo sta nel `Tooltip`, mai nel `title` nativo; il nome leggibile del legame resta il nome del comando. +- `PopoverItem` accetta `description`, una seconda riga a spaziatura fissa + (l'inizio di un modello di prompt salvato). - Nessuna riga di elenco, etichetta di legame o voce di menu scritta a mano nei componenti. @@ -149,12 +188,21 @@ interattivi. ### TabStrip +Dentro una linguetta di colonna che raccoglie più viste (Memoria, Revisione, +Documento nello Studio di traduzione) la fila sta in `SubTabsPanel`: ferma in +cima, nome della vista aperta accanto in `font-display text-sm italic`, un solo +corpo che scorre. Mai `SegmentedControl` per questo: è per le scelte con nome +nelle impostazioni. + Fila di linguette icona con la propria navigazione da tastiera: frecce, Home ed End, con il focus che segue la linguetta scelta come vuole il modello ARIA. Usarla per ogni gruppo di linguette che non sia già dentro `InspectorShell` — sotto-schede di una finestra di impostazioni, linguette di un pannello. - `tabs`: `{ id, label, icon }`; l'etichetta vive nel tooltip, non a schermo. +- `disabled` per linguetta: stessa regola dei tab di `InspectorShell` — + visibile, spenta, motivo nell'etichetta, saltata da frecce e Home/End + (sottolinguette della Revisione nello Studio di traduzione). - `idPrefix`: da cui derivano `-tab-` e `-panel-`, così il pannello si collega con `aria-labelledby`. - Il pannello attivo lo monta il chiamante, con `role="tabpanel"`. @@ -194,15 +242,52 @@ solo**. lungo, come la colonna dei lavori in Panoramica. **Un solo contenitore che scorre per colonna**: due aree annidate dividono rotellina e tasti fra due destinazioni e nessuna delle due si comporta come ci si aspetta. +- `beforeTabs`: blocco fisso fra intestazione e linguette, in vista con + qualunque scheda (l'esecuzione nello Studio di traduzione). +- Larghezza: `INSPECTOR_WIDTH` per tutti, incluso lo Studio di traduzione + con cinque gruppi di linguette: minimo 320 px, iniziale 400 px, massimo 560 px. - `tabRowHeightClassName`: altezza fissa della barra tab quando accanto c'è un'altra intestazione (la casella della Ricerca, la barra di un visore): le due righe hanno la stessa altezza e lo stesso filetto, e la linea sotto è una sola da una colonna all'altra. +### PanelSection, PageHeader, ResizeHandle + +- `PanelSection`: sezione di una colonna a schede — `SectionLabel` con icona, + filetto `rule` sotto, comandi della sezione a destra. Il corpo della scheda + usa `PANEL_BODY_CLASSNAME` e gli elenchi etichetta–valore + `STAT_LIST_CLASSNAME` (`ui/panelStyles.ts`); le larghezze della colonna sono + `INSPECTOR_WIDTH` (Biblioteca e Studio uguali). +- `PageHeader`: la riga `h-14` in cima a una pagina di dettaglio — ritorno, + segno dell'area nel suo inchiostro, identità, comandi a destra. `center` + resta per gruppi centrati. `titleAccessory` affianca opera / pipeline: nomi troncati, menu e opzioni, tipo con icona; coppia soltanto in DeepL. Il nome della pipeline è l’unico ingresso al menu (piccola freccia muted dentro lo stesso comando, nessun pulsante separato); nel menu i comandi Rinomina/Nuova stanno in un solo gruppo sotto un solo filetto. Nessuna pillola o pannello aggiuntivo nella barra. +- `ResizeHandle`: l'unico divisore trascinabile fra colonne, con nome per chi + legge con la voce; `layer="shell"` fra colonne dell'applicazione. + +### ChoiceDots + +Scelta esclusiva fra poche opzioni a cerchietti da 24 px con icona o lettera +(immagine inviata all'OCR, livello di ragionamento, modalità della pipeline, modo Struttura / Frammento aperto dell'anteprima prompt). +Un'opzione può essere `disabled` (la modalità DeepL senza chiave): resta +visibile, il motivo è nell'etichetta, le frecce la saltano; `role="radiogroup"`, +frecce/Home/End spostano scelta e fuoco, suggerimento per opzione, la scelta in +accento pieno con `text-on-accent`. L'icona di categoria accanto è muted dentro +un `Hint`, mai in ocra. Nessun cerchietto scritto a mano. + +### Barre strette + +Una barra che non ci sta non taglia i comandi: sfoglio, salto a pagina e zoom +restano, i secondari passano nel menu con i tre puntini (una sola misura +decide barra e menu insieme). Una parola di stato accanto al suo segno (la +provenienza della pagina, lo stato del salvataggio) si riduce al segno, con il +testo nel suggerimento e per chi legge con la voce. + ### SettingRow e campi - Ogni impostazione usa `SettingRow` dentro una lista con `divide-y` e - `border-y`. + `border-y`. Subito sotto il titolo di una `PanelSection` la lista usa + `SECTION_SETTING_LIST_CLASSNAME` (solo `border-b`): il filetto del titolo fa + già da bordo superiore e un secondo bordo lo raddoppierebbe. - Riga `py-2.5`, label `text-sm`, una sola icona nel comando a destra. - L'etichetta prende lo spazio disponibile, il comando non lo ruba: `SettingRow` incapsula i figli in un contenitore che non si allarga. Un campo a larghezza @@ -227,20 +312,54 @@ solo**. dell'etichetta accanto si legge come una nota a margine invece che come la scelta fatta. La larghezza si lascia al contenuto, senza numeri fissi, salvo un tetto per i testi lunghi. -- Scelte esclusive con nome usano `SegmentedControl`. +- Scelte esclusive con nome: `ChoiceDots` (o `SettingChoiceRow` nelle + impostazioni) e `Select`. +- Ogni campo di ricerca usa `CatalogSearchField` (anche nei pannelli e nei + fogli dello Studio): `onKeyDown` per Esc, `focusOnMount` quando si apre da un + comando esplicito. - Interruttori booleani usano `ToggleRow`. ### Dialog +- Finestre di lavoro grandi (anteprima d'importazione): `Dialog` con + `widthClassName="max-w-6xl"`, `panelClassName="h-[90vh]"`, corpo senza + padding a due colonne: impostazioni a sinistra (`w-96`, scorrimento proprio, + sezioni con `SectionLabel`), contenuto a destra a tutta altezza con la sua + barra di conteggi e la scelta di vista su `ChoiceDots`. +- Lingue di un'opera: `WorkLanguagesFields`, due sezioni (Partenza, Arrivo) di + `SettingRow`: Lingua e Varietà con valore in `font-display`, codice mono + piccolo, `SearchPicker` e «x»; Nota con campo in linea. Nella riga in cima + allo Studio la coppia sta a destra, prima dei comandi dell'opera: + `LanguagePairLabel` (nomi e varietà; le note nel suggerimento) e un + `IconButton` lingue che apre la finestra con Annulla/Conferma, mai salvata + all'uscita; poi `CommandRule`. Il tipo di pipeline è solo un'icona con nome e + spiegazione nel suggerimento. +- Riferimenti della memoria: una sola riga di comandi (soglia con − cursore + + e valore, filetto, globo, aggiorna) senza titolo ripetuto; ogni risultato: + percentuale e cerchio, originale in `font-display` e traduzione in sans con + il codice lingua a margine, una riga di provenienza e i dettagli in un `Hint`. + Le coppie della Memoria del frammento (salvate, estratte, scritte a mano) + usano la stessa forma (`PhraseLine`), con campi modificabili al posto del testo. + - Finestre modali tramite `Dialog`; conferme distruttive tramite `AlertDialog`. - Conferma e annullamento usano i pulsanti dialog condivisi. - Niente overlay, focus trap o gestione Escape implementati localmente. - Comandi di conferma testuali sono ammessi solo dentro dialog. +- Una finestra a linguette (Impostazioni, configurazione della pipeline) mette + la fila `TabStrip` nello slot `tabBar`, con il nome della linguetta aperta in + `font-display text-sm italic` accanto. Un'azione distruttiva sull'insieme + (azzerare le traduzioni) è un `IconButton` danger a sinistra del footer, + sempre visibile e spento con il motivo, mai un pulsante a scritta. +- Contenuto bloccato durante un lavoro: il velo comune `PagePendingOverlay` + (`components/common`) con la sua riga di stato; rende inerti i comandi + coperti. Nessun velo scritto a mano. `tone="running"` lo fa oro, per il + frammento che la pipeline sta traducendo (solo il testo, non la colonna + delle fasi). ### Badge numerici Conteggi compatti non interattivi: cerchio `h-5 w-5`, testo -`text-[10px] font-bold text-white`, tooltip e `aria-label`. Il colore deriva +`text-[10px] font-bold text-on-accent`, tooltip e `aria-label`. Il colore deriva dalla mappa semantica esistente. Conteggi cliccabili usano una primitiva interattiva. @@ -249,7 +368,8 @@ interattiva. ### Intestazione di un'area Ogni area globale apre con il **titolo grande** in `font-display` corsivo -(`text-4xl md:text-5xl`): Traduzioni, Trascrizioni, Analisi, Biblioteca. I +(`text-4xl md:text-5xl`): Traduzioni, Trascrizioni, Analisi, Biblioteca; le tre +aree con un inchiostro usano `AreaHeading`, che aggiunge il filetto. I comandi propri dell'elenco (vista, ordinamento) stanno in fondo alla stessa riga, allineati alla base del titolo. La Biblioteca usava una `SectionLabel` piccola con icona: era l'unica area a non somigliare alle altre. @@ -291,6 +411,25 @@ precedente si accendeva solo entro trenta secondi dall'ultima risposta — contando anche le immagini lette dal disco — e su un libro tutto online restava spento quasi sempre. +### Coppia di lingue + +Una coppia di lingue si scrive sempre con `LanguagePairLabel` +(`components/languages`): «Italiano (Old Italian) → Inglese», `font-sans +text-xs`, mai corsivo. Nomi in inchiostro, varietà fra parentesi e freccia in +muted; tronca da sola e il testo intero va nel suggerimento. Accanto a un +comando lascia `mr-2` di respiro. Nelle righe di frase sta nella riga di testa, +non nel margine dei codici, così le frasi restano allineate. + +### Costi + +Stima e consumo usano `CostTable` (`components/pipeline`): una riga per fase con +nome, modello mono sotto, token (cache sotto) e costo, totale in fondo; per il +consumo anche le chiamate. Nessun carosello, nessuna carta. Nello Studio la riga +sotto i comandi è un solo pulsante senza suggerimento (le cifre si leggono già) +che apre un `ClickPopover` con le due tabelle: solo titoli e tabelle, né +paragrafi né suggerimenti dentro (sarebbe un secondo riquadro sopra il primo). +Il consumo non si colora: il verde resta per scelta e stato attivo. + ### Completamento in una riga di elenco Quanto di una cosa è già disponibile si dice con una **riga di dati a @@ -332,6 +471,28 @@ non in maiuscoletto spaziato: il maiuscolo su elenchi lunghi si legge peggio e rallenta. Vale per i dati che arrivano da fuori (le voci di un manifesto), dove i nomi li sceglie la biblioteca e possono essere lunghi. +### Catalogo: un solo modello + +Biblioteca e Trascrizioni sono lo stesso catalogo, e le Traduzioni lo +seguiranno: titolo grande con `CatalogViewSwitch` (elenco, copertine, tabella) +e i comandi d'insieme in fondo alla riga, `CatalogSearchField` e filtri rapidi +sotto, elenco con `CATALOG_LIST_CLASSNAME`/`CATALOG_GRID_CLASSNAME`, gruppi +con `CATALOG_GROUP_HEADER_CLASSNAME`, colonna di `ShelfItem` a destra. I +comandi di riga sono sempre `CommandBar`: icone neutre con la descrizione al +passaggio, in gruppi divisi da un filetto, `inline` nella riga a elenco, +`menu` nelle copertine e nella tabella. L'avanzamento è `CompletionBar`. Le +classi condivise stanno in `ui/catalogStyles.ts`. Un catalogo nuovo copia la +Biblioteca pezzo per pezzo: nessuna riga, scaffale o barra di comandi scritta +a mano. + +Nelle Trascrizioni la riga apre con il **nome** della trascrizione in +`font-display` corsivo `text-lg`, e sotto l'opera in forma compatta +(`WorkIdentity` `header`: autore · anno · luogo, tipografo, titolo su una +riga). Se il nome è il titolo dell'opera non si ripete, e la riga torna quella +della Biblioteca (`WorkIdentity` `row`). Il nome si cambia nel posto in cui +sta: il campo sostituisce la riga del nome, Invio salva, Esc o uscire +annullano. + ### Scaffali e filtri rapidi Un catalogo si organizza con una **colonna di scaffali** a destra (larghezza @@ -440,21 +601,100 @@ colonna vivono in `uiStore` e sopravvivono alla chiusura. - Grip sempre visibile; stato hover, drag e focus riconoscibile. - Animazioni usano i token di motion condivisi. - Fly-out non coprono il controllo che li ha aperti e si chiudono con Escape. +- Barra di sinistra (`WorkspaceRailNext`): solo navigazione (Dashboard con le + sue viste, aree nell'ordine del lavoro, workspace), gruppi aperti da un + filetto (`ShellNavSection`), una riga per voce con la spiegazione nel + suggerimento. Voce scelta: velatura `bg-editorial-accent/6`, nome e + cerchietto in accento, nessuna barretta. Due misure di cerchietto, uguali aperta e chiusa: voci + principali `h-7` icona 14, viste della Dashboard `h-5` icona 11. In fondo il + menu generale (`ShellNavFooter`): salva, risorse linguistiche, impostazioni, + guida, lingua; in fila da aperta, in colonna da chiusa. Chiusa: restano le + icone, niente titoli di gruppo; una voce porta alla pagina senza riaprire. +- Collasso: entrambi i pannelli hanno `PANEL_FLEX_TRANSITION_CLASS`; il + contenuto della barra ha subito la larghezza finale (64 px chiusa, minimo + 198 aperta, cioè la larghezza esatta del menu generale, che è anche quella + iniziale; massimo 420) e il pannello `overflow-hidden` lo scopre senza ricomporlo. ### Impostazioni - Radice: `space-y-10`, `role="tabpanel"`, `aria-labelledby`. -- Sezione: `space-y-4` con `SectionLabel`. -- Elenchi: righe piatte separate, niente card o pill. +- Sezione: `PanelSection` (titolo con il filetto sotto, comandi della sezione + nello slot `actions`), come nella scheda dell'opera della Biblioteca. Sotto, + l'elenco usa `SECTION_SETTING_LIST_CLASSNAME`: il filetto del titolo fa da + bordo superiore, quindi un filetto fra le righe e uno solo in fondo. Mai due + elenchi a filetti attaccati: righe della stessa sezione stanno in un elenco + solo; un `ToggleRow` dentro un elenco va in un contenitore `py-2.5`. +- Elenchi: righe piatte separate, niente card o pill. Una riga isolata senza + titolo sopra ha filetto sopra e sotto. - Una scheda che raccoglie argomenti diversi si divide in **sotto-linguette** - (`TabStrip`) invece di diventare un rotolo unico: accanto alla fila, - l'etichetta della linguetta attiva in `font-display italic`. + con `SettingsSubTabs` invece di diventare un rotolo unico: fila `TabStrip`, + accanto l'etichetta della linguetta attiva in `font-display italic`. +- Ogni impostazione è una riga (`SettingRow`) dentro una lista a filetti: + niente caselle affiancate, riquadri cliccabili o paragrafi fissi (la + spiegazione va nel `hint` della riga o del `SectionLabel`). +- Scelta esclusiva: `SettingChoiceRow` (cerchietti `ChoiceDots` con il nome + della scelta accanto) quando le opzioni hanno un'icona naturale, `Select` + `md` quando sono solo parole. Mai i riquadri larghi di `SegmentedControl`. +- Una cartella è una `FolderRow`, riga di un elenco: nome, percorso mono sotto, + icona cartella a destra per sceglierne un'altra. +- Sotto-linguette (`SettingsSubTabs`) senza filetto proprio: la separazione la + dà il titolo della prima sezione. - Salvataggio al cambio, salvo input intermedi che richiedono conferma esplicita. In quel caso mostrare stato non salvato e comando di ripristino. -- Ordine generale: modalità di traduzione, coppia linguistica, persona. +- Ordine generale: modalità, coppia DeepL sempre visibile (disabilitata fuori DeepL), Descrizione comune del lavoro. ### Elenchi di versioni e comandi per riga +Le anteprime dei prompt della pipeline, nelle opzioni e nel frammento, usano +la stessa superficie `surface-resource` delle risorse linguistiche e il +contrasto `linguistic-resource`: un fondo per messaggio/blocco, testo aperto +senza fondo annidato. Accento verde sul bordo sinistro richiesto esplicitamente dall’utente. Gli editor aprono una bozza con conferma/annullamento a icona; rifinitura e applicazione modelli non salvano prima della conferma. I prompt di sistema nell’anteprima (`SystemTextCard`): carta come gli altri pezzi con lucchetto `IconButton` (aperto in tono accent, `ariaPressed`); aperto diventa `PipelinePromptEditor` con il lucchetto fra i comandi; interruttore di un pezzo facoltativo `ToggleRight`/`ToggleLeft` con `ariaPressed` nel gruppo comandi, anche sulla riga grigia del pezzo spento; i pezzi con contenuto altrove hanno solo `ArrowUpRight` che apre la scheda, senza lucchetto. Testo dei prompt sempre `font-mono`, in lettura e in modifica. L’anteprima delle opzioni è una vista sola: per fase, due gruppi `SectionLabel` (messaggio di sistema / utente) e una carta per pezzo, a destra tre gruppi divisi dal filetto verticale (`CommandRule`, `h-4 w-px bg-editorial-border`): informazioni come icone muted con `Hint` (tipo del pezzo: `Settings2` di sistema, `Braces` automatico; provenienza `PromptSourceIcon`: `BookMarked` template, `PenLine` personalizzato, niente se predefinito) | comandi (apertura della scheda `ArrowUpRight`, lucchetto) | occhio e copia; nessuna scritta lunga, nessun badge; i pezzi assenti sono una riga muted con filetto sinistro neutro, titolo in corsivo con `Hint` e motivo. Ogni testo si espande e si copia integralmente. Le fasi non usate restano visibili ma disabilitate; audit/coerenza restano disponibili in DeepL. +Gli editor prompt/descrizione usano `FIELD_MONO_CLASSNAME` nella bozza; un solo fondo tenue per la carta. + +Le Risorse linguistiche usano `TabStrip` nella fila della finestra e ricerca +`CatalogSearchField`. Ogni voce ha un unico fondo `surface-resource`, carta +calda appena distinta dalla finestra, separata dalle altre voci con spazio. +Anteprima e dettagli aperti ereditano questo fondo: niente superfici annidate, +bianco acceso o separatori fra i campi di un modulo. I campi modificabili +conservano il proprio fondo `editorial-textbox`. La classe `linguistic-resource` +usa `resource-muted` per etichette e metadati: contrasto sulla carta 5,01:1 +in chiaro e 6,13:1 in scuro; testo principale 5,37:1 e 11,10:1. Rinomina del dizionario al posto del +nome, nello stesso flex della testata, senza aggiungere un campo a tutta riga. Titoli editoriali distinguono +le voci, testo e spaziatura distinguono le informazioni al loro interno. +La variante `compact` di `Dialog` riduce testata e footer delle risorse e +finestre collegate, conservando focus, tastiera e semantica della primitiva. +Copia usa `PopoverItem` con proprietario come seconda riga e scelta segnata; +importazione usa `IconButton`, spiegazioni nei `Hint`, anteprima su carta leggera; +esportazione offre CSV/Excel affiancati. Nessun selettore o comando raw locale. +Riferimento per raggruppamento e superfici: [NN/g, Common Region](https://www.nngroup.com/articles/common-region/). + +Dizionari: icona condivisione o scudo per originale condiviso/correzioni locali, +ambito nel `Hint`; il più vicino allo scudo spiega che le nuove voci sono +condivise. Le coppie si leggono come testo affiancato; solo la voce in modifica +mostra i campi comuni. Note a richiesta. Il più inserisce in cima e porta il +fuoco al termine; spunta = termina modifica della voce, dischetto = salva il +dizionario. Più e dischetto stanno nella stessa intestazione sticky delle voci. + +Modelli di prompt: titolo, icone per ambito/flusso/modello, anteprima breve e +comando per leggere tutto. Il testo condivide la carta della voce, senza +riquadro interno: fondo sul contenitore, clamp sull'anteprima interna. Modifica +nella propria voce, creazione in cima. +Il modulo usa una griglia responsive a due colonne con `FieldLabel` e campi +comuni, senza `SettingRow` o filetti ripetuti: è un modulo, non un elenco di +impostazioni. Nome e testo hanno tutta la larghezza; ambito/flusso e +servizio/modello sono affiancati. I motivi e le spiegazioni restano nei `Hint`. + +Memorie: origine e lingue in una testata compatta, coppia affiancata, tag come +riassunto breve. Il comando dei dettagli apre una sola voce per volta. Anche +aperti, i dettagli sono compatti: provenienza a due colonne con `ResourceFact` +(icona esplicativa e valore che va a capo), misure e tag affiancati, senza +filetti interni. Numero delle misure visibile; ogni modello disponibile ha una riga propria, +con dimensione e relativo `Hint`. Durante il calcolo il comando usa `Loader2` +animato, accompagnato da uno stato accessibile. +La provenienza estesa nei Riferimenti dello Studio conserva `StatBlock` per +nomi e titoli lunghi. Ricerca realmente disabilitata con motivo durante le +modifiche; non ignorare silenziosamente la digitazione. + Quando una riga descrive una cosa su cui si può agire — una versione locale di un libro, un profilo, un file — i comandi che la riguardano stanno **su quella riga**, non nell'intestazione della sezione: nell'intestazione non si capisce su @@ -468,13 +708,19 @@ separati da «·» non si leggono. ### Pannelli modello + prompt Ogni pannello che configura una chiamata a un modello (fase di traduzione, -scheda OCR della trascrizione) ha la stessa forma: una sezione **Modello** -(bordo sinistro neutro, fornitore + modello + lucchetto su una riga, comandi -di taratura sotto) e una sezione **Prompt** (bordo sinistro verde, pillola -«Personalizzato», solo ripristino e modifica fuori dalla modifica). Nessun -testo di spiegazione fisso: il perché sta nei suggerimenti dei comandi. -L'editor prompt è uno solo, `AuditPromptEditor`, con `variant="stage"` per -questa resa; la variante predefinita resta quella del giudizio traduzione. +giudizio, scheda OCR della trascrizione) ha la stessa forma: una `PanelSection` +con il **Modello** (fornitore + modello + lucchetto su una riga, comandi di +taratura sotto, opzioni del fornitore come righe con interruttore) e una sezione +**Prompt** (carta tenue e bordo verde; espandi e modifica fuori dalla bozza, conferma/annulla nella bozza). Nella configurazione della pipeline la sezione +Modello è una sola, `ModelSection`, per fasi e giudizio. Nessun testo di +spiegazione fisso: il perché sta nei suggerimenti dei titoli, delle righe e dei +comandi, e un comando spento dice il motivo («Modifica prompt — esistono già +traduzioni», «Rifinisci… — manca la chiave di X»). Le icone di categoria della +taratura (ragionamento, temperatura) sono muted dentro un `Hint`, mai in ocra. +La pipeline usa `PipelinePromptEditor`; OCR mantiene `AuditPromptEditor`. La spiegazione la porta il titolo della carta (`Hint` con `children`, come `SectionLabel`), nessuna «i» separata; vista di lettura fino a 12 righe. Accanto al titolo `PromptSourceLabel`: Predefinito / Personalizzato / Template «nome», riconosciuto confrontando il testo (`describePromptSource`); niente pillola «Personalizzato». Campi di modifica dei prompt alti 16 righe. Un suggerimento Radix non intercetta mai il puntatore (regola CSS sul contenitore) e quello sul nome della pipeline resta chiuso a menu aperto. I modelli salvati stanno in +due `ClickPopover` (`PromptTemplateMenus`), il libro con `CatalogSearchField` e +`PopoverItem`, il segnalibro con `RenameField`; niente eliminazione lì, si +elimina nelle risorse linguistiche. La conferma applica testo ed eventuali impostazioni del modello insieme; annullare scarta entrambi. Le scelte di taratura sotto il modello (livello di ragionamento, immagine inviata dall'OCR) sono cerchietti da 24 px con icona e suggerimento, preceduti da un'icona di categoria. Il comando principale di un pannello che si chiude @@ -486,6 +732,29 @@ riapertura. Tre zone stabili: contesto a sinistra, stato centrale, comandi globali a destra. Un'informazione non cambia posizione passando tra sezioni. +Lo stato del salvataggio di uno Studio (traduzione o trascrizione) vive qui, +in fondo a destra: pallino e parola, suggerimento con l'ora dell'ultimo +salvataggio e un messaggio tradotto in caso di errore; dettagli tecnici nel log. Nella testata del foglio resta solo il +dischetto, spento senza niente da salvare (motivo nel suggerimento se è +bloccato), `danger` con «Riprova» dopo un errore. La barra è una regione +`aria-live`: si annuncia solo l'errore, mai «da salvare» o «salvato». + +### Verifica di un testo + +Verificato = `CircleCheck` in un `IconButton` accanto al titolo del foglio, +`success` quando acceso, `ariaPressed`; mai un lucchetto. Il lucchetto non +serve a dire «non modificabile»: lo dice il comando che rende modificabile +(matita accesa o spenta). Uno stato di cautela («da aggiornare») si affianca +alla spunta, non la sostituisce, ed è ocra, non oro (oro = lavoro in corso). + +### Comandi nel margine della pagina + +Due pagine affiancate (originale e traduzione) devono restare in linea: una +testata non cresce per ospitare comandi in più. I comandi di vista di una +pagina (fasi, confronto) stanno in colonna nel margine destro della pagina, +`IconButton` xs con suggerimento a sinistra, fermi mentre il testo scorre; il +margine è uguale sulle due pagine anche dove la colonna non c'è. + ## Accessibilità - Focus visibile: `focus-visible:ring-2 focus-visible:ring-editorial-accent`. @@ -541,3 +810,14 @@ Niente colori neon o valori locali. quelle esistenti. Il riferimento visivo live è nella guida di stile interna dell'app. + + +Risorse Memorie: elenco piatto, dettagli compatti con ResourceFact, Select e +IconButton neutri. Bozze protette durante cambio scheda, +filtri e chiusura. Le correzioni ai testi riguardano la memoria: evidenza iniziale +congelata e riga esplicita dopo modifica dell’originale. Costi su Popover comune, +nessun pannello con portal/posizionamento/timer propri. Opzioni vista su righe +comuni; importazione a vuoto con IconButton. Tab indisponibili aria-disabled, +focusabili per il motivo e mai attivabili. Dettagli tecnici di errore nel log. + +Le ricerche non disponibili durante modifica sono disabilitate con motivo accessibile. Nelle provenienze, nomi e titoli lunghi vanno a capo con `StatBlock`, senza invadere la colonna. diff --git a/docs/en/guides/annotations.md b/docs/en/guides/annotations.md index c529321a..62e3eb91 100644 --- a/docs/en/guides/annotations.md +++ b/docs/en/guides/annotations.md @@ -17,17 +17,17 @@ note can be edited or removed without rewriting the translation. | Problem | An error requiring action | | Approved | A note recording the outcome of review | -The Approved type does not replace **Lock translation**. Annotations describe -review work; locking controls whether a segment can be reprocessed. +The Approved type does not replace the **Mark as verified** check. Annotations +describe review work; verification controls whether a segment can be reprocessed. ## Creating an annotation Select a passage in the translation and choose **Add annotation** from the context menu. The selected text becomes the note’s anchor. You can also add -an unanchored note in the segment’s **Notes** tab or convert an audit finding +an unanchored note with **+** in the **Notes** sub-tab of **Review** or convert an audit finding into an annotation. -Segment notes are in the project sidebar. They are separate from notes about +Segment notes are in the studio’s Tools column. They are separate from notes about a bibliographic work, which belong to its Library record. ## Display and export diff --git a/docs/en/guides/audit-review.md b/docs/en/guides/audit-review.md index 019fad94..2dda69f6 100644 --- a/docs/en/guides/audit-review.md +++ b/docs/en/guides/audit-review.md @@ -31,10 +31,10 @@ assessments; schema compliance concerns the response format. 3. Edit the text manually or rerun the relevant stage. 4. Use **Re-evaluate** to update the assessment without translating again. 5. Record decisions and unresolved questions in **Notes**. -6. Lock the translation when review is complete. +6. Mark the translation as verified when review is complete. An audit finding can be converted into an annotation. Passage lookup uses -text supplied by the model and may not find the exact location. Locking a +text supplied by the model and may not find the exact location. Verifying a translation is a reviewer decision, separate from the automated rating and annotation type. @@ -43,8 +43,8 @@ annotation type. After completing the segments, run the consistency check. It examines translations with neighbouring translated segments as context, without comparing them with the source. It uses the dedicated prompt under -**Quality Control** and displays results in the Insight panel’s -**Coherence** tab. +**Quality Control** and displays results in **Document** → **Coherence**, in +the Tools column. This check can identify terminology or style variations between passages. It does not replace a segment-level accuracy assessment. diff --git a/docs/en/guides/context-and-caching.md b/docs/en/guides/context-and-caching.md index 6931c823..fcdf659c 100644 --- a/docs/en/guides/context-and-caching.md +++ b/docs/en/guides/context-and-caching.md @@ -23,7 +23,7 @@ rather than the original text. For translation and revision, the system message preserves this order: -1. Static instructions: persona, structural rules, glossary and examples. +1. Static instructions: role and Translation context, structural rules, glossary and examples. 2. Shared document context. 3. Stage-specific instructions, including any selected memory references. @@ -57,7 +57,8 @@ same document, without the image ever breaking the reusable prefix. ## Configuration and inspection -Anthropic caching is disabled by default. Enable it when you expect to reuse +Anthropic caching is disabled by default; its switch sits under the model of +each Anthropic stage, in the Stages tab. Enable it when you expect to reuse a prefix, and consider extended retention in light of the interval between requests. Cache writes may incur a charge, so a prefix that is never reused does not necessarily save money. diff --git a/docs/en/guides/document-pipeline.md b/docs/en/guides/document-pipeline.md index c7839d96..d96cbdd1 100644 --- a/docs/en/guides/document-pipeline.md +++ b/docs/en/guides/document-pipeline.md @@ -4,6 +4,10 @@ title: Translating a document # Translating a document +During execution you cannot switch pipelines or create a new one. Memory, +Review and Document remember the selected subtab when you visit another tab; +if that view is unavailable, an available view is shown instead. + Translation operates on text segments, also called *chunks* in the interface. Each segment retains its source text, stage outputs, editable translation, assessment and annotations. A document containing only one segment uses the @@ -19,16 +23,70 @@ Check the boundaries before confirming: they define the units used for translation and review. See the [import and export reference](../reference/import-export) for formats, size limits and imported footnotes. +## Languages of the work + +Languages belong to the translated work, not to its pipelines: they apply to +every pipeline of the work, which no longer have a language pair. The +translation prompt is guided by the **Translation context**; only DeepL keeps +its own pair in the options of its stage, because the service requires codes. + +Source and Target each have three optional fields: + +- **Language**: an ISO 639-3 code from the official SIL register, the one used + by library catalogues, including historical languages (Latin, Ancient Greek, + Old and Middle French, Old Provençal/Occitan, Old Spanish, Old and Middle + English, Old and Middle High German, Anglo-Norman and others). Search by + Italian name, English name or code; with an empty search the list shows + “Already used in the workspace” and “Historical languages”, and typing + searches all roughly 7,900 languages. Names are in Italian where a standard + Italian name exists, otherwise the official English name. +- **Variety**: a Glottolog variety, offered only among the varieties of the + chosen language (Latin: Late Latin, Medieval Latin, Vulgar Latin; Italian: + Old Italian, Fiorentino, Laziale, Cicolano-Reatino-Aquilano). It stays + disabled until you choose a language. +- **Note**: free text for what codes cannot say (period, area, hand), for + example “Po Valley vernacular, 15th c.”. + +Every field can be left empty (“not specified”) and has an X to clear it. A new +work starts with the source not specified and the target Italian. + +Limits: the standards have no code for Medieval Latin, Old Italian or Old +Catalan as languages; Glottolog lists no “classical Latin” variety; regional +Italo-Romance languages (Venetian, Lombard, Ligurian, Neapolitan, Sicilian…) +are separate languages with modern varieties only; there are no medieval area +varieties, so use the note. For a 15th-century vernacular fencing treatise: +Italian + Old Italian + note, or the regional language + note. The language +list ships with the app and works offline; **Settings → Languages** shows the +list in use and downloads it again from the official sources, keeping languages +that are gone as retired. + +Where to set them: in the import window and in the top row of the Studio. +There, on the right before the commands, the pair appears with its varieties +(“Italian (Old Italian) → English”) and the tooltip adds the notes. The languages +icon next to it opens the **Languages of the work** window with the same fields, Cancel +and Confirm: nothing is saved unless you confirm. While the pipeline is +running the control is visible but blocked, with the reason in the tooltip. If +the work has no source language and comes from a Library book whose language +matches a known language, the source is proposed pre-filled (still to be +confirmed); the import window pre-fills it the same way. + +After every Confirm, if memory holds phrases saved from this work with +different languages, the app asks whether to give them the work’s languages: +text and similarity measures stay the same and the previous version stays in +the history. This also aligns phrases saved earlier. See +[phrase memory](./phrase-memory). + ## Configuration -Open pipeline configuration and set the languages, mode, providers, models -and instructions. Pipeline modes define these sequences: +Open the pipeline configuration (the gear in the Studio's top row) and set the +Translation context, mode, providers, models and instructions; languages belong to the work (see below) and only DeepL has a pair of its own: its tabs are described in +[pipeline configuration](../reference/pipeline-config). Pipeline modes define these sequences: | Mode | Processing | | --- | --- | | Standard | Translation and automated assessment | | Editorial | Translation, draft revision (*Refine*), formatting (*Format*) and assessment | -| DeepL Hybrid | DeepL translation, optional LLM revision and LLM assessment | +| DeepL Hybrid | DeepL translation, LLM refinement and LLM assessment | Each LLM stage has an independent provider and model selection. DeepL requires its own API key and does not act as the evaluator. The pipeline mode cannot @@ -44,24 +102,95 @@ count limits the group to process. Processing advances through the segments and updates their states. Cancelling stops the current work without removing completed results. Resuming and rerunning serve different purposes: resuming processes outstanding work, while rerunning -recalculates the unlocked segments covered by the selected action. +recalculates the unverified segments covered by the selected action. +If models, stage or audit prompts, the Translation context or DeepL options change after the +interruption, resuming warns that the configuration is no longer the one the +work started with. + +While a segment is being translated, its translation text is covered by a gold +veil, “Translation in progress…”, and cannot be edited: the text does not +appear as it is generated, it arrives when the stage finishes. The stage column +in the margin stays usable; the original is read-only and the pencil says why. ## Reading and review -The document view places source and translation side by side. The segment -sidebar contains **References**, **Preview**, **Audit**, **Memory** and **Notes**. -The **Insight** panel provides the document index, search, statistics, -consistency review and glossary. +The translation studio opens inside the application frame: the main bar on +the left stays visible and leads to any area, closing the translation. While +the pipeline is running its entries are off, like the way back to the +catalogue. + +At the top, the bar shows work / pipeline and Simple, Editorial or DeepL mode. Long names are truncated; hover reveals the full name. Click the work name to rename it. The pipeline name, with its small arrow, opens the menu to select, create, rename or delete; the gear opens options. Next to them the languages of the work appear in short form (see [Languages of the work]). Import, export, language resources and deletion are on the right. + +In the middle the two sheets place source and translation side by side. Above +them, on the left, the number of the open segment; in the middle a window of +seven dots, one per segment with its state: the open segment stays still under +the central mark while the others slide to the sides. The single arrows move +to the neighbouring segment, the double ones jump by seven; the mouse wheel +over the dots also scrolls through segments, and clicking a dot opens it. To +the right of the dots, the stage indicators show where the open segment stands +(translation, revision, formatting, audit): a click opens that stage's detail. +The magnifier next to them opens the whole-document search below the row; a +result leads to its segment, Esc closes it. + +On the right, the **Tools** column starts with execution — translate, the **Multiple chunks** switch with the number of +segments to process (always visible, off when translating a single segment) — +and costs, then the tabs, in this order: **Glossary**, +**Memory**, **Preview**, **Review** and **Document**, which gathers the +whole-document summaries as three sub-tabs: **Index**, **Statistics** and +**Coherence**. Memory gathers two sub-tabs: **similar phrases +in memory**, to use while translating, and **Extract phrases**, which saves the +segment's pairs; extraction turns on once the segment is translated. Review gathers three sub-tabs, +**Audit**, **Notes** and **Source notes** (the footnotes imported with the +source, shown only when the segment has some): icon tabs with name and count +in the tooltip, each with its own list; it opens on the audit when it has open findings, +otherwise on the notes. Audit turns on once the segment is translated, +Glossary once a glossary is assigned; the reason stays in the +tooltip. Collapsed to icons, the column +keeps only the translate button visible, or stop while running. Intermediate stage outputs help identify where a change was introduced. +Their controls sit in a column in the right margin of the translation page, +next to the scrollbar: at the top the stages in pipeline order, then +comparison and the pairs to compare; the one you are viewing is highlighted. After a manual edit, **Re-evaluate** runs the quality assessment alone. -**Lock translation** protects an approved result from reprocessing. If the -source text changes, the interface flags the translation for updating. +If you correct the source with the pencil, the ochre **Source changed** label +appears next to the translation title and the segment’s circle gets an ochre +mark: the translation needs updating. + +The check next to the **Candidate translation** title marks the translation as +verified: it turns green, the text locks and rerunning unverified segments +skips it. Verifying also clears “needs updating”, because it means you checked +it against the current source; the same control returns it to draft. The check +is off while the segment is being translated or when there is no translation +yet. Every disabled control gives its reason in the tooltip. + +Current limit: “needs updating” is not kept when the translation is closed; +reopening it, the mark is gone. + +### Segment history + +**Review → History** lists the open segment’s versions, newest first, with the +author (**Pipeline** or **Manual**), date and time, and the **Current** and +**Verified** marks. A version is written at every pipeline pass (including the +rewrite after the audit), at every save with the disk or `Ctrl + S` if the +segment’s text changed since its last version, and on verification when the +verified text differs. Automatic saving writes no versions, so the history does +not fill up at every pause. + +Restore puts a version’s text back on the page and writes it as a new version: +the earlier ones stay. It is off on a verified translation (return it to draft +first) and while the segment is being translated. Current limits: versions +cannot be deleted or named, and the history does not show the model used, +because a pipeline uses more than one. History is per pipeline and is lost if +the document is split again into different segments. ## Request previews -Pipeline configuration shows the structure of the prompts. The segment’s -**Preview** tab builds the selected stage’s request for the current text. +The **Prompt preview** in pipeline options has two modes. **Structure** shows +the pieces of each request with placeholders where data goes. **Open chunk** +fills them with the chunk open in the Studio: text, neighbouring chunks, +checked memory phrases, previous translation. Audit and Coherence use the +chunk’s current translation and say so when there is none. This action does not call a model or generate a translation. ## Export @@ -70,3 +199,24 @@ Check incomplete segments before exporting: standard formats may use the source text where a segment has no translation. Bilingual export explicitly distinguishes the source from a missing translation. See [export formats and contents](../reference/import-export). + + +## Saving, navigation and cost details + +While translation is running, the Studio remains open. Wait until it finishes +or stop processing before leaving. The save icon and Ctrl/⌘+S save changes; +after a failure the command offers **Retry**. User feedback is translated; +technical details remain in the log. History loading and restoration follow the same rule. + +Unavailable tabs remain reachable with Tab so their labels explain why they +cannot be opened. Arrow keys skip them. Circular selectors remain reachable +even when the currently selected option is unavailable. + +In the Tools column, under the commands, one line shows the next run estimate +and what was spent on the open segment. A click opens the cost panel: two tables +with one row per stage (model, tokens, cost). The estimate follows the selected +mode and block count and is indicative. The spent table also shows model +**calls** per stage: a re-translated segment or a review loop adds calls, so the +number does not count runs. The block count does not represent repeated +translations. The Statistics tab shows the same per-stage table for the whole +document. diff --git a/docs/en/guides/glossary-and-memory.md b/docs/en/guides/glossary-and-memory.md index 8a011393..b3246450 100644 --- a/docs/en/guides/glossary-and-memory.md +++ b/docs/en/guides/glossary-and-memory.md @@ -15,19 +15,48 @@ dictionaries, prompt templates and phrases. The dictionary tab lets you create, rename, duplicate and delete resources, edit entries and import CSV or TSV data. -Import provides a preview and a choice between merging data and replacing -the contents. Check the source, target and notes field mapping before -confirming. Replacement removes the dictionary’s previous entries. +Import from resources creates a new dictionary in the selected workspace, +named after the file. The icon command opens the file picker; the preview +shows the first entries. For Excel, you can map source, translation and notes +columns before confirming. ## Sharing and local overrides +**General language resources** offer the same management. The workspace +filter shows every linked dictionary, including shared dictionaries; it also +supports all dictionaries or those without a workspace. This filter narrows +the list: general resources always read and edit originals. Select a +destination workspace to create, import or copy a dictionary. Search checks +dictionary names. + +In an open dictionary the sharing icon identifies the **Shared original**; +the shield identifies **Local corrections**. Its hint explains where changes apply; in guest +workspaces the adjacent plus explains that new entries are added to the shared +original. Existing source terms cannot be renamed through a local override. + +Entries show source and translation side by side, without opening fields for +the entire list. The pencil edits a pair; the notebook shows its notes. +**+** inserts an entry at the top and focuses the term, making it immediately +visible. The check finishes editing the entry: changes still need to be saved +with the dictionary’s disk command. + +Plus and the disk remain in the entry header while you scroll. +The disk saves edited entries. Copy lists dictionaries with their source +workspace and marks the selected one; export offers CSV and Excel side by side. **Save and close** respects the same scope; a +failed save keeps the window open. **Close without saving** discards changes. +Rename and delete apply to the shared original; deletion requires +confirmation. CSV and Excel exports contain saved original entries. + A dictionary can be linked to multiple workspaces. Links share the same resource; a copy creates an independent dictionary. Workspace overrides and exclusions change the local view of entries without altering the shared original. Use the assignment control to attach a dictionary to the project. The -**Glossary** tab in the Insight panel shows the full assigned glossary; -pipeline configuration provides access to its terminology register. +**Glossary** tab in the Tools column shows the full assigned glossary, with the +number of terms in its title; the highlight command colours the terms in the +sheets and, while on, shows the colour legend; +in the Glossary tab of the pipeline configuration you assign the dictionary and +edit its terms, saving them with the disk icon. ## Use during translation @@ -51,6 +80,27 @@ Highlighting identifies textual matches; it does not interpret context or replace linguistic review. A missing target term may require a correction or a justified variant recorded in the glossary notes. +## Prompt templates + +The **Prompt Templates** tab searches names and prompt text and filters by +Stages, Quality control, Translation context, System, Memory, or OCR. Templates are shared across +the application without belonging to a workspace. Use **+** to create a +template or its pencil command to edit it. + +The form stores a name, scope, workflow, text, and an optional default provider +and model. Label hints explain each field. Select a provider and model to +refine the text; disabled commands explain missing requirements. The disk +saves and the cross cancels. A name already used in the same scope and +workflow requires editing that template or choosing a different name. The +trash command deletes only after confirmation. + +Each template has a distinct title and a short preview on the same muted surface as the entry. The eye opens the +full text; the collapse command returns to the preview. Icons beside the title +explain scope, workflow and model on hover or when pressed. The pencil opens +the form within the same entry; search and filters stay disabled while editing +to preserve the draft. The form pairs scope with workflow and provider with +model, without separators between fields. + ## Glossary, memory and examples | Resource | Role | @@ -61,3 +111,5 @@ or a justified variant recorded in the glossary notes. See [Phrase memory and examples](./phrase-memory) for extraction, selection and saving of bilingual pairs. + +Dictionary and prompt lists distinguish each entry with one muted background shared by its text and expanded details. The dictionary title pencil replaces the name in place, keeping commands on the same row; Enter saves, Escape cancels. diff --git a/docs/en/guides/keyboard-shortcuts.md b/docs/en/guides/keyboard-shortcuts.md index ea924b63..446803c9 100644 --- a/docs/en/guides/keyboard-shortcuts.md +++ b/docs/en/guides/keyboard-shortcuts.md @@ -11,7 +11,8 @@ field, select control or editable text area. | Shortcut | Action | Conditions | | --- | --- | --- | | `Ctrl + Enter` | Starts the selected translation action | Segment mode requires a selected segment; also works inside text fields | -| `Ctrl + S` | Saves the project and modified language resources | Requires an existing project; project saving is deferred while processing | +| `Ctrl + S` | Saves the open translation (with a history version of changed segments) and modified language resources | With a translation open it also works while typing on the pages; the translation is not saved while processing | +| `Ctrl + S` in the transcription Studio | Saves a version of the page to the history right away | Also works while typing on the page; does nothing when there is nothing new to save | | `Ctrl + E` | Opens export | Requires at least one segment | | `Ctrl + ,` | Opens pipeline configuration | Outside editable fields | | `Ctrl + H` | Opens this section of the in-app guide | Outside editable fields | diff --git a/docs/en/guides/llm-and-pipelines.md b/docs/en/guides/llm-and-pipelines.md index a491d00e..411cc143 100644 --- a/docs/en/guides/llm-and-pipelines.md +++ b/docs/en/guides/llm-and-pipelines.md @@ -28,8 +28,9 @@ context, terminology, examples and references must be included in the request. | Judge | Source, translation and assessment criteria | Assessment and structured findings | | Coherence | Translations and neighbouring translated segments | Findings about consistency across segments | -Format uses a separate prompt. It does not receive the translation persona, -glossary or source context. Its instructions limit it to formatting repairs, +Format uses a separate prompt. It does not receive the Translation context, +glossary, memory phrases or source context. The glossary appears once, in the +rules at the start of the translation and Refine instructions. Its instructions limit it to formatting repairs, but the output still requires review. In DeepL Hybrid mode, the initial stage uses the DeepL API and its language, diff --git a/docs/en/guides/phrase-memory.md b/docs/en/guides/phrase-memory.md index 8453907b..375a53af 100644 --- a/docs/en/guides/phrase-memory.md +++ b/docs/en/guides/phrase-memory.md @@ -4,8 +4,11 @@ title: Phrase memory and examples # Phrase memory and examples +While editing text or metadata, search stays disabled with its reason. +Long provenance titles wrap. + Phrase memory stores source text and approved translation pairs for reuse -within a workspace. Finding matches, selecting them for a prompt and saving +in translations. Finding matches, selecting them for a prompt and saving new phrases are separate operations. ## Retrieving references @@ -14,8 +17,16 @@ When enabled, Glossa searches for matches for the document’s segments using resources available to the workspace. This search changes neither translations nor saved phrases. -The **References** tab displays matches and lets you adjust the similarity -threshold. Only selected pairs are included in the next request for that +The **References** sub-tab of the **Memory** tab displays matches. At the top +there is a single row: the similarity threshold (slider, or − and + one +hundredth at a time), the globe and refresh. Each result shows the original in +book type and the translation below, with the language code in the margin; at +the top, next to the similarity, the phrase's language pair with its varieties +(“Italian (Old Italian) → English”); below, one line with work and chunk; the “i” icon adds workspace, book and model. +Each result says where it comes from: “this document” +(highlighted), “another work in the workspace” or “another workspace”. The +order is this document first, then the workspace, then the rest; within each +group the highest similarity comes first. Only selected pairs are included in the next request for that segment. If matches exist but none are selected, starting translation warns that those references will not be used. @@ -26,23 +37,91 @@ current context. ## Creating and reviewing phrases -1. Review the translation and lock the segment. -2. Open **Memory**: previously saved pairs are loaded and selected. +1. Review the translation and mark it as verified. +2. Open **Memory** → **Memory**: saved pairs are loaded read-only. 3. Use **Extract phrases** to generate proposals, or add pairs manually. 4. Edit the text and select the pairs to retain. -5. Save to apply the selection. +5. Use the disk command **Add checked pairs to memory**. -Extraction does not save automatically. Deselecting a previously saved pair -and saving removes it from the collection. Unconfirmed edits remain in the +Extraction does not save automatically. Saving only adds selected new pairs, +preserving existing ones. To remove a saved pair, use its own trash command +and confirm. Unconfirmed edits remain in the segment’s draft when you switch segments during review; this is not a permanent save. -## Scope +## Scope and compatibility + +New saved phrases take the language, variety and note of the work they come +from; automatic phrase-pair extraction tells the model the work’s real +languages, with variety and note. After every Confirm in the +[Languages of the work](./document-pipeline) window, if +memory holds phrases from this work with different languages, the app asks +whether to give them the work’s languages: text and similarity measures stay +the same and the previous version stays in the history. Extracted phrases follow the current workspace of their source translation. +Provenance separately preserves the workspace at extraction time. Viewing a +phrase from another workspace does not create links or copies. + +The globe icon next to the References refresh button extends the search to +other workspaces and unassigned phrases; the choice is remembered per workspace +and is no longer in the workspace settings. Only the target language filters: +phrases translated into a different target language are not suggested; the +source language does not filter. If the work has no target language, no +language filter applies. Search always requires the same embedding model, +dimensions and input profile. Texts without the requested embedding remain +in the catalogue but are excluded from similarity results. Different models +are never compared. Already saved chunk pairs are not inserted as zero-distance matches. + +## Texts, provenance and multiple embeddings + +Each pair links to a textual unit with revisions of its source and translation. +Adding a model preserves embeddings from other models without duplicating the +pair. Equal texts from different chunks remain distinct units. + +When creating a translation, select **Source book and version** from the Library (searchable by title or copy). +Without this explicit choice, the book remains **Not specified**; the filename +does not determine it. Extracted phrases record the book, version, translation, +chunk and original workspace. Unknown page numbers are never invented. + +## Managing the collection + +Open **Linguistic resources → Memories**. General resources start with all +phrases; workspace resources start with its collection. Filter by workspace, +**All**, **Unassigned** and tag, or search texts and tags. -Extracted phrases retain their link to the originating translation. Moving -that translation to another workspace moves its phrases with it. Imported, -linked resources can be shared through workspace links. The **Phrases** tab -in Language Resources provides access to the collection. +Each entry shows source and translation side by side. Languages with varieties, an origin +title and tags help identify it without opening details. **Provenance, tags and +measurements** opens the full information for the selected entry: book, +workspace, translation, chunk, date, tags and embedding models. + +Select a model in the details and use plus to add its embedding or the circular +arrow to recalculate it. Calculation requires an OpenAI key and incurs costs, +described in the confirmation. Other models are preserved. Search and filters +are disabled while editing text or tags to preserve the draft; finish or cancel +the edit to use them again. +Workspace settings can calculate the selected model for all its memory entries. +Saving settings applies the active search model independently of calculation. + +The pencil edits both memory texts. Editing only the translation does not +recalculate source embeddings. Editing the source recalculates every available +embedding after cost confirmation: text and vectors are saved together, or +nothing changes if a calculation fails. Earlier revisions remain archived; +search uses only the current revision. These edits do not change the source document. + +**Text tags** are reusable manual labels, separated by semicolons. They are +not automatic classifications. After confirmation, the bin removes the pair, +its revisions and embeddings. The dictionary command copies it into the +selected dictionary; save pending dictionary edits first. **CSV export** +includes text, tags, provenance and models for visible phrases only. The app +backup also preserves revisions and vectors; CSV does not replace it. + +The underlying structure supports variable-length units and untranslated texts. +Selecting pages or sections, automatic classification and semantic corpus +analysis are future work, not tools currently available on this screen. + +Deleting a translation preserves archived texts, book and provenance. +They become unassigned entries and indicate that the translation is unavailable. +After a source edit, provenance still describes the initial extraction; a row +identifies the edit without silently reassigning the citation. ## Style examples @@ -50,10 +129,22 @@ Translation examples are complete source and translation segment pairs used to guide a pipeline’s register and style. They are not retrieved according to similarity with the current segment. -For a locked segment, **Use as a style example** in the Audit tab adds the -pair to pipeline settings, where it can be edited or removed. The limit is +For a verified segment, **Use as a style example** in the Audit tab adds the +pair to the Memory tab of the pipeline configuration, where it can be edited or removed. The limit is five examples. Since they form part of the static context, their length contributes to request size. Use the [glossary](./glossary-and-memory) for mandatory terminology and memory references for wording relevant to an individual passage. + +## In the Translation Studio + +In the Memory tab, **References** shows similar phrases already in memory: for +each, under the pair, where it comes from (workspace, translation, segment, or +“imported”). The circled check decides which to use in the translation. +**Memory** opens only once the translation is verified: saved pairs are removed +one by one with the bin; new ones are checked and added with the disk, which +never deletes the others. The search excludes phrases translated into a +different target language than the work’s. + +Collection entries have one muted background distinct from the window; expanded details share that same surface. Details list embeddings one per row; the calculation command shows a spinner until the request completes. diff --git a/docs/en/guides/projects-and-workspace.md b/docs/en/guides/projects-and-workspace.md index 3be569a7..968bd4f5 100644 --- a/docs/en/guides/projects-and-workspace.md +++ b/docs/en/guides/projects-and-workspace.md @@ -23,15 +23,95 @@ produces an independent dictionary. ## Creation and saving -You can create projects from a workspace page or the **Translations** area. -Translations lists projects across all workspaces. A workspace’s name, -description and icon identify it throughout the application. - -Autosave operates on projects that have already been created. It detects changes -and saves them after a short idle interval; while processing is active, it waits -for a stable state. The status bar distinguishes unsaved changes, saving, -successful saves and errors. `Ctrl + S` requests a manual save, subject to the -conditions described under [keyboard shortcuts](./keyboard-shortcuts). +You can create projects from a workspace page or the **Translations** area: +the «+» next to the Translations title opens a window asking for the name, the +workspace and, optionally, the file to translate, in the same formats as the +editor import. The file is read as soon as you choose it: if it cannot be read +(a scanned PDF with no text, a file not in UTF-8) the reason appears under the +field and nothing is created. **Create** opens the translation in the editor +with the import preview, where you choose languages and segments; closing the +preview leaves the translation empty, and you can import the file later from +the editor. Under **Source book and version** the row shows the chosen book title +(truncated; the full title and copy appear in the tooltip): the book icon opens +a searchable list (search by title or copy) where each entry shows the title +and, below it, the copy; the X removes the book. The window does not widen with +long titles. + +The import preview is the app’s standard window, large and almost full height. +At the top, the file name and title. On the left, in a column that scrolls on +its own: **Languages of the work** (Source, then Target), **Model** (provider and +model) and **Splitting into chunks** (automatic splitting, headings for +Markdown, carrying trailing short blocks, size presets, recalculate). On the +right, the word, paragraph and chunk counts, the choice between card view and +segment view (small round icons) and the chunk preview at full height. At the +bottom, the result of the word-count check, Cancel and Import. A workspace’s name, description and icon identify it throughout +the application. + +## The Translations catalogue + +The Translations area gathers the translations of every workspace and is laid +out like the [Transcriptions catalogue](./transcription#the-transcriptions-catalogue): +shelves on the right, search and quick filters above the list, three views +(list, covers, table). + +- **Row**: name in italics, below it the names of the source and target languages, then + workspace, translated segments out of the total, verified segments and the + completion bar, green when every segment is verified. +- **Shelves**: All, Recent (edited in the last 30 days), Not started (no + segment translated), In progress, Verified (every segment verified). +- **Quick filters**: workspace (kept when you leave the page and come back) + and language pair; sort by name, last edit or progress; group by workspace + or language pair. +- **Row commands**: rename and delete; in the cover and table views they sit + in the three-dot menu. Clicking a row opens the editor. + +Current limits: the counts cover the project's first pipeline, the one the +editor opens; a translation is not yet linked to the work or transcription it +starts from, so there are no commands to open them, no library or century +filters and no archiving. + +A translation saves itself shortly after the last edit, always as a whole: +source text and every segment. During automatic translation saving waits for +the end, because the pipeline saves by itself. The status bar, bottom right, +distinguishes unsaved changes, saving, saved and error; its tooltip gives the +time of the last save and, after an error, the reason. + +To save right away there is the disk at the top of the translation page (on +the source page when only the source is open), or `Ctrl + S`, which also works +while typing on the pages. The disk also writes a version to the +[history](./document-pipeline#segment-history) of every changed segment. It is +off when there is nothing to save and no new version to write, and during +automatic translation; if a save fails it turns red and clicking it +retries. + +Leaving the translation — back to the catalogue, main bar, path at the top, +switching workspace — saves before closing. If that save fails, the +translation stays open with the error in view: no edit is lost by leaving. One +limit remains: closing the Glossa window an instant after the last edit can +lose it. + +## Left bar + +The left bar holds navigation: at the top the Dashboard with Overview and +Search, then the areas in working order (Library, Transcriptions, Translations, +Analysis) and finally the workspaces, with the plus to create one. A rule +separates each group; an item’s explanation appears on hover. + +At the bottom is the general menu: save all, general language resources, +settings, help and interface language. It is always available, even without +an active workspace. + +The icon at the top collapses the bar: the item icons and the general menu +remain, and a click opens the page without expanding it. The Glossa mark at +the top expands it again. + +## Areas and their ink + +Library, Transcriptions and Translations each have their own ink — petrol, +sepia and indigo — so you can tell them apart at a glance: the area icon in the +left sidebar, a short rule under the large title and a slightly different +background paper. State colours stay the same in every area: green marks what +is selected or active, red errors, ochre cautions, gold running jobs. ## Dashboard diff --git a/docs/en/guides/storage-and-jobs.md b/docs/en/guides/storage-and-jobs.md index 6e8d45f8..12cd33e4 100644 --- a/docs/en/guides/storage-and-jobs.md +++ b/docs/en/guides/storage-and-jobs.md @@ -16,7 +16,7 @@ screens to remain open. | File repository | Manifests, images, thumbnails and local versions | Selectable directory, including an external drive | | Network cache | Reusable search responses and images | Temporary storage with a configurable limit | -View and change locations under **Settings → Data**. Changing the data +View and change locations under **Settings → Data → Folders**. Changing the data directory copies the database, checks the copy and records the new location for the next restart; the original is not deleted automatically. Changing the file repository selects an empty directory or reconnects an existing @@ -108,7 +108,7 @@ the job reports an error and retains the pages produced successfully. The network cache reuses responses and images. Its default size limit is 512 MB, and search responses are valid for 24 hours by default. Images are subject to the size limit without the same time-based expiry. Change these -values or clear the cache under **Settings → Data**. Cached data does not +values or clear the cache under **Settings → Data → Network cache**. Cached data does not increase downloaded-page counts and is excluded from [backups](../reference/backup-and-restore). diff --git a/docs/en/guides/transcription.md b/docs/en/guides/transcription.md index d1331f14..49cd9042 100644 --- a/docs/en/guides/transcription.md +++ b/docs/en/guides/transcription.md @@ -8,9 +8,37 @@ The Transcription Studio is the focused mode for writing and correcting a document's text: viewer on the left, text in the middle, tools on the right. It follows the same idea as Translation, applied to transcription. -## Creating a document - -From the **Transcriptions** area, choose "New document" and give it a title. +## The Transcriptions catalogue + +The **Transcriptions** area gathers the transcriptions of every workspace and +is organised like the Library: shelves on the right, search and quick filters +above the list, three views — list, covers, table — at the end of the title +row. + +Each row shows the transcription name in italics and, below it, the +transcribed work: author, year, place and printer, title. When the name is the +work's own title, which is the name suggested at creation, it is not repeated. +At the end of the row are the workspace, the pages written out of the work's +total, the verified pages and the completion bar. A transcription with no +linked work has no total: it only counts the written pages. + +The shelves are **All**, **Recent** (edited in the last 30 days), **Not +started**, **In progress**, **Verified** (every page verified), **No work** +(started from scratch) and **Archived**; archived transcriptions only appear +on their own shelf. Search looks at the name and at every field of the work. +Quick filters narrow by workspace, library and century of the work, with the +count next to each value; the list sorts by name, author, year, last edit or +progress, and groups by workspace or library. + +Clicking a row opens the Studio. Hovering a row shows its commands: open the +work in the Library, rename — the name turns into a field, Enter saves, Esc +cancels —, archive or restore, delete. In covers and table view the same +commands are in the three-dots menu. Selecting several transcriptions at once +is not available yet. + +## Creating a transcription + +From the **Transcriptions** area, use the **+** command next to the title and give it a name. You can also link it right away to a work already in the Library, searching by title in the same dialog — optional: without it, the document has no viewer, a single block of text. From a work's page in the **Library** you @@ -34,22 +62,29 @@ viewer, not an error, and it stays a single block of text. At the top, a document tied to a work presents it as the Library page does — author, year, place, and printer above, title below — with the link out to the library's site. -The three-dot command offers only "Remove transcription": downloading, -verifying, or archiving the work stay commands of the Library page, not of -the Studio. +On the right, the bin deletes the transcription after confirmation: +downloading, verifying, or archiving the work stay commands of the Library +page, not of the Studio. ## Writing and saving -Text saves after 30 seconds without changes. The indicator in the top right -distinguishes unsaved text from a save in progress, completed, or failed, with -a command to retry on failure. Changing page or leaving the Studio normally +Text saves after 30 seconds without changes. The indicator at the bottom +right, in the status bar as for translations, distinguishes unsaved text from +a save in progress, completed, or failed; its tooltip gives the time of the +last save. To save a version to the history right away, use the disk at the +top of the page, or **Ctrl + S** even while typing +on the page: the command stays off when there is nothing new to save, and the +version is created without a name — pin it in the history to give it one. +Changing page or leaving the Studio normally saves pending text immediately. A forced shutdown before a save can lose the latest edits. -When the text is ready, mark it as **verified** with the lock next to the -"Page N" title: the text becomes locked, so an already-checked transcription +When the text is ready, mark it as **verified** with the check next to the +page title: the text becomes locked, so an already-checked transcription doesn't get overwritten by accident. You can return it to draft at any time -with the same command. +with the same command. When the check is off, its tooltip says why: empty +page, loading, or being read. If a save fails, the disk turns red and retries +it. While the viewer is still opening the chosen page, a veil covers the text and history with a spinner in the middle: writing or restoring stay @@ -63,8 +98,11 @@ between them. If the two declare the same number of pages, the switch is smooth: same numbering, the text follows. If the number doesn't match, switching to the secondary copy detaches the viewer from the text — browse it freely to find what you need, while the text pages through its own -arrows, next to the "Page N" title. Switching back to the main copy -restores the link on its own. +arrows, which appear next to the page title only when text and viewer are +detached. Switching back to the main copy restores the link on its own. In a +narrow window the viewer's secondary commands — copy switch, detach, local +files only, open the page on the site — move into the three-dots menu; +paging, go-to-page and zoom always stay visible. A third command detaches the link **regardless** of page counts, even while staying on the main copy: handy for glancing at a different page @@ -115,7 +153,7 @@ the reading command stays under the reopen command. While a page is being read its sheet is veiled and stays read-only, even after a cancellation request, until the job actually stops. Text you are still editing on another page is not replaced when reading finishes. -A pill in the top row of the text column names the page being read, and stays visible even if you page +A gold label in the top row of the text column names the page being read, and stays visible even if you page ahead in the meantime. The reading starts in the queue, like a download: you'll find it in the @@ -133,11 +171,12 @@ was sent (expandable row by row), and the outcome with duration, tokens, estimated cost and the number of the revision created. On top there are search, filters by row type and by level, and grouping by page. -## History, summary and metadata +## History and summary -Every save stays in that page's history, in the panel on the right: it -shows who wrote that version — manual correction, automatic recognition, or -import — and when. The command on each history entry brings that version's +The right column has three tabs — **History**, **OCR**, **Summary** — and +opens on History. Every save stays in that page's history: for each version +it shows who wrote it — manual correction, automatic recognition, or import +—, when, and whether it is the current or the verified version. The command on each history entry brings that version's text back as a new save, without overwriting earlier versions, even after a restore. Changing page changes the history shown, too. You can consolidate any version and give it a name, which you can edit later. @@ -145,16 +184,14 @@ Consolidated versions appear above ordinary saves without duplicating their text. Remove a name to move a version back to ordinary history. You can delete older versions one by one, or use **Clear history** to delete older ordinary saves. Consolidated, current, and verified versions survive the -cleanup. The current and verified versions cannot be deleted. Deleting a +cleanup. The current and verified versions cannot be deleted: the command stays off +and its tooltip says why. Deleting a version removes the ability to restore its text. -The **Summary** tab follows the Translation summary layout: it shows pages -with text, word count, verified pages, and progress. It also shows completed -OCR readings, tokens, and estimated cost. - -The **Metadata** tab, next to History, shows the raw data saved for the -current page: position, label, status, and revision count. It's there to -show what's recorded today; how it's presented will change. +The **Summary** starts from the open page — number, the library's page label +if any, status and saved versions — then the document: pages with text, word +count, verified pages and progress, completed OCR readings, tokens and +estimated cost. ## Current limits diff --git a/docs/en/intro/getting-started.md b/docs/en/intro/getting-started.md index 0a09fe94..e65a5677 100644 --- a/docs/en/intro/getting-started.md +++ b/docs/en/intro/getting-started.md @@ -21,14 +21,14 @@ RPM packages for Linux. Check the assets attached to your chosen release. Remote translation services require their own credentials. Local processing requires a running Ollama server and a downloaded model. Configure your -connection under **Settings → Provider**. +connection under **Settings → Models**. ## Your first translation project 1. Create a workspace: a group of projects and shared resources. 2. Create a project in that workspace and import a document. 3. Check the extracted text and segment boundaries in the import preview. -4. Set the languages, pipeline mode, providers and models for the active stages. +4. Set the languages of the work, pipeline mode, providers and models for the active stages. 5. Run a test on a representative segment and compare the result with the source. 6. Process the remaining segments, review the translations and export the document. diff --git a/docs/en/reference/backup-and-restore.md b/docs/en/reference/backup-and-restore.md index 19541e3d..5b0c710e 100644 --- a/docs/en/reference/backup-and-restore.md +++ b/docs/en/reference/backup-and-restore.md @@ -5,7 +5,7 @@ title: Backup and restore # Backup and restore A backup saves application data to a `.glossa-backup` file. It is available -under **Settings → Backup** and includes all workspaces. Restoring replaces +under **Settings → Data → Backup & Restore** and includes all workspaces. Restoring replaces the existing application data with the backup’s contents; it does not merge datasets or import an individual workspace. diff --git a/docs/en/reference/import-export.md b/docs/en/reference/import-export.md index e96061f6..260dc906 100644 --- a/docs/en/reference/import-export.md +++ b/docs/en/reference/import-export.md @@ -23,7 +23,7 @@ add support for binary or structured formats such as ODT or RTF. Non-UTF-8 text is rejected with an encoding error. The system dialog can select files from any accessible directory, including -external drives. The preview lets you check extraction and segmentation +external drives. The preview lets you check extraction and segmentation, and set languages and model, before confirming. A PDF containing only images does not provide text through this extraction path; import does not perform OCR. diff --git a/docs/en/reference/pipeline-config.md b/docs/en/reference/pipeline-config.md index 145a48ee..add4eb72 100644 --- a/docs/en/reference/pipeline-config.md +++ b/docs/en/reference/pipeline-config.md @@ -7,30 +7,117 @@ title: Pipeline configuration Configuration belongs to a project’s pipeline. Service credentials and connections belong to application settings. -## Sections +## The window -| Section | Parameters | +The configuration opens from the gear in the Translation Studio's top row, or +with Ctrl + comma. The title is the pipeline name: rename it from the Studio's +top row, not here. Explanations are not written under the fields: they appear +when you hover a section title or a row name. + +The bar places **work / pipeline** together, truncating long names and showing +the full name on hover. The pipeline name, with its small arrow, opens the menu to +select, create, rename or delete; the gear opens options. **Simple, Editorial +or DeepL** mode is always visible with an icon. The languages of the work, shared by all its +pipelines, appear in short form in the bar (see +[Languages of the work](../guides/document-pipeline)). Pipeline operations stay blocked while processing. + +| Tab | Parameters | | --- | --- | -| Settings | Mode, languages, persona, examples and general options | -| Translation | Providers, models, prompts and generation-stage options | -| Quality Control | Evaluator and consistency check | -| Term registry | Assigned glossary | -| Prompt Preview | Request structure for active stages | +| General | Mode, DeepL languages, Translation context | +| Stages | Service, model, prompt and options for each stage; context memory | +| Quality control | Refinement loop, assessment model, assessment and consistency prompts | +| Memory | Phrase memory and translation examples (off in DeepL mode) | +| Glossary | Assigned dictionary and its terms | +| Prompt preview | Every request piece by piece: stages, audit, coherence, DeepL request | + +While the pipeline runs the window stays open and readable, but a veil locks +its controls. At the bottom, the red **Reset all translations** icon deletes +translations and their audits after a confirmation; it is off, with the reason, +during a run or when there is nothing to reset. + +Once translations exist, mode, stage prompts and context memory cannot be +changed; a stage's model is closed by a lock. Opening it lets you change the +model, but chunks already translated keep the previous one. Standard, Editorial and DeepL Hybrid modes are described in the [translation workflow](../guides/document-pipeline). -## Languages and persona - -Set source and target languages. A persona is free text that replaces the -default opening of the system message. It can specify role, subject area, -languages and register. When enabled, it should state the intended language -pair and instructions accurately. - -Prompts can be saved as reusable templates, organised by context. Prompt +## Translation context and DeepL languages + +Pipelines have no language pair: languages belong to the work. For LLMs the +translation is guided by the **Translation context**: languages, +historical varieties, audience, register and goal. It is required and never +empty: a new pipeline starts from a default text (English into Italian) that +you can rewrite or replace with a saved template; the circular arrow restores +it and confirming an empty text is disabled. Translation, refinement, audit and +coherence receive it alongside their own instructions; formatting remains +limited to syntax. No language pair is added automatically to prompts: the +languages live in the context. + +Next to each prompt title is where its text comes from: +**Default**, **Custom** or **Template “name”** when it matches a saved +template of the same category. Recognition compares the text: a template +edited after it was applied is no longer recognised, because the pipeline +keeps the applied copy. + +The pencil opens a draft; the checkmark confirms and X discards it. Model-based +refinement also changes only the draft. Prompts use a subdued surface with a +green accent, icon commands and expandable previews. The context is saved with +the pipeline and copied when duplicated. + +The **DeepL · languages** pair remains visible in General, disabled in LLM +modes and enabled in DeepL. Unused stages remain visible as disabled tabs. +Switching modes preserves all stage configurations. DeepL retains LLM +refinement, audit and coherence. Phrase memory uses the languages of the work, not the DeepL ones. + +In **Prompt preview**, choose a stage, Audit or Coherence: you see the request +piece by piece in sending order, split into the **system message** +(instructions, rules and resources, identical for every chunk) and the **user +message** (the chunk and the final request). The pieces come from the backend, +the same ones that compose the real request: preview and sending match. Each +piece says whether it is **fixed in the program**, **your text**, **data** or +**automatic**, and where it is changed; its title explains its purpose. Pieces +absent from this pipeline stay visible, disabled, with the reason (for example +“only for documents imported as Markdown”). Placeholders in double braces show +where chunk data goes: it is built from the current settings, not a historical +request. **Open chunk** mode fills the same pieces with the real data of the +chunk open in the Studio, including checked memory phrases. + +All the text the program adds around your prompts — roles, rules, headings, +result requests, response format, user message — is a **system prompt** and is +edited here, in Structure mode. Each system prompt is locked: unlocking makes +it like the other prompts (draft with confirm and discard, refine wand, +templates of the **System** category, circular arrow back to the default). The +lock closes again when you close the window. The audit and coherence response +format asks for an extra confirmation, because the app reads the answer by that +format. In the user message and frames, placeholders in double braces (for +example `{{TEXT}}`) are required: without them confirming stays disabled. +Pieces whose content lives in another tab (Translation context, phase prompts, +glossary table, examples) are not edited here: they only have the arrow that +opens that tab, and their heading stays the default. Glossary rules are a +system prompt instead and are edited here. Prompts are always read and edited +in a monospaced font. Role, rules and frames of Translation and Refine are shared: +editing them once applies to both. Changed texts are saved with the pipeline +and carried to new pipelines created by copy. + +Optional pieces can be **switched off per phase** with the switch on their +card: role, structural rules, glossary rules, Markdown rules, examples, +neighbouring chunks, glossary table, review method, result request. A switched +off piece stays in the list as a grey “switched off in this phase” line with +the switch to turn it back on; the choice is saved with the pipeline and also +applies when running. The Translation context, phase prompt, user message and +response format are always on. Switching neighbouring chunks off also drops the +chunk identifier, which only serves to find the chunk among them. DeepL shows the request body. + +Prompts can be saved as reusable templates, organised by context: while +editing, the bookmark saves the prompt under a name and the book opens the +list of saved templates, with search. Templates are deleted from the language +resources. A template whose scope is no longer recognised is left out of the +list and named in a notice, without hiding the others. Prompt refinement sends the current text to a configured model and places a revised version in the field. It requires a connection and any credentials needed -by the selected provider. +by the selected provider: without a key the command is off and its tooltip +says which one is missing. ## Model parameters @@ -58,18 +145,26 @@ translation mode and remote glossary. Its API key is separate from the keys used by revision and assessment LLMs. DeepL quota or glossary errors must be resolved with that service before the sequence can complete. +Choose the DeepL pair in General, using the service's language lists. +Automatic source detection is available; a glossary requires an explicit source. +Changing either language unlinks the remote glossary, and changing the target +resets formality to its default. The Translation context is not sent to DeepL; its +Context field is separate. + +The target must be selected explicitly; requests without a target are blocked before contacting DeepL. Both the configuration preview and the chunk preview display the API +body, built by the backend as for execution; the configuration preview uses a +placeholder for chunk text. Logs retain the actual request without credentials. + ## Examples and context -A pipeline can hold up to five translation examples, added from a locked -segment’s Audit tab and edited in settings. [Phrase memory](../guides/phrase-memory) +A pipeline can hold up to five translation examples, added from a verified +segment’s Audit tab and edited in the Memory tab. [Phrase memory](../guides/phrase-memory) instead supplies references selected for an individual segment. [Prompt caching](../guides/context-and-caching) has provider-specific rules. ## Cost estimates -The configuration estimate covers the entire document, including consistency -review when configured. In the document sidebar, the estimate follows the -selected action. Details distinguish stages and models. +In the Studio's tools column, the estimate follows the selected action. Details distinguish stages and models. Estimates use an approximate word-to-token conversion and the model prices recorded in Glossa. Usage shown after execution uses token counts returned diff --git a/docs/en/reference/provider-support.md b/docs/en/reference/provider-support.md index 880d88fa..12b6bac7 100644 --- a/docs/en/reference/provider-support.md +++ b/docs/en/reference/provider-support.md @@ -20,7 +20,7 @@ describes implemented roles without ranking model quality by brand. ## Credentials -Open **Settings → Provider**. Keys are stored in the operating system’s +Open **Settings → Models**. Keys are stored in the operating system’s credential store when available; otherwise Glossa uses an encrypted local store. Keys are not included in application backups. diff --git a/docs/en/reference/troubleshooting.md b/docs/en/reference/troubleshooting.md index 892b0222..45596060 100644 --- a/docs/en/reference/troubleshooting.md +++ b/docs/en/reference/troubleshooting.md @@ -10,7 +10,7 @@ service involved. Keep other parameters unchanged while investigating one cause. ## Connection or credentials If a stage does not start, check its key, model and, for Custom, selected -profile under **Settings → Provider**. Authentication errors, unavailable +profile under **Settings → Models**. Authentication errors, unavailable models and exhausted quotas require different actions. Custom remote endpoints must use HTTPS; HTTP is accepted only for supported local addresses. diff --git a/docs/guides/annotations.md b/docs/guides/annotations.md index 5355f341..cf6e3986 100644 --- a/docs/guides/annotations.md +++ b/docs/guides/annotations.md @@ -17,18 +17,18 @@ così una nota può essere modificata o rimossa senza riscrivere la traduzione. | Problema | Errore che richiede un intervento | | Approvato | Nota che registra l’esito della revisione | -Il tipo Approvato non sostituisce il comando **Blocca traduzione**. Le -annotazioni descrivono il lavoro di revisione; il blocco controlla la +Il tipo Approvato non sostituisce la spunta **Segna come verificata**. Le +annotazioni descrivono il lavoro di revisione; la verifica controlla la possibilità di rielaborare il frammento. ## Creazione Seleziona un passaggio nella traduzione e usa **Aggiungi annotazione** dal menu contestuale. Il testo selezionato diventa il riferimento della nota. -Puoi anche aggiungere una nota senza selezione dalla scheda **Note** del -frammento, oppure convertire una segnalazione dell’audit in annotazione. +Puoi anche aggiungere una nota senza selezione con **+** nella sottolinguetta +**Note** di **Revisione**, oppure convertire una segnalazione dell’audit in annotazione. -Le note del frammento si trovano nella barra laterale del progetto. Non sono +Le note del frammento si trovano nella colonna Strumenti dello Studio. Non sono le note bibliografiche dell’opera, che appartengono alla scheda della Biblioteca. ## Visualizzazione ed esportazione diff --git a/docs/guides/audit-review.md b/docs/guides/audit-review.md index ee363583..db66b63b 100644 --- a/docs/guides/audit-review.md +++ b/docs/guides/audit-review.md @@ -32,11 +32,11 @@ identici né valutazioni corrette; il rispetto dello schema riguarda il formato. 3. Correggi il testo manualmente o riesegui la fase pertinente. 4. Usa **Rivaluta** per aggiornare il giudizio senza ritradurre. 5. Registra le decisioni e i dubbi nelle **Note**. -6. Blocca la traduzione quando la revisione è conclusa. +6. Segna la traduzione come verificata quando la revisione è conclusa. Un problema dell’audit può essere convertito in annotazione. La ricerca del passaggio usa il testo fornito dal modello e può non trovare la posizione -esatta. Il blocco della traduzione è una scelta del revisore, distinta +esatta. La verifica della traduzione è una scelta del revisore, distinta dall’esito automatico e dal tipo di annotazione. ## Coerenza del documento @@ -44,7 +44,7 @@ dall’esito automatico e dal tipo di annotazione. Dopo aver completato i frammenti, avvia il controllo di coerenza. Esamina le traduzioni con il contesto dei frammenti vicini, senza confrontarle con il sorgente. Usa il prompt dedicato in **Controllo qualità** e presenta -i risultati nella scheda **Coerenza** del pannello Insight. +i risultati in **Documento** → **Coerenza**, nella colonna Strumenti. Questo controllo può evidenziare variazioni terminologiche o stilistiche tra passaggi. Non sostituisce l’audit di fedeltà del singolo frammento. diff --git a/docs/guides/context-and-caching.md b/docs/guides/context-and-caching.md index 928047f0..abbd4ce4 100644 --- a/docs/guides/context-and-caching.md +++ b/docs/guides/context-and-caching.md @@ -24,7 +24,7 @@ costruisce il contesto dalle traduzioni, anziché dal testo originale. Per traduzione e revisione, il messaggio di sistema mantiene questo ordine: -1. Istruzioni statiche: persona, regole strutturali, glossario ed esempi. +1. Istruzioni statiche: ruolo e Contesto di traduzione, regole strutturali, glossario ed esempi. 2. Contesto documentale condiviso. 3. Istruzioni specifiche della fase, con gli eventuali riferimenti di memoria selezionati. @@ -60,7 +60,8 @@ riutilizzabile. ## Configurazione e verifica -La cache Anthropic è disattivata per impostazione predefinita. Attivala quando +La cache Anthropic è disattivata per impostazione predefinita; l’interruttore sta +sotto il modello di ogni fase Anthropic, nella linguetta Fasi. Attivala quando prevedi di riutilizzare un prefisso e valuta l’opzione di durata estesa in base agli intervalli tra richieste. La scrittura in cache può avere un costo, quindi un prefisso mai riutilizzato non produce necessariamente un risparmio. diff --git a/docs/guides/document-pipeline.md b/docs/guides/document-pipeline.md index 928395eb..e9858f06 100644 --- a/docs/guides/document-pipeline.md +++ b/docs/guides/document-pipeline.md @@ -4,6 +4,10 @@ title: Traduzione di un documento # Traduzione di un documento +Durante l’esecuzione non puoi cambiare pipeline o crearne una nuova. I gruppi +Memoria, Revisione e Documento ricordano la sottoscheda scelta quando cambi +scheda; se non è disponibile, viene mostrata una vista disponibile. + La traduzione opera su frammenti di testo, chiamati *chunk* in alcune parti dell’interfaccia. Ogni frammento conserva sorgente, risultati delle fasi, traduzione modificabile, valutazione e annotazioni. Anche un documento composto @@ -20,16 +24,71 @@ Verifica i confini prima di confermare: essi determinano le unità di traduzione e revisione. I dettagli su formati, limiti e note importate sono nel [riferimento per importazione ed esportazione](../reference/import-export). +## Lingue dell’opera + +Le lingue appartengono all’opera tradotta, non alle sue pipeline: valgono per +tutte le pipeline dell’opera, che non hanno più una coppia di lingue. Il +prompt di traduzione è guidato dal **Contesto di traduzione**; solo DeepL +conserva una coppia propria nelle opzioni della sua fase, perché il servizio +richiede codici. + +Partenza e Arrivo hanno ciascuno tre campi facoltativi: + +- **Lingua**: codice ISO 639-3 del registro ufficiale SIL, quello dei cataloghi + bibliotecari, comprese lingue storiche (latino, greco antico, francese antico + e medio, provenzale/occitano antico, spagnolo antico, inglese antico e + medio, alto tedesco antico e medio, anglo-normanno e altre). Si cerca per + nome italiano, nome inglese o codice; a ricerca vuota l’elenco mostra + «Già usate nel workspace» e «Lingue storiche», scrivendo si cercano tutte le + circa 7.900 lingue. I nomi sono in italiano dove esiste un nome standard, + altrimenti l’inglese ufficiale. +- **Varietà**: una varietà di Glottolog, offerta solo tra quelle della lingua + scelta (latino: tardo latino, latino medievale, latino volgare; italiano: + italiano antico, fiorentino, laziale, cicolano-reatino-aquilano). Resta + disabilitata finché non scegli una lingua. +- **Nota**: testo libero per ciò che i codici non dicono (epoca, area, mano), + per esempio «volgare padano, sec. XV». + +Ogni campo può restare vuoto («non indicata») e ha una X per svuotarlo. Una +nuova opera parte con partenza non indicata e arrivo italiano. + +Limiti: gli standard non hanno un codice per latino medievale, italiano antico +o catalano antico come lingue; Glottolog non elenca una varietà «latino +classico»; le lingue regionali italo-romanze (veneto, lombardo, ligure, +napoletano, siciliano…) sono lingue separate con sole varietà moderne; non +esistono varietà d’area medievali, per cui si usa la nota. Per un trattato di +scherma volgare del Quattrocento: italiano + italiano antico + nota, oppure la +lingua regionale + nota. L’elenco delle lingue è incluso nell’app e funziona +senza rete; **Impostazioni → Lingue** mostra l’elenco in uso e lo riscarica +dalle fonti ufficiali, conservando come ritirate le lingue che non ci sono più. + +Dove si impostano: nella finestra di importazione e nella riga in cima allo +Studio. Lì, a destra prima dei comandi, la coppia compare con le varietà +(«Italiano (Old Italian) → Inglese»); il suggerimento aggiunge le note. L’icona delle +lingue accanto apre la finestra **Lingue dell’opera** con gli stessi campi, Annulla e +Conferma: senza conferma non si salva nulla. Mentre la pipeline lavora il +comando è visibile ma bloccato, con il motivo nel suggerimento. Se l’opera non +ha lingua di partenza e viene da un libro della Biblioteca la cui lingua +corrisponde a una lingua nota, la partenza è proposta già compilata (da +confermare); nella finestra di importazione si precompila allo stesso modo. + +Dopo ogni conferma, se la memoria contiene frasi salvate da questa opera con +lingue diverse, l’app chiede se darle le lingue dell’opera: testo e misure di +somiglianza restano uguali e la versione precedente resta nella cronologia. Vale +anche per allineare frasi salvate prima. Vedi la +[memoria di frasi](./phrase-memory). + ## Configurazione -Apri la configurazione della pipeline e imposta lingue, modalità, provider, -modelli e istruzioni. Le modalità definiscono questa sequenza: +Apri la configurazione della pipeline (l’ingranaggio nella riga in cima allo +Studio) e imposta Contesto di traduzione, modalità, provider, modelli e istruzioni; le lingue stanno nell’opera (vedi sotto) e solo DeepL ha una propria coppia: le +linguette sono descritte nella [configurazione della pipeline](../reference/pipeline-config). Le modalità definiscono questa sequenza: | Modalità | Elaborazione | | --- | --- | | Standard | Traduzione e valutazione automatica | | Editoriale | Traduzione, revisione della bozza (*Refine*), formattazione (*Format*) e valutazione | -| DeepL Hybrid | Traduzione DeepL, revisione LLM facoltativa e valutazione LLM | +| DeepL Hybrid | Traduzione DeepL, revisione LLM e valutazione LLM | Provider e modelli delle fasi LLM sono indipendenti. DeepL richiede una propria chiave API e non svolge il ruolo di valutatore. La modalità della pipeline non @@ -45,25 +104,100 @@ frammenti; il numero impostato limita il gruppo da elaborare. L’elaborazione procede per frammenti e ne aggiorna lo stato. L’annullamento interrompe il lavoro corrente senza eliminare i risultati già completati. La ripresa e la rielaborazione hanno scopi diversi: la prima completa il lavoro -restante, la seconda ricalcola i frammenti non bloccati selezionati dall’azione. +restante, la seconda ricalcola i frammenti non verificati selezionati dall’azione. +Se dopo l’interruzione cambi modelli, prompt delle fasi o dell’audit, la +Contesto di traduzione o le opzioni DeepL, la ripresa avvisa che la configurazione non è +più quella con cui il lavoro era cominciato. + +Mentre un frammento si traduce, il testo della sua traduzione è coperto da un +velo oro, «Traduzione in corso…», e non si modifica: il testo non compare man +mano, arriva quando la fase finisce. La colonna delle fasi nel margine resta +usabile; l’originale è in sola lettura e la matita dice perché. ## Lettura e revisione -La vista documento affianca originale e traduzione. La barra del frammento -contiene **Riferimenti**, **Anteprima**, **Audit**, **Memoria** e **Note**. -Il pannello **Insight** raccoglie indice, ricerca, statistiche, coerenza e -glossario dell’intero documento. +Lo Studio di traduzione si apre dentro la cornice dell’applicazione: la barra +principale a sinistra resta in vista e porta a qualunque area, chiudendo la +traduzione. Mentre la pipeline lavora le sue voci sono spente, come il ritorno +al catalogo. + +In cima, la barra mostra opera / pipeline e il tipo Semplice, Editoriale o DeepL. I nomi lunghi si troncano; il suggerimento mostra il nome completo. Il nome dell’opera si rinomina con un clic. Il nome della pipeline, con la piccola freccia accanto, apre il menu per scegliere, creare, rinominare o eliminare; l’ingranaggio apre le opzioni. Accanto compaiono le lingue dell’opera in forma breve (vedi [Lingue dell’opera]). A destra stanno importa, esporta, risorse linguistiche ed eliminazione. + +Al centro i due fogli affiancano originale e traduzione. Sopra di loro, a +sinistra, il numero del frammento aperto; al centro una finestra di sette +pallini, uno per frammento con il suo stato: il frammento aperto resta fermo +sotto il segno centrale e gli altri scorrono ai lati. Le frecce singole passano +al frammento vicino, quelle doppie saltano di sette; anche la rotella del mouse +sopra i pallini scorre i frammenti, e un clic su un pallino lo apre. A destra +dei pallini, le spie delle fasi dicono a che punto è il frammento aperto +(traduzione, revisione, formattazione, audit): un clic apre il dettaglio della +fase. La lente accanto apre, sotto la fila, la ricerca in tutto il documento; +un risultato porta al suo frammento, Esc la chiude. + +A destra, la colonna **Strumenti** tiene in cima l’esecuzione — traduci, l’interruttore **Blocchi multipli** con il +numero di frammenti da elaborare (sempre in vista, spento quando si traduce un +frammento solo) — e i costi, poi le linguette, in quest’ordine: +**Glossario**, **Memoria**, **Anteprima**, **Revisione** e **Documento**, che +raccoglie in tre sottolinguette i riepiloghi del documento intero: **Indice**, +**Statistiche** e **Coerenza**. Memoria raccoglie in due +sottolinguette le **frasi simili in memoria**, da usare traducendo, e +**Estrai frasi**, che salva le coppie del frammento; l’estrazione si accende a +frammento tradotto. Revisione +raccoglie in tre sottolinguette **Audit**, **Note** e **Note del testo** (le +note a piè di pagina importate con l’originale, presenti solo se il frammento +ne ha), linguette a icona con nome e conteggio nel suggerimento, ognuna con il +suo elenco; si apre +sull’audit se ha segnalazioni aperte, altrimenti sulle note. L’audit si +accende a frammento tradotto, il Glossario con un glossario assegnato; il +motivo resta nel suggerimento. Chiusa a icone, la +colonna lascia in vista il solo pulsante traduci, o lo stop durante +l’esecuzione. I risultati intermedi delle fasi permettono di individuare dove è stata -introdotta una modifica. Dopo una correzione manuale, **Rivaluta** esegue il -solo controllo qualità. **Blocca traduzione** protegge un risultato approvato -dalla rielaborazione. Se cambia il testo sorgente, l’interfaccia segnala che -la traduzione richiede un aggiornamento. +introdotta una modifica. I comandi stanno in colonna nel margine destro del +foglio della traduzione, accanto alla barra di scorrimento: in alto le fasi +nell’ordine della pipeline, poi il confronto e le coppie da confrontare; quella +che stai guardando è evidenziata. Dopo una correzione manuale, **Rivaluta** esegue il +solo controllo qualità. Se correggi l’originale con la matita, accanto al titolo della traduzione +compare l’etichetta ocra **Sorgente modificata** e il pallino del frammento ha +un segno ocra: la traduzione va aggiornata. + +La spunta accanto al titolo **Traduzione candidata** segna la traduzione come +verificata: diventa verde, il testo si blocca e la rielaborazione dei soli +frammenti non verificati la salta. Verificare toglie anche il «da aggiornare», +perché vuol dire averla controllata sull’originale di adesso; lo stesso comando +la riporta in bozza. La spunta è spenta mentre il frammento è in traduzione o +quando non c’è ancora una traduzione. Ogni comando spento dice il motivo nel +suggerimento. + +Limite attuale: il «da aggiornare» non si conserva chiudendo la traduzione; +riaprendola, il segno non c’è più. + +### Storico del frammento + +**Revisione → Storico** elenca le versioni del frammento aperto, dalla più +recente, con l’autore (**Pipeline** o **Manuale**), data e ora, e i segni +**Corrente** e **Verificata**. Una versione nasce a ogni passata della pipeline +(anche la riscrittura dopo l’audit), a ogni salvataggio col dischetto o con +`Ctrl + S` se il testo del frammento è cambiato dall’ultima versione, e alla +verifica quando il testo verificato è diverso. Il salvataggio automatico non +scrive versioni, per non riempire lo storico a ogni pausa. + +Il comando di ripristino riporta il testo di una versione nel foglio e lo +scrive come versione nuova: le precedenti restano. È spento su una traduzione +verificata (prima va riportata in bozza) e mentre il frammento è in +traduzione. Limiti attuali: le versioni non si eliminano, non si possono +nominare, e lo storico non indica il modello usato, perché una pipeline ne usa +più d’uno. Lo storico è per pipeline e si perde se il documento viene diviso di +nuovo in frammenti diversi. ## Anteprima delle richieste -La configurazione mostra la struttura dei prompt. La scheda **Anteprima** del -frammento costruisce invece la richiesta della fase scelta per il testo corrente. +L’**Anteprima prompt** nelle opzioni della pipeline ha due modi. **Struttura** +mostra i pezzi di ogni richiesta con i segnaposto dove entrano i dati. +**Frammento aperto** li riempie con il frammento aperto nello Studio: testo, +frammenti vicini, frasi della memoria spuntate, traduzione precedente. Audit e +Coerenza usano la traduzione attuale del frammento; senza traduzione lo dicono. Questa operazione non chiama il modello e non produce una traduzione. ## Esportazione @@ -72,3 +206,24 @@ Controlla anche i frammenti incompleti prima di esportare: nei formati ordinari, un frammento senza traduzione può essere esportato con il testo sorgente. Il formato bilingue distingue esplicitamente originale e traduzione assente. Vedi [formati e contenuto esportato](../reference/import-export). + + +## Salvataggio, navigazione e costi + +Durante una traduzione in corso lo Studio resta aperto: per uscire, attendi la fine +o interrompi l’elaborazione. Il dischetto e Ctrl/⌘+S salvano; dopo un errore il +comando propone **Riprova**. Gli errori visibili sono messaggi tradotti; i dettagli +tecnici sono nel log. Lo stesso vale per caricamento e ripristino dello storico. + +Le schede disattivate restano raggiungibili col tabulatore, così puoi leggere il +motivo nell’etichetta. Non si attivano; le frecce passano alle schede disponibili. +I selettori circolari restano raggiungibili anche se la scelta corrente non è disponibile. + +Nella colonna Strumenti, sotto i comandi, una riga riporta la stima del +prossimo lancio e lo speso sul frammento aperto. Un clic apre il pannello dei +costi: due tabelle con una riga per fase (modello, token, costo). La stima segue +modalità e numero di blocchi selezionati ed è indicativa. Lo speso riporta anche +le **chiamate** al modello per fase: un frammento ritradotto o un ciclo di +revisione aggiungono chiamate, quindi il numero non conta le esecuzioni. Il +numero di blocchi non indica ripetizioni della traduzione. La scheda Statistiche +mostra la stessa tabella per fase sull’intero documento. diff --git a/docs/guides/glossary-and-memory.md b/docs/guides/glossary-and-memory.md index b3fbb0d0..6e5c9fff 100644 --- a/docs/guides/glossary-and-memory.md +++ b/docs/guides/glossary-and-memory.md @@ -14,10 +14,38 @@ Apri **Risorse linguistiche** nel workspace. La finestra distingue dizionari, modelli di prompt e frasi. Nella scheda dei dizionari puoi creare, rinominare, duplicare ed eliminare risorse, modificare voci e importare dati da CSV o TSV. -L’importazione mostra un’anteprima e permette di scegliere tra integrazione -e sostituzione del contenuto. Verifica l’associazione dei campi sorgente, -destinazione e note prima di confermare. La sostituzione elimina le voci -precedenti del dizionario. +L’importazione dalle risorse crea un nuovo dizionario nel workspace scelto, +con il nome del file. Il comando a icona apre la scelta del file; l’anteprima +mostra le prime voci. Per Excel puoi associare le colonne di termine, +traduzione e note prima di confermare. + +Le **Risorse linguistiche generali** offrono la stessa gestione. Il filtro per +workspace mostra tutti i dizionari collegati, compresi quelli condivisi; puoi +anche scegliere tutti o quelli senza workspace. Il filtro restringe l’elenco: +nelle risorse generali si leggono e modificano sempre gli originali. +Per creare, importare o copiare un dizionario scegli il workspace di +destinazione. La ricerca controlla i nomi dei dizionari. + +Nel dizionario aperto la condivisione indica l’**Originale condiviso** e lo +scudo le **Correzioni locali**. Il suggerimento dell’icona spiega dove valgono +le modifiche; nei workspace ospiti il più accanto ricorda che le nuove voci +entrano nell’originale condiviso. Un termine sorgente esistente non si rinomina +attraverso una correzione locale. + +Le voci affiancano termine e traduzione, senza campi aperti per tutta la lista. +La matita apre la modifica di una coppia; il quaderno mostra le sue note. +**+** inserisce una nuova voce in cima e porta il cursore al termine, rendendola +subito visibile. La spunta termina la modifica della voce: i cambiamenti +restano da salvare con il dischetto del dizionario. + +Più e dischetto restano nell’intestazione delle voci mentre scorri. +Il dischetto salva le voci modificate. La copia propone i dizionari con il +workspace di provenienza e segna quello scelto; l’esportazione offre CSV ed +Excel affiancati. **Salva e chiudi** rispetta lo stesso +ambito; se il salvataggio fallisce la finestra resta aperta. **Chiudi senza +salvare** scarta davvero le modifiche. Rinomina ed elimina riguardano il +dizionario originale condiviso; eliminare richiede conferma. Le esportazioni +CSV ed Excel contengono le voci originali salvate. ## Condivisione e correzioni locali @@ -27,8 +55,11 @@ correzioni o esclusioni applicate nel workspace modificano la vista locale delle voci senza alterare l’originale condiviso. Assegna il dizionario al progetto con il comando dedicato. La scheda -**Glossario** del pannello Insight mostra l’intero glossario assegnato; -la configurazione della pipeline ne espone il registro terminologico. +**Glossario** della colonna Strumenti mostra l’intero glossario assegnato, con +il numero dei termini nel titolo; il comando di evidenziazione colora i termini +nei fogli e, finché è acceso, mostra la legenda dei colori; +nella linguetta Glossario della configurazione della pipeline si assegna il +dizionario e se ne modificano i termini, salvandoli con il dischetto. ## Applicazione alla traduzione @@ -52,6 +83,27 @@ L’evidenziazione segnala corrispondenze testuali: non interpreta il contesto e non sostituisce la verifica linguistica. Una resa assente può richiedere una correzione o una variante motivata nelle note del glossario. +## Modelli di prompt + +La scheda **Modelli Prompt** delle Risorse linguistiche cerca nel nome e nel +testo e filtra per Fasi, Controllo qualità, Contesto di traduzione, Sistema, Memoria oppure OCR. +I modelli sono comuni all’applicazione, senza appartenenza a un workspace. +Usa **+** per crearne uno o la matita sulla sua riga per modificarlo. + +Il modulo conserva nome, ambito, flusso, testo e, facoltativamente, servizio e +modello predefinito. I suggerimenti delle etichette spiegano l’uso dei campi. +Scegli servizio e modello per rifinire il testo; il comando spento indica ciò +che manca. Il dischetto salva, la croce annulla. Un nome già usato nello stesso +ambito e flusso richiede di modificare il modello esistente o scegliere un +altro nome. Il cestino elimina soltanto dopo conferma. + +Ogni modello ha un titolo riconoscibile e un’anteprima breve sulla stessa carta tenue della voce. L’occhio apre +il testo completo, il comando di riduzione torna all’anteprima. Le icone +accanto al titolo spiegano ambito, flusso e modello al passaggio del mouse o +premendole. La matita apre il modulo nella stessa voce; durante la modifica +ricerca e filtri restano bloccati per conservare la bozza. Il modulo affianca +ambito e flusso, servizio e modello, senza separatori fra i campi. + ## Glossario, memoria ed esempi | Risorsa | Ruolo | @@ -62,3 +114,5 @@ una correzione o una variante motivata nelle note del glossario. Per estrazione, selezione e salvataggio delle coppie bilingui, consulta [Memoria di frasi ed esempi](./phrase-memory). + +Le liste di dizionari e prompt distinguono ogni voce con un unico sfondo tenue, condiviso da testo e dettagli aperti. La matita del titolo del dizionario apre il nome al suo posto, mantenendo i comandi sulla stessa riga; Invio salva, Esc annulla. diff --git a/docs/guides/keyboard-shortcuts.md b/docs/guides/keyboard-shortcuts.md index f361c677..438ebd28 100644 --- a/docs/guides/keyboard-shortcuts.md +++ b/docs/guides/keyboard-shortcuts.md @@ -11,7 +11,8 @@ un campo di testo, una selezione o un editor modificabile. | Scorciatoia | Azione | Condizioni | | --- | --- | --- | | `Ctrl + Invio` | Avvia l’azione di traduzione selezionata | In modalità frammento richiede un frammento selezionato; funziona anche nei campi di testo | -| `Ctrl + S` | Salva il progetto e le risorse linguistiche modificate | Richiede un progetto esistente; il progetto non viene salvato mentre è in elaborazione | +| `Ctrl + S` | Salva la traduzione aperta (con una versione nello storico dei frammenti cambiati) e le risorse linguistiche modificate | A traduzione aperta funziona anche mentre scrivi nei fogli; la traduzione non viene salvata mentre è in elaborazione | +| `Ctrl + S` nello Studio di trascrizione | Salva subito una versione della pagina nello storico | Funziona anche mentre scrivi nel foglio; non fa niente se non c’è niente di nuovo da salvare | | `Ctrl + E` | Apre l’esportazione | Richiede almeno un frammento | | `Ctrl + ,` | Apre la configurazione della pipeline | Fuori dai campi modificabili | | `Ctrl + H` | Apre questa sezione della guida in-app | Fuori dai campi modificabili | diff --git a/docs/guides/llm-and-pipelines.md b/docs/guides/llm-and-pipelines.md index b74f23b3..0b995865 100644 --- a/docs/guides/llm-and-pipelines.md +++ b/docs/guides/llm-and-pipelines.md @@ -29,8 +29,9 @@ nella richiesta. | Judge | Sorgente, traduzione e criteri di valutazione | Valutazione e problemi strutturati | | Coherence | Traduzioni e contesto dei frammenti vicini | Segnalazioni di incoerenza tra frammenti | -La fase Format usa un prompt separato: non riceve persona, glossario o contesto -sorgente della traduzione. Le istruzioni ne limitano il compito alle correzioni +La fase Format usa un prompt separato: non riceve il Contesto di traduzione, il glossario, le frasi +della memoria né il contesto sorgente della traduzione. Il glossario compare una +sola volta, nelle regole all’inizio delle istruzioni di traduzione e Refine. Le istruzioni ne limitano il compito alle correzioni di formattazione, ma il risultato deve comunque essere controllato. In DeepL Hybrid, la prima fase usa l’API DeepL e i relativi parametri di lingua, diff --git a/docs/guides/phrase-memory.md b/docs/guides/phrase-memory.md index d43e01f9..f8236a35 100644 --- a/docs/guides/phrase-memory.md +++ b/docs/guides/phrase-memory.md @@ -4,8 +4,11 @@ title: Memoria di frasi ed esempi # Memoria di frasi ed esempi +Durante la modifica di testi o metadati, la ricerca resta disabilitata +con il motivo. I titoli lunghi della provenienza vanno a capo. + La memoria di frasi conserva coppie di testo sorgente e traduzione approvata -per riutilizzarle nei progetti del workspace. La ricerca di corrispondenze, +per riutilizzarle nelle traduzioni. La ricerca di corrispondenze, la loro selezione per il prompt e il salvataggio di nuove frasi sono operazioni separate. @@ -15,8 +18,17 @@ Quando la funzione è attiva, Glossa cerca corrispondenze per i frammenti del documento. La ricerca usa le risorse accessibili al workspace e non modifica né traduzioni né frasi salvate. -La scheda **Riferimenti** mostra i risultati e permette di regolare la soglia -di somiglianza. Solo le coppie selezionate vengono incluse nella successiva +La sottolinguetta **Riferimenti** della scheda **Memoria** mostra i risultati. +In alto, una sola riga: la soglia di somiglianza (cursore, oppure − e + per un +centesimo alla volta), il globo e l’aggiornamento. Ogni risultato mostra +l’originale in carattere da libro e la traduzione sotto, con il codice della +lingua a margine; in testa, accanto alla somiglianza, la coppia di lingue della +frase con le varietà («Italiano (Old Italian) → Inglese»); sotto, una riga con +opera e frammento; l’icona «i» aggiunge +workspace, libro e modello. Ogni risultato dice da dove viene: «questo documento» +(evidenziato), «altra opera del workspace» o «altro workspace». L’ordine è +prima questo documento, poi il workspace, poi il resto; in ogni gruppo prima la +somiglianza più alta. Solo le coppie selezionate vengono incluse nella successiva richiesta per quel frammento. Se esistono risultati ma nessuno è selezionato, l’avvio segnala che la traduzione procederà senza quei riferimenti. @@ -27,24 +39,94 @@ semantica o l’adeguatezza della resa al contesto corrente. ## Creazione e revisione delle frasi -1. Rivedi la traduzione e blocca il frammento. -2. Apri **Memoria**: le coppie già salvate vengono caricate e selezionate. +1. Rivedi la traduzione e segnala come verificata. +2. Apri **Memoria** → **Memoria**: le coppie già salvate vengono caricate in sola lettura. 3. Usa **Estrai frasi** per ottenere nuove proposte, oppure aggiungi coppie manualmente. 4. Correggi i testi e seleziona le coppie da conservare. -5. Salva per applicare la selezione. +5. Usa il dischetto **Aggiungi alla memoria le coppie spuntate**. -L’estrazione non salva automaticamente. Togliere la selezione a una coppia -già salvata e confermare ne provoca la rimozione dalla raccolta. Le modifiche +L’estrazione non salva automaticamente. Il dischetto aggiunge soltanto le coppie +nuove selezionate e conserva quelle già salvate. Per rimuovere una coppia +salvata usa il suo cestino e conferma. Le modifiche non confermate restano nella bozza del frammento quando si passa a un altro frammento durante la revisione; non equivalgono a un salvataggio permanente. -## Ambito +## Ambito e compatibilità + +Le nuove frasi salvate prendono lingua, varietà e nota dell’opera da cui +provengono; l’estrazione automatica delle coppie comunica al modello le lingue +reali dell’opera, con varietà e nota. Dopo ogni conferma nella finestra +[Lingue dell’opera](./document-pipeline), se la memoria ha frasi +di quest’opera con lingue diverse, l’app chiede se darle le lingue dell’opera: +testo e misure di somiglianza restano uguali e la versione precedente resta +nella cronologia. Le frasi estratte seguono il workspace corrente della traduzione di origine. +La provenienza conserva separatamente il workspace al momento dell’estrazione. +Consultare una frase da un altro workspace non crea collegamenti né copie. + +L’icona del globo accanto al comando di aggiornamento dei Riferimenti estende la +ricerca agli altri workspace e alle frasi senza workspace; la scelta è ricordata +per ogni workspace e non sta più nelle impostazioni del workspace. Filtra solo la +lingua di arrivo: frasi tradotte in un’altra lingua di arrivo non vengono +suggerite; la lingua di partenza non filtra. Se l’opera non ha lingua di arrivo, +nessun filtro linguistico. La ricerca confronta sempre embedding dello stesso modello, dimensione e profilo di input. +Testi privi della misura richiesta restano nel catalogo, senza entrare nei risultati. +Non si confrontano embedding di modelli diversi. Le coppie già salvate nel +frammento non diventano automaticamente corrispondenze a distanza zero. + +## Testi, provenienza e più embedding + +Ogni coppia è collegata a un’unità testuale con versioni dell’originale e della +traduzione. Aggiungere un modello conserva gli embedding degli altri modelli e +non duplica la coppia. Testi uguali provenienti da frammenti diversi restano distinti. + +Creando una traduzione puoi indicare **Libro e versione di origine** scegliendo +una versione della Biblioteca (con la ricerca per titolo o copia). È una scelta esplicita: senza, il libro resta +**Non specificato**. Non viene dedotto dal nome del file. Le frasi estratte +registrano questo libro, la versione, la traduzione, il frammento e il workspace +di origine; non ricevono numeri di pagina se questi non sono noti. + +## Gestire la raccolta + +Apri **Risorse linguistiche → Memorie**. Le risorse generali partono da tutte +le frasi; quelle del workspace dalla sua raccolta. Puoi filtrare per workspace, +**Tutti**, **Senza workspace** ed etichetta, oppure cercare nei testi e nei tag. -Le frasi estratte mantengono il collegamento alla traduzione di origine. -Spostando quella traduzione in un altro workspace, le frasi la seguono. -Le risorse importate e collegate possono essere condivise secondo i collegamenti -del workspace. La scheda **Frasi** delle Risorse linguistiche permette di -consultare la raccolta. +Ogni voce affianca originale e traduzione; lingue con varietà, titolo di provenienza e tag +permettono di orientarsi senza aprire i dettagli. Il comando **Provenienza, tag +e misure** apre le informazioni complete della sola voce scelta: libro, +workspace, traduzione, frammento, data, etichette e modelli di misura. + +Nei dettagli scegli un modello e usa il più per aggiungere il suo embedding +o la freccia circolare per ricalcolarlo. Il calcolo richiede la chiave OpenAI e +comporta costi, segnalati nella conferma. Gli altri modelli restano disponibili. +Ricerca e filtri sono bloccati durante la modifica di testi o tag, per +conservare la bozza; termina o annulla la modifica per usarli di nuovo. +Nelle impostazioni del workspace puoi calcolare il modello selezionato su tutta +la sua memoria senza cancellare gli altri modelli. La scelta del modello attivo +si applica salvando le impostazioni, indipendentemente dal calcolo. + +La matita apre la correzione dei due testi. Cambiare solo la traduzione non +ricalcola gli embedding dell’originale. Cambiare l’originale richiede il ricalcolo +di tutti gli embedding disponibili, con conferma dei costi: testi e vettori si +salvano insieme, oppure nessuna modifica viene applicata se un calcolo fallisce. +Le versioni precedenti rimangono archiviate, mentre la ricerca usa quella corrente. +Queste correzioni non modificano il documento di origine. + +Le **Etichette del testo** sono manuali e riutilizzabili; separale con un punto +e virgola. Non sono classificazioni automatiche. Il cestino elimina la coppia, +le sue revisioni ed embedding dopo conferma. Il comando del dizionario copia la +coppia nel dizionario scelto: salva prima eventuali sue modifiche pendenti. +**Esporta CSV** include testi, tag, provenienza e modelli delle sole frasi visibili. +Il backup dell’app conserva anche le revisioni e i vettori; il CSV non lo sostituisce. + +La struttura permette unità di diversa lunghezza e testi senza traduzione. +La selezione di pagine o sezioni, la classificazione automatica e lo studio +semantico del corpus sono sviluppi successivi, non funzioni disponibili in questa schermata. + +Eliminare una traduzione conserva i testi già archiviati, con libro e provenienza. +Diventano memorie senza workspace e indicano che la traduzione non è più disponibile. +Dopo una correzione dell’originale, la provenienza continua a descrivere +l’estrazione iniziale; una riga segnala la correzione, senza riattribuire la citazione. ## Esempi di stile @@ -52,10 +134,23 @@ Gli esempi di traduzione sono coppie di frammenti completi usate per orientare registro e stile della pipeline. Non vengono recuperati in base alla somiglianza del frammento corrente. -Da un frammento bloccato, il comando **Usa come esempio di stile** nella scheda -Audit aggiunge la coppia alle impostazioni della pipeline. Qui puoi modificarla +Da un frammento verificato, il comando **Usa come esempio di stile** nella scheda +Audit aggiunge la coppia alla linguetta Memoria della configurazione della +pipeline. Lì puoi modificarla o rimuoverla. Il limite è cinque esempi; poiché entrano nel contesto statico, la loro lunghezza contribuisce alla dimensione delle richieste. Usa il [glossario](./glossary-and-memory) per le rese obbligatorie e i riferimenti di memoria per formulazioni pertinenti al singolo passaggio. + +## Nello Studio di traduzione + +Nella linguetta Memoria, **Riferimenti** mostra le frasi simili già in memoria: +per ognuna, sotto la coppia, da dove viene (workspace, traduzione, frammento, +oppure «importata»). La spunta in cerchio decide quali usare nella traduzione. +**Memoria** si apre solo a traduzione verificata: le coppie già salvate si +tolgono una a una con il cestino; quelle nuove si spuntano e si aggiungono con +il dischetto, che non cancella mai le altre. La ricerca esclude le frasi tradotte +in una lingua di arrivo diversa da quella dell’opera. + +Le voci della raccolta hanno un unico sfondo tenue distinto dalla finestra; i dettagli aperti condividono lo stesso fondo. Nei dettagli gli embedding sono elencati uno per riga; il comando di calcolo mostra una rotellina fino al termine della richiesta. diff --git a/docs/guides/projects-and-workspace.md b/docs/guides/projects-and-workspace.md index 26dcf571..551cfec8 100644 --- a/docs/guides/projects-and-workspace.md +++ b/docs/guides/projects-and-workspace.md @@ -23,16 +23,101 @@ originale; creare una copia produce invece un dizionario indipendente. ## Creazione e salvataggio -La pagina del workspace e l’area **Traduzioni** consentono di creare progetti. -L’area Traduzioni raccoglie i progetti di tutti i workspace. Nome, descrizione e -icona del workspace aiutano a riconoscerne l’appartenenza nelle diverse viste. - -Il salvataggio automatico opera su progetti già creati. Le modifiche vengono -rilevate e salvate dopo un breve intervallo di inattività; durante l’elaborazione -il salvataggio automatico attende uno stato stabile. La barra di stato distingue -modifiche da salvare, salvataggio in corso, completamento ed errore. -`Ctrl + S` richiede un salvataggio manuale, con i limiti descritti nelle -[scorciatoie](./keyboard-shortcuts). +La pagina del workspace e l’area **Traduzioni** consentono di creare progetti: +il «+» accanto al titolo Traduzioni apre una finestra che chiede il nome, il +workspace e, facoltativo, il file da tradurre, negli stessi formati dell’import +dall’editor. Il file si legge appena scelto: se non è leggibile (un PDF +scansionato senza testo, un file non in UTF-8) il motivo compare sotto il campo +e non si crea nulla. Con **Crea** la traduzione si apre nell’editor con +l’anteprima dell’import, dove si scelgono lingue e frammenti; chiudendo +l’anteprima la traduzione resta vuota e il file si importa poi dall’editor. +Sotto **Libro e versione di origine** la riga mostra il titolo del libro scelto +(troncato; titolo completo e copia compaiono nel suggerimento): l’icona del +libro apre un elenco con ricerca per titolo o copia, dove ogni voce mostra il +titolo e, sotto, la copia; la X toglie il libro. La finestra non si allarga con +i titoli lunghi. + +L’anteprima dell’import è la finestra standard dell’app, ampia e quasi a tutta +altezza. In cima il nome del file e il titolo. A sinistra, in una colonna che +scorre da sola: **Lingue dell’opera** (Partenza, poi Arrivo), **Modello** +(servizio e modello) e **Suddivisione in frammenti** (suddivisione automatica, +titoli per il Markdown, accorpamento dei blocchi brevi finali, misure +predefinite, ricalcolo). A destra i conteggi di parole, paragrafi e frammenti, +la scelta fra vista a schede e vista a segmenti (piccole icone rotonde) e +l’anteprima dei frammenti a tutta altezza. In fondo l’esito del controllo del +conteggio delle parole, Annulla e Importa. +Nome, descrizione e icona del workspace aiutano a riconoscerne l’appartenenza +nelle diverse viste. + +## Il catalogo delle Traduzioni + +L’area Traduzioni raccoglie le traduzioni di tutti i workspace ed è organizzata +come il [catalogo delle Trascrizioni](./transcription#il-catalogo-delle-trascrizioni): +scaffali a destra, ricerca e filtri rapidi sopra l’elenco, tre viste (elenco, +copertine, tabella). + +- **Riga**: nome in corsivo, sotto i nomi delle lingue di partenza e di arrivo, poi + workspace, frammenti tradotti sul totale, frammenti verificati e la barretta + di completamento, verde quando tutti i frammenti sono verificati. +- **Scaffali**: Tutte, Recenti (modificate negli ultimi 30 giorni), Da iniziare + (nessun frammento tradotto), In corso, Verificate (tutti i frammenti + verificati). +- **Filtri rapidi**: workspace (resta anche uscendo e rientrando nella pagina) + e coppia di lingue; ordine per nome, ultima modifica o avanzamento; + raggruppamento per workspace o coppia di lingue. +- **Comandi di riga**: rinomina ed elimina; nelle copertine e nella tabella + stanno nel menu con i tre puntini. Un click sulla riga apre l’editor. + +Limiti attuali: i conteggi riguardano la prima pipeline del progetto, quella +che l’editor apre; una traduzione non è ancora legata all’opera o alla +trascrizione da cui parte, quindi mancano i comandi per aprirle, i filtri per +biblioteca e secolo e l’archiviazione. + +Una traduzione si salva da sola poco dopo l’ultima modifica, sempre per intero: +testo di partenza e tutti i frammenti. Durante la traduzione automatica il +salvataggio attende la fine, perché la pipeline salva da sé. La barra di stato, +in basso a destra, distingue modifiche non salvate, salvataggio in corso, +salvato ed errore; il suggerimento riporta l’ora dell’ultimo salvataggio e, +dopo un errore, il motivo. + +Per salvare subito c’è il dischetto in cima al foglio della traduzione (su +quello dell’originale quando è aperto solo l’originale), oppure `Ctrl + S`, +che funziona anche mentre scrivi nei fogli. Il dischetto scrive anche una +versione nello [storico](./document-pipeline#storico-del-frammento) di ogni +frammento cambiato. È spento quando non c’è niente da salvare né versioni nuove +da scrivere, e durante la traduzione automatica; se un salvataggio +fallisce diventa rosso e il suo clic riprova. + +Uscire dalla traduzione — ritorno al catalogo, barra principale, percorso in +alto, cambio di workspace — salva prima di chiudere. Se quel salvataggio +fallisce, la traduzione resta aperta con l’errore in vista: nessuna modifica +si perde uscendo. Resta un limite: chiudere la finestra di Glossa entro un +istante dall’ultima modifica può perderla. + +## Barra di sinistra + +La barra di sinistra raccoglie la navigazione: in alto la Dashboard con +Panoramica e Ricerca, poi le aree nell’ordine del lavoro (Biblioteca, +Trascrizioni, Traduzioni, Analisi) e infine i workspace, con il più per +crearne uno. Ogni gruppo è separato da un filetto; la spiegazione di una voce +compare passando il puntatore. + +In fondo c’è il menu generale: salva tutto, risorse linguistiche generali, +impostazioni, guida e lingua dell’interfaccia. È sempre disponibile, anche +senza un workspace attivo. + +L’icona in alto chiude la barra: restano le icone delle voci e del menu +generale, e un clic porta alla pagina senza riaprirla. Il segno di Glossa in +alto la riapre. + +## Aree e loro inchiostro + +Biblioteca, Trascrizioni e Traduzioni hanno ognuna un inchiostro proprio — +petrolio, seppia e indaco — per riconoscerle a colpo d'occhio: l'icona +dell'area nella barra di sinistra, un filetto corto sotto il titolo grande e una +carta di fondo appena diversa. I colori di stato restano gli stessi in ogni +area: il verde segna ciò che è scelto o attivo, il rosso gli errori, l'ocra le +cautele, l'oro i lavori in corso. ## Dashboard diff --git a/docs/guides/storage-and-jobs.md b/docs/guides/storage-and-jobs.md index 904b6d0b..9e1029dc 100644 --- a/docs/guides/storage-and-jobs.md +++ b/docs/guides/storage-and-jobs.md @@ -16,7 +16,7 @@ richiedere che la relativa schermata resti aperta. | Deposito | Manifesti, immagini, miniature e versioni locali | Cartella selezionabile, anche su disco esterno | | Cache di rete | Risposte di ricerca e immagini riutilizzabili | Spazio temporaneo con limite configurabile | -In **Impostazioni → Dati** puoi consultare e cambiare le posizioni. +In **Impostazioni → Dati → Cartelle** puoi consultare e cambiare le posizioni. Cambiare la cartella dati copia il database, verifica la copia e registra la nuova posizione per il riavvio; l’originale non viene cancellato automaticamente. Cambiare deposito seleziona una cartella vuota o ricollega un deposito Glossa @@ -115,7 +115,7 @@ correttamente. La cache di rete riutilizza risposte e immagini; il limite predefinito è 512 MB e la validità predefinita delle ricerche è 24 ore. Le immagini sono soggette al limite di spazio, senza la stessa scadenza temporale. In -**Impostazioni → Dati** puoi modificare questi valori e svuotare la +**Impostazioni → Dati → Cache di rete** puoi modificare questi valori e svuotare la cache. La cache non aumenta il conteggio delle pagine scaricate e non entra nel [backup](../reference/backup-and-restore). diff --git a/docs/guides/transcription.md b/docs/guides/transcription.md index 4a5b8bd6..b12909aa 100644 --- a/docs/guides/transcription.md +++ b/docs/guides/transcription.md @@ -8,9 +8,37 @@ Lo Studio di trascrizione è la modalità concentrata per scrivere e correggere il testo di un documento: visore a sinistra, testo al centro, strumenti a destra. È la stessa idea della Traduzione, applicata alla trascrizione. -## Creare un documento - -Dall'area **Trascrizioni** scegli "Nuovo documento" e dai un titolo. Puoi +## Il catalogo delle Trascrizioni + +L'area **Trascrizioni** raccoglie le trascrizioni di tutti i workspace ed è +organizzata come la Biblioteca: scaffali a destra, ricerca e filtri rapidi +sopra l'elenco, tre viste — elenco, copertine, tabella — in fondo alla riga +del titolo. + +Ogni riga mostra il nome della trascrizione in corsivo e, sotto, l'opera +trascritta: autore, anno, luogo e tipografo, titolo. Se il nome è lo stesso +titolo dell'opera, che è il nome proposto alla creazione, non si ripete. In +fondo alla riga stanno il workspace, le pagine scritte sul totale dell'opera, +le pagine verificate e la barretta di completamento. Una trascrizione senza +opera collegata non ha un totale: conta solo le pagine scritte. + +Gli scaffali sono **Tutte**, **Recenti** (modificate negli ultimi 30 giorni), +**Da iniziare**, **In corso**, **Verificate** (tutte le pagine verificate), +**Senza opera** (nate da zero) e **Archiviate**; le archiviate compaiono solo +nel loro scaffale. La ricerca guarda il nome e tutti i dati dell'opera. I +filtri rapidi restringono per workspace, biblioteca e secolo dell'opera, con +il conteggio accanto a ogni valore; l'elenco si ordina per nome, autore, anno, +ultima modifica o avanzamento, e si raggruppa per workspace o biblioteca. + +Un click sulla riga apre lo Studio. Passando sulla riga compaiono i comandi: +apri l'opera in Biblioteca, rinomina — il nome diventa un campo, Invio salva, +Esc annulla —, archivia o ripristina, elimina. Nelle copertine e nella +tabella gli stessi comandi stanno nel menu con i tre puntini. Non c'è ancora +la scelta di più trascrizioni insieme. + +## Creare una trascrizione + +Dall'area **Trascrizioni** usa il comando **+** accanto al titolo e dai un nome. Puoi anche collegarlo subito a un'opera già in Biblioteca cercandola per titolo nello stesso dialogo — facoltativo: senza, il documento resta senza visore, un solo blocco di testo. Dalla scheda di un'opera in **Biblioteca** puoi @@ -34,23 +62,30 @@ visore compare un avviso, non un errore, e resta un solo blocco di testo. In alto, un documento legato a un'opera la presenta come la scheda in Biblioteca — autore, anno, luogo e tipografo sopra, titolo sotto — con l'uscita verso il sito della -biblioteca. Il comando con i tre puntini offre solo "Rimuovi trascrizione": +biblioteca. A destra il cestino elimina la trascrizione, dopo conferma: scaricare, verificare o archiviare l'opera restano comandi della scheda in Biblioteca, non dello Studio. ## Scrivere e salvare -Il testo si salva dopo 30 secondi senza modifiche. L'indicatore in alto a -destra distingue il testo ancora da salvare dal salvataggio in corso, riuscito -o fallito, con un comando per riprovare in caso di errore. Cambiando pagina o +Il testo si salva dopo 30 secondi senza modifiche. L'indicatore in basso a +destra, nella barra di stato come per le traduzioni, distingue il testo ancora +da salvare dal salvataggio in corso, riuscito o fallito; il suggerimento dà +l'ora dell'ultimo salvataggio. Per salvare subito una versione nello storico +usa il dischetto in cima al foglio, oppure **Ctrl + S** anche mentre scrivi nel foglio: il comando +resta spento quando non c'è niente di nuovo da salvare, e la versione nasce +senza nome — per fissarla con un nome si usa la puntina nello storico. +Cambiando pagina o uscendo normalmente dallo Studio, il testo ancora da salvare viene scritto subito. Una chiusura forzata prima del salvataggio può perdere le ultime modifiche. -Quando il testo è pronto, segnalo come **verificato** con il lucchetto -accanto al titolo "Pagina N": il testo diventa bloccato, per non -sovrascrivere per sbaglio una trascrizione già controllata. Puoi tornare in -bozza in qualsiasi momento con lo stesso comando. +Quando il testo è pronto, segnalo come **verificato** con la spunta accanto +al titolo della pagina: il testo diventa bloccato, per non sovrascrivere per +sbaglio una trascrizione già controllata. Puoi tornare in bozza in qualsiasi +momento con lo stesso comando. Se la spunta è spenta, il suggerimento dice +perché: pagina vuota, in caricamento o in lettura. Se un salvataggio non +riesce, il dischetto diventa rosso e lo riprova. Mentre il visore sta ancora aprendo la pagina scelta, un velo copre il testo e lo storico con una rotellina al centro: scrivere o ripristinare restano @@ -64,8 +99,12 @@ comandi per passare dall'una all'altra. Se le due dichiarano lo stesso numero di pagine, il cambio è fluido: stessa numerazione, il testo segue. Se il numero non combacia, passando sulla copia secondaria il visore si stacca dal testo — si sfoglia liberamente cercando quel che serve, mentre il -testo si sfoglia con le proprie frecce, accanto al titolo "Pagina N". -Tornando sulla copia principale l'aggancio si ripristina da solo. +testo si sfoglia con le proprie frecce, che compaiono accanto al titolo della +pagina solo quando testo e visore sono sganciati. Tornando sulla copia +principale l'aggancio si ripristina da solo. Con la finestra stretta i comandi +secondari del visore — cambio di copia, sgancio, solo file locali, apertura +della pagina nel sito — passano nel menu con i tre puntini; sfoglio, salto a +una pagina e zoom restano sempre in vista. Un terzo comando stacca il collegamento **a prescindere** dal numero di pagine, anche restando sulla copia principale: comodo per guardare una @@ -119,7 +158,7 @@ Mentre una pagina viene letta il suo foglio si vela e resta in sola lettura, anche dopo la richiesta di annullamento, finché il lavoro si ferma davvero. Se stai modificando un'altra pagina, il testo non viene rimpiazzato quando la lettura finisce. Nella riga in alto -della colonna del testo una pastiglia dice quale pagina è in lettura, e resta +della colonna del testo una scritta color oro dice quale pagina è in lettura, e resta visibile anche se nel frattempo sfogli avanti. La lettura parte in coda, come uno scaricamento: la trovi nel pannello @@ -138,11 +177,13 @@ modello, l'immagine inviata con misura reale, peso e provenienza (libro scaricat della revisione creata. In testa ci sono ricerca, filtri per tipo di riga e per livello, e il raggruppamento per pagina. -## Storico, riepilogo e metadati +## Storico e riepilogo -Ogni salvataggio resta nello storico della pagina, nel pannello a destra: -mostra chi ha scritto quella versione — correzione manuale, riconoscimento -automatico o importazione — e quando. Il comando su ogni voce dello storico +La colonna a destra ha tre schede — **Storico**, **OCR**, **Riepilogo** — e si +apre sullo Storico. Ogni salvataggio resta nello storico della pagina: per ogni +versione mostra chi l'ha scritta — correzione manuale, riconoscimento +automatico o importazione —, quando, e se è la versione corrente o quella +verificata. Il comando su ogni voce dello storico riporta il testo di quella versione come nuovo salvataggio, senza sovrascrivere le versioni precedenti. Cambiando pagina lo storico mostrato cambia con lei. @@ -152,16 +193,14 @@ senza occupare spazio aggiuntivo. Puoi togliere il nome per riportarle nello storico ordinario. Puoi eliminare una vecchia versione singolarmente, oppure usare **Svuota storico** per cancellare i salvataggi ordinari precedenti. Le versioni consolidate, quella corrente e quella verificata restano dopo -la pulizia. La versione corrente e quella verificata non si possono eliminare. +la pulizia. La versione corrente e quella verificata non si possono eliminare: il comando +resta spento e il suggerimento dice perché. L'eliminazione di una versione toglie la possibilità di ripristinarne il testo. -La scheda **Riepilogo** segue il layout del riepilogo della traduzione: -mostra pagine con testo, parole, pagine verificate e avanzamento; sotto -riporta letture OCR completate, token e costo stimato. - -La scheda **Metadati**, accanto allo Storico, mostra i dati grezzi salvati -per la pagina corrente: posizione, etichetta, stato e numero di revisioni. -Serve a vedere cosa viene registrato oggi; la sua presentazione cambierà. +Il **Riepilogo** parte dalla pagina aperta — numero, etichetta della +biblioteca se c'è, stato e versioni salvate — e prosegue con il documento: +pagine con testo, parole, pagine verificate e avanzamento, letture OCR +completate, token e costo stimato. ## Limiti attuali diff --git a/docs/intro/getting-started.md b/docs/intro/getting-started.md index 74967470..d05a8210 100644 --- a/docs/intro/getting-started.md +++ b/docs/intro/getting-started.md @@ -21,14 +21,14 @@ verifica gli allegati della release scelta. Per usare un servizio di traduzione remoto occorrono le relative credenziali. Per l’elaborazione locale occorrono un server Ollama in esecuzione e un modello -già scaricato. La scelta si configura in **Impostazioni → Provider**. +già scaricato. La scelta si configura in **Impostazioni → Modelli**. ## Primo progetto di traduzione 1. Crea un workspace, cioè un gruppo di progetti e risorse condivise. 2. Crea un progetto nel workspace e importa un documento. 3. Controlla il testo estratto e la suddivisione in frammenti nell’anteprima. -4. Configura lingue, modalità della pipeline, provider e modelli per le fasi attive. +4. Indica le lingue dell’opera e configura modalità della pipeline, provider e modelli per le fasi attive. 5. Esegui una prova su un frammento rappresentativo e confronta il risultato con l’originale. 6. Avvia l’elaborazione degli altri frammenti, rivedi le traduzioni ed esporta il documento. diff --git a/docs/reference/backup-and-restore.md b/docs/reference/backup-and-restore.md index 81cf76c6..60dc8a3a 100644 --- a/docs/reference/backup-and-restore.md +++ b/docs/reference/backup-and-restore.md @@ -5,7 +5,7 @@ title: Backup e ripristino # Backup e ripristino Il backup salva i dati dell’applicazione in un file `.glossa-backup`. -L’operazione è disponibile in **Impostazioni → Backup** e comprende tutti +L’operazione è disponibile in **Impostazioni → Dati → Backup e ripristino** e comprende tutti i workspace. Il ripristino sostituisce i dati applicativi presenti con quelli del backup; non esegue un’unione e non importa un singolo workspace. diff --git a/docs/reference/import-export.md b/docs/reference/import-export.md index 2bd1fe0a..fb450f76 100644 --- a/docs/reference/import-export.md +++ b/docs/reference/import-export.md @@ -24,7 +24,7 @@ o RTF. Un testo non UTF-8 viene rifiutato con un errore di codifica. La finestra di sistema permette di scegliere file da qualsiasi cartella accessibile, anche su dischi esterni. L’anteprima consente di controllare -estrazione e segmentazione prima di confermare. Un PDF composto soltanto da +estrazione e segmentazione, e di indicare lingue e modello, prima di confermare. Un PDF composto soltanto da immagini non fornisce testo tramite questa estrazione: l’importazione non esegue OCR. diff --git a/docs/reference/pipeline-config.md b/docs/reference/pipeline-config.md index d553380b..02448490 100644 --- a/docs/reference/pipeline-config.md +++ b/docs/reference/pipeline-config.md @@ -7,30 +7,125 @@ title: Configurazione della pipeline La configurazione appartiene alla pipeline del progetto. Le credenziali e le connessioni ai servizi appartengono invece alle impostazioni dell’applicazione. -## Sezioni - -| Sezione | Parametri | +## La finestra + +La configurazione si apre con l’ingranaggio nella riga in cima allo Studio di +traduzione, o con Ctrl + virgola. Il titolo è il nome della pipeline: si +rinomina dalla riga in cima allo Studio, non da qui. Le spiegazioni delle voci +non sono scritte sotto i campi: compaiono passando sopra il titolo di una +sezione o il nome di una voce. + +La barra affianca **opera / pipeline**, con i nomi lunghi troncati e il nome +completo nel suggerimento. Il nome della pipeline, con la piccola freccia accanto, apre il menu +per scegliere, creare, rinominare o eliminare; l’ingranaggio apre le opzioni. +Il tipo **Semplice, Editoriale o DeepL** è sempre visibile con un’icona. Le +lingue dell’opera, comuni a tutte le sue pipeline, compaiono in forma breve nella barra +(vedi «Lingue dell’opera» nel [flusso di traduzione](../guides/document-pipeline)). Le operazioni sulla pipeline +restano bloccate durante l’elaborazione. + +| Linguetta | Parametri | | --- | --- | -| Impostazioni | Modalità, lingue, persona, esempi e opzioni generali | -| Traduzione | Provider, modelli, prompt e opzioni delle fasi di generazione | -| Controllo qualità | Valutatore e controllo di coerenza | -| Registro termini | Glossario assegnato | -| Anteprima Prompt | Struttura delle richieste delle fasi attive | +| Generale | Modalità, lingue DeepL, Contesto di traduzione | +| Fasi | Servizio, modello, prompt e opzioni di ogni fase; memoria di contesto | +| Controllo qualità | Ciclo di raffinamento, modello del giudizio, prompt di giudizio e coerenza | +| Memoria | Memoria delle frasi ed esempi di traduzione (spenta in modalità DeepL) | +| Glossario | Dizionario assegnato e suoi termini | +| Anteprima prompt | Ogni richiesta pezzo per pezzo: fasi, audit, coerenza, richiesta DeepL | + +Mentre la pipeline lavora la finestra resta aperta e leggibile, ma un velo ne +blocca i comandi. In fondo, l’icona rossa **Azzera tutte le traduzioni** cancella +le traduzioni e i relativi audit dopo una conferma; è spenta, con il motivo, +durante l’esecuzione o quando non c’è niente da azzerare. + +Quando esistono già traduzioni, modalità, prompt delle fasi e memoria di contesto +non si cambiano; il modello di una fase è chiuso da un lucchetto. Aprirlo +permette di cambiarlo, ma i frammenti già tradotti restano fatti con il modello +precedente. Le modalità Standard, Editoriale e DeepL Hybrid sono descritte nel [flusso di traduzione](../guides/document-pipeline). -## Lingue e persona - -Imposta lingua sorgente e destinazione. La persona è un testo libero che -sostituisce l’introduzione predefinita del messaggio di sistema: può specificare -ruolo, ambito, lingue e registro. Se è attiva, deve descrivere correttamente -la coppia linguistica e le istruzioni che intendi applicare. +## Contesto di traduzione e lingue DeepL + +Le pipeline non hanno una coppia di lingue: le lingue appartengono all’opera. +Per gli LLM la traduzione è guidata dal **Contesto di traduzione**: lingue e varietà storiche, destinatari, registro e obiettivo. È +obbligatorio e non è mai vuoto: una pipeline nuova parte da un testo +predefinito (dall’inglese all’italiano), che puoi riscrivere o sostituire con +un template salvato; la freccia circolare lo ripristina e la conferma di un +testo vuoto è spenta. Traduzione, revisione, audit e coerenza lo ricevono +insieme alle proprie istruzioni; la formattazione resta limitata alla +sintassi. Nessuna coppia linguistica viene aggiunta automaticamente ai prompt: +le lingue stanno nel Contesto. + +Accanto al titolo di ogni prompt compare da dove viene il testo: +**Predefinito**, **Personalizzato** oppure **Template «nome»** quando coincide +con un template salvato della stessa categoria. Il riconoscimento confronta il +testo: un template modificato dopo l’applicazione non è più riconosciuto, +perché la pipeline conserva la copia applicata. + +La matita apre una bozza; la spunta conferma, la X annulla. Anche la rifinitura +con un modello modifica solo la bozza. I prompt usano una carta tenue con +accento verde, comandi a icona e anteprima espandibile. Il Contesto viene +salvato con la pipeline e copiato nella duplicazione. + +La coppia **DeepL · lingue** resta visibile in Generale, disabilitata nelle +modalità LLM e attiva in DeepL. Le fasi non usate restano visibili come +linguette disabilitate. Cambiando modalità conservi la configurazione di tutte +le fasi. In DeepL restano attive revisione LLM, audit e coerenza. +Le lingue usate dalle memorie sono quelle dell’opera, non quelle DeepL. + +In **Anteprima prompt** scegli una fase, Audit o Coerenza: vedi la richiesta +pezzo per pezzo, nell’ordine in cui parte, divisa in **messaggio di sistema** +(istruzioni, regole e risorse, uguale per tutti i frammenti) e **messaggio +utente** (il frammento e la consegna finale). I pezzi vengono dal backend, gli +stessi che compongono la richiesta vera: anteprima e invio coincidono. Ogni +pezzo dice se è **fisso nel programma**, **testo tuo**, **dati** o +**automatico**, e dove si modifica; il titolo spiega a cosa serve. I pezzi che +in questa pipeline non ci sono restano visibili, spenti, con il motivo (per +esempio «solo per documenti importati come Markdown»). I segnaposto tra doppie +graffe indicano dove entrano i dati del frammento: è una costruzione con la +configurazione attuale, non la richiesta storica di un’esecuzione. Il modo +**Frammento aperto** riempie gli stessi pezzi con i dati veri del frammento +aperto nello Studio, comprese le frasi della memoria spuntate. + +Tutto il testo che il programma aggiunge attorno ai tuoi prompt — ruoli, +regole, intestazioni, consegne, formato della risposta, messaggio utente — è un +**prompt di sistema** e si modifica qui, in modo Struttura. Ogni prompt di +sistema ha il lucchetto chiuso: aprendolo diventa come gli altri prompt (bozza +con conferma e annullamento, bacchetta, template della categoria **Sistema**, +freccia di ripristino al predefinito). Il lucchetto si richiude quando chiudi +la finestra. Il formato della risposta di audit e coerenza chiede una conferma +in più, perché l’app legge la risposta secondo quel formato. Nel messaggio +utente e nelle cornici i segnaposto tra doppie graffe (per esempio +`{{TEXT}}`) sono obbligatori: senza, la conferma resta spenta. I pezzi il cui +contenuto sta in un’altra scheda (Contesto di traduzione, prompt delle fasi, +tabella del glossario, esempi) non si modificano qui: hanno solo la freccia che +apre quella scheda, e la loro intestazione resta quella predefinita. Le regole +del glossario sono invece un prompt di sistema e si modificano qui. I prompt si +leggono e si modificano sempre in carattere a spaziatura fissa. Ruolo, regole e cornici di Traduzione e Refine +sono in comune: modificarli una volta vale per entrambe. I testi cambiati si +salvano con la pipeline e passano alle pipeline nuove create per copia. + +I pezzi facoltativi si possono **spegnere fase per fase** con l’interruttore +nella loro carta: ruolo, regole strutturali, regole del glossario, regole +Markdown, esempi, frammenti vicini, tabella del glossario, metodo di controllo, +consegna del risultato. Un pezzo spento resta nell’elenco come riga grigia +«spento in questa fase», con l’interruttore per riaccenderlo; la scelta si salva +con la pipeline e vale anche nell’esecuzione. Restano sempre accesi il Contesto +di traduzione, il prompt della fase, il messaggio utente e il formato della +risposta. Spegnendo i frammenti vicini sparisce anche l’identificativo del +frammento, che serve solo a riconoscerlo fra quelli. DeepL +mostra il corpo della richiesta. I prompt possono essere salvati come modelli riutilizzabili, separati per -contesto. Il comando di rifinitura del prompt invia il testo a un modello +contesto: durante la modifica il segnalibro salva il prompt con un nome e il +libro apre l’elenco dei modelli salvati, con la ricerca. I modelli si +eliminano dalle risorse linguistiche. Un modello con un ambito non più +riconosciuto viene escluso dall’elenco e segnalato per nome, senza nascondere +gli altri. Il comando di rifinitura del prompt invia il testo a un modello configurato e ne propone una riscrittura nel campo. Richiede la connessione -e le eventuali credenziali del provider scelto. +e le eventuali credenziali del provider scelto: senza chiave il comando è +spento e il suggerimento dice quale manca. ## Parametri dei modelli @@ -53,23 +148,33 @@ Il significato delle opzioni del servizio va verificato per il modello utilizzat ## DeepL Hybrid -La prima fase usa le impostazioni DeepL: lingua, registro dove supportato, +DeepL esegue la traduzione iniziale, poi revisione LLM e audit. La prima fase usa le impostazioni DeepL: lingua, registro dove supportato, modalità di traduzione e glossario remoto. La chiave DeepL è distinta da quella degli LLM usati per revisione e valutazione. Un errore di quota o del glossario DeepL deve essere risolto su quel servizio prima di completare la sequenza. +La coppia DeepL si sceglie in Generale, dagli elenchi del servizio. La sorgente +può essere rilevata automaticamente; per usare o caricare un glossario serve +una sorgente esplicita. Cambiare coppia scollega il glossario remoto e cambiare +destinazione ripristina il registro predefinito. La destinazione va scelta esplicitamente: una richiesta senza destinazione viene bloccata prima di contattare DeepL. Il Contesto di traduzione non viene +inviato a DeepL: il suo campo Contesto è distinto. + +Anteprima prompt nelle opzioni e anteprima del +frammento mostrano il corpo API, costruito dal backend come durante l’esecuzione; +nelle opzioni il testo del frammento è un segnaposto. I log conservano la +richiesta effettiva senza credenziali. + ## Esempi e contesto Puoi mantenere fino a cinque esempi di traduzione nella pipeline, aggiunti -dalla scheda Audit di un frammento bloccato. Sono modificabili nelle impostazioni. +dalla scheda Audit di un frammento verificato. Sono modificabili nella linguetta Memoria. La [memoria di frasi](../guides/phrase-memory) fornisce invece riferimenti selezionati per il singolo frammento. La [cache dei prompt](../guides/context-and-caching) ha regole specifiche per provider. ## Stima dei costi -Il preventivo della configurazione copre l’intero documento, compresa la -coerenza se configurata. Nella barra del documento, la stima segue l’azione +Nella colonna degli strumenti dello Studio la stima segue l’azione selezionata. I dettagli distinguono le fasi e i modelli. La stima usa una conversione approssimativa da parole a token e i prezzi diff --git a/docs/reference/provider-support.md b/docs/reference/provider-support.md index 62fe8731..45ea6e10 100644 --- a/docs/reference/provider-support.md +++ b/docs/reference/provider-support.md @@ -21,7 +21,7 @@ dei modelli per marca. ## Configurazione delle credenziali -Apri **Impostazioni → Provider**. Le chiavi sono conservate nel portachiavi +Apri **Impostazioni → Modelli**. Le chiavi sono conservate nel portachiavi del sistema quando disponibile; se non è accessibile, Glossa usa un archivio locale cifrato. Non sono incluse nei backup dell’applicazione. diff --git a/docs/reference/troubleshooting.md b/docs/reference/troubleshooting.md index 88fe2784..b5faafee 100644 --- a/docs/reference/troubleshooting.md +++ b/docs/reference/troubleshooting.md @@ -10,7 +10,7 @@ verifichi una causa. ## Connessione o credenziali -Se una fase non parte, controlla in **Impostazioni → Provider** la chiave, +Se una fase non parte, controlla in **Impostazioni → Modelli** la chiave, il modello e, per Custom, il profilo selezionato. Un errore di autenticazione, un modello non disponibile e una quota esaurita richiedono interventi diversi. Per un endpoint remoto personalizzato è obbligatorio HTTPS; HTTP è ammesso diff --git a/e2e/support/tauriMock.js b/e2e/support/tauriMock.js index 1bd3d50b..71015b55 100644 --- a/e2e/support/tauriMock.js +++ b/e2e/support/tauriMock.js @@ -24,8 +24,12 @@ id: project.id, name: project.name, workspace_id: project.workspaceId, - source_language: project.sourceLanguage, - target_language: project.targetLanguage, + source_language: project.sourceLanguage ?? '', + source_language_variety: project.sourceLanguageVariety ?? null, + source_language_note: project.sourceLanguageNote ?? '', + target_language: project.targetLanguage ?? '', + target_language_variety: project.targetLanguageVariety ?? null, + target_language_note: project.targetLanguageNote ?? '', source_display_text: project.sourceDisplayText, source_processing_text: project.sourceProcessingText, source_footnotes: '[]', @@ -43,8 +47,6 @@ id: pipeline.id, project_id: pipeline.projectId, name: 'Default', - source_language: pipeline.sourceLanguage, - target_language: pipeline.targetLanguage, pipeline_mode: 'standard', stages: '[]', judge_prompt: '', @@ -126,9 +128,13 @@ state.projects.push({ id: params[0], name: params[1], - sourceLanguage: params[2], - targetLanguage: params[3], - workspaceId: params[4], + workspaceId: params[2], + sourceLanguage: params[3], + sourceLanguageVariety: params[4], + sourceLanguageNote: params[5], + targetLanguage: params[6], + targetLanguageVariety: params[7], + targetLanguageNote: params[8], sourceDisplayText: '', sourceProcessingText: '', createdAt: timestamp, @@ -140,8 +146,6 @@ state.pipelines.push({ id: params[0], projectId: params[1], - sourceLanguage: params[2], - targetLanguage: params[3], }); return; } @@ -150,7 +154,7 @@ return; } if (sql.startsWith('UPDATE PROJECTS SET SOURCE_DISPLAY_TEXT')) { - const project = state.projects.find((candidate) => candidate.id === params[9]); + const project = state.projects.find((candidate) => candidate.id === params[13]); if (project) { project.sourceDisplayText = params[0]; project.sourceProcessingText = params[1]; diff --git a/e2e/workspace.spec.ts b/e2e/workspace.spec.ts index 5d10a9c2..55355679 100644 --- a/e2e/workspace.spec.ts +++ b/e2e/workspace.spec.ts @@ -34,6 +34,6 @@ test('crea un progetto e apre la sua schermata di importazione', async ({ page } await page.getByPlaceholder('Nome del progetto...').fill('Manoscritto E2E'); await page.getByRole('button', { name: 'Crea', exact: true }).click(); - await expect(page.getByRole('heading', { name: 'Manoscritto E2E' })).toBeVisible(); + await expect(page.getByRole('heading', { name: 'Manoscritto E2E', level: 1 })).toBeVisible(); await expect(page.getByRole('button', { name: 'Importa documento' })).toBeVisible(); }); diff --git a/scripts/update-languages.ts b/scripts/update-languages.ts new file mode 100644 index 00000000..7e7a55a3 --- /dev/null +++ b/scripts/update-languages.ts @@ -0,0 +1,41 @@ +#!/usr/bin/env -S npx tsx +// Rigenera l'elenco delle lingue incluso nell'app dalle fonti ufficiali, con la +// stessa costruzione del comando «Aggiorna elenco lingue»: i codici spariti dalle +// fonti restano, segnati come ritirati. Uso: npx tsx scripts/update-languages.ts +import { readFile, writeFile } from 'node:fs/promises'; +import { dirname, join } from 'node:path'; +import { fileURLToPath } from 'node:url'; +import { + buildLanguageListUpdate, + cldrItalianNames, + GLOTTOLOG_SOURCE_URL, + ISO_SOURCE_URL, + parseIsoFile, + parseVarietiesFile, +} from '../src/languages/build'; + +const OUT_DIR = join(dirname(fileURLToPath(import.meta.url)), '..', 'src', 'languages', 'data'); +const ISO_PATH = join(OUT_DIR, 'iso639-3.json'); +const VARIETIES_PATH = join(OUT_DIR, 'glottolog-varieties.json'); + +async function fetchText(url: string): Promise { + const response = await fetch(url); + if (!response.ok) throw new Error(`${url}: HTTP ${response.status}`); + return response.text(); +} + +const [isoTab, glottologCsv, currentIso, currentVarieties] = await Promise.all([ + fetchText(ISO_SOURCE_URL), + fetchText(GLOTTOLOG_SOURCE_URL), + readFile(ISO_PATH, 'utf8'), + readFile(VARIETIES_PATH, 'utf8'), +]); +const update = buildLanguageListUpdate( + { isoTab, glottologCsv }, + { iso: parseIsoFile(currentIso), varieties: parseVarietiesFile(currentVarieties) }, + cldrItalianNames(), + new Date().toISOString().slice(0, 10), +); +await writeFile(ISO_PATH, JSON.stringify(update.iso)); +await writeFile(VARIETIES_PATH, JSON.stringify(update.varieties)); +console.log(`${update.iso.languages.length} lingue (${update.added} nuove, ${update.retired} ritirate ora) in ${OUT_DIR}`); diff --git a/src-tauri/migrations.lock b/src-tauri/migrations.lock index 57cedb46..c6f15bd0 100644 --- a/src-tauri/migrations.lock +++ b/src-tauri/migrations.lock @@ -30,3 +30,11 @@ # della trascrizione nella baseline, prima del merge. Database locale ricreato. 0001_baseline_2_0.sql 78fd5b2f9c0d3588 +0002_workspace_memory_scope.sql c55dcc7e778fe168 + +0003_text_corpus.sql a37f19b0640b0551 +0004_pipeline_work_brief.sql ef82af7e7b69dd3d +0005_pipeline_shared_context.sql 5c739564ba240f04 +0006_pipeline_prompt_composition.sql fb72574825df2ba7 +0007_work_languages.sql 676a265773629b03 +0008_text_language_variety.sql bcf89a39ab85c6ba diff --git a/src-tauri/migrations/0002_workspace_memory_scope.sql b/src-tauri/migrations/0002_workspace_memory_scope.sql new file mode 100644 index 00000000..9489b15b --- /dev/null +++ b/src-tauri/migrations/0002_workspace_memory_scope.sql @@ -0,0 +1,4 @@ +-- La ricerca nella memoria di frasi di un workspace può includere anche le +-- frasi degli altri workspace e delle traduzioni senza workspace (#485 N, T7e). +-- Spenta per impostazione predefinita: la memoria resta del workspace. +ALTER TABLE workspaces ADD COLUMN memory_search_all_workspaces INTEGER NOT NULL DEFAULT 0; diff --git a/src-tauri/migrations/0003_text_corpus.sql b/src-tauri/migrations/0003_text_corpus.sql new file mode 100644 index 00000000..f63c7441 --- /dev/null +++ b/src-tauri/migrations/0003_text_corpus.sql @@ -0,0 +1,212 @@ +-- Incremental corpus upgrade. Existing baseline and its checksum stay unchanged. +-- SQLx runs the entire migration in a transaction; no reset or paid API calls. +ALTER TABLE phrase_memory RENAME TO phrase_memory_previous; +DROP INDEX idx_phrase_memory_chunk_project; + +-- Unità testuali e revisioni indipendenti dalla memoria traduttiva. +CREATE TABLE IF NOT EXISTS text_units ( + id TEXT PRIMARY KEY, + source_id TEXT REFERENCES sources(id) ON DELETE SET NULL, + source_version_id TEXT REFERENCES source_versions(id) ON DELETE SET NULL, + parent_unit_id TEXT REFERENCES text_units(id) ON DELETE SET NULL, + workspace_id TEXT REFERENCES workspaces(id) ON DELETE SET NULL, + provenance TEXT NOT NULL DEFAULT '{}' CHECK (json_valid(provenance)), + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP +); + +CREATE TABLE IF NOT EXISTS text_unit_revisions ( + id TEXT PRIMARY KEY, + unit_id TEXT NOT NULL REFERENCES text_units(id) ON DELETE CASCADE, + role TEXT NOT NULL CHECK (role IN ('source', 'translation', 'normalized')), + language TEXT NOT NULL CHECK (length(trim(language)) > 0), + text TEXT NOT NULL CHECK (length(trim(text)) > 0), + content_hash TEXT NOT NULL CHECK (length(content_hash) > 0), + revision_number INTEGER NOT NULL CHECK (revision_number > 0), + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, + UNIQUE(unit_id, role, language, revision_number), + UNIQUE(unit_id, id) +); + +CREATE TRIGGER IF NOT EXISTS text_unit_revisions_immutable +BEFORE UPDATE ON text_unit_revisions +BEGIN + SELECT RAISE(ABORT, 'Text revisions are immutable; create a new revision'); +END; + +CREATE TABLE IF NOT EXISTS text_embeddings ( + revision_id TEXT NOT NULL REFERENCES text_unit_revisions(id) ON DELETE CASCADE, + provider TEXT NOT NULL CHECK (length(trim(provider)) > 0), + model TEXT NOT NULL CHECK (length(trim(model)) > 0), + dimensions INTEGER NOT NULL CHECK (dimensions > 0), + profile TEXT NOT NULL CHECK (length(trim(profile)) > 0), + embedding BLOB NOT NULL CHECK (length(embedding) = dimensions * 4), + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, + PRIMARY KEY(revision_id, provider, model, dimensions, profile) +); +CREATE INDEX IF NOT EXISTS idx_text_embeddings_profile + ON text_embeddings(provider, model, dimensions, profile, revision_id); + +CREATE TABLE IF NOT EXISTS text_unit_tags ( + unit_id TEXT NOT NULL REFERENCES text_units(id) ON DELETE CASCADE, + name TEXT NOT NULL COLLATE NOCASE CHECK (length(trim(name)) > 0), + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, + PRIMARY KEY(unit_id, name) +); + +CREATE TABLE IF NOT EXISTS phrase_memory ( + id TEXT PRIMARY KEY, + unit_id TEXT NOT NULL UNIQUE REFERENCES text_units(id) ON DELETE CASCADE, + source_revision_id TEXT NOT NULL, + target_revision_id TEXT NOT NULL, + confidence REAL NOT NULL DEFAULT 1.0 CHECK (confidence BETWEEN 0 AND 1), + author TEXT, + work TEXT, + domain TEXT, + notes TEXT, + chunk_id TEXT, + project_id TEXT REFERENCES projects(id) ON DELETE SET NULL, + created_at TEXT NOT NULL DEFAULT CURRENT_TIMESTAMP, + FOREIGN KEY(unit_id, source_revision_id) REFERENCES text_unit_revisions(unit_id, id), + FOREIGN KEY(unit_id, target_revision_id) REFERENCES text_unit_revisions(unit_id, id) +); +CREATE INDEX IF NOT EXISTS idx_phrase_memory_chunk_project ON phrase_memory(chunk_id, project_id); + + +-- Provenance is reconstructed only from recorded links; no filename inference. +INSERT INTO text_units(id, source_id, source_version_id, workspace_id, provenance, created_at) +SELECT 'unit:' || pm.id, sv.source_id, sv.id, p.workspace_id, + json_object('projectId', pm.project_id, 'projectName', p.name, + 'chunkId', pm.chunk_id, 'chunkPosition', t.position, + 'workspaceId', p.workspace_id, 'workspaceName', w.name, + 'sourceId', sv.source_id, 'sourceVersionId', sv.id, + 'sourceTitle', s.title, 'sourceVersionLabel', sv.label, + 'quote', pm.source_phrase), pm.created_at +FROM phrase_memory_previous pm +LEFT JOIN projects p ON p.id = pm.project_id +LEFT JOIN workspaces w ON w.id = p.workspace_id +LEFT JOIN translations t ON t.id = pm.chunk_id AND t.project_id = pm.project_id +LEFT JOIN translation_origins o ON o.project_id = pm.project_id +LEFT JOIN transcription_documents d ON d.id = o.transcription_document_id +LEFT JOIN source_versions sv ON sv.id = COALESCE(o.source_version_id, d.source_version_id) +LEFT JOIN sources s ON s.id = sv.source_id; + +-- FNV-1a over UTF-8 bytes, identical to the native revision writer. +-- Two unsigned 32-bit halves avoid SQLite's signed overflow-to-REAL behavior. +WITH RECURSIVE +texts(id, unit_id, role, language, text, created_at) AS ( + SELECT 'source:' || id, 'unit:' || id, 'source', source_language, source_phrase, created_at + FROM phrase_memory_previous + UNION ALL + SELECT 'target:' || id, 'unit:' || id, 'translation', target_language, target_phrase, created_at + FROM phrase_memory_previous +), +bytes(id, unit_id, role, language, text, created_at, encoded) AS ( + SELECT *, hex(CAST(text AS BLOB)) FROM texts +), +hashes(id, position, low, high) AS ( + SELECT id, 0, 2216829733, 3421674724 FROM bytes + UNION ALL + SELECT h.id, h.position + 1, + (((h.low | ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1)) + - (h.low & ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1))) * 435) % 4294967296, + (h.high * 435 + + ((h.low | ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1)) + - (h.low & ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1))) * 256 + + (((h.low | ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1)) + - (h.low & ((instr('0123456789ABCDEF', substr(b.encoded, h.position*2+1, 1))-1)*16 + + instr('0123456789ABCDEF', substr(b.encoded, h.position*2+2, 1))-1))) * 435) / 4294967296 + ) % 4294967296 + FROM hashes h JOIN bytes b ON b.id = h.id + WHERE h.position < length(b.encoded)/2 +) +INSERT INTO text_unit_revisions(id, unit_id, role, language, text, content_hash, revision_number, created_at) +SELECT b.id, b.unit_id, b.role, b.language, b.text, + printf('%08x%08x', h.high, h.low), 1, b.created_at +FROM bytes b JOIN hashes h ON h.id = b.id AND h.position = length(b.encoded)/2; + +INSERT INTO phrase_memory(id, unit_id, source_revision_id, target_revision_id, + confidence, author, work, domain, notes, chunk_id, project_id, created_at) +SELECT id, 'unit:' || id, 'source:' || id, 'target:' || id, + confidence, author, work, domain, notes, chunk_id, project_id, created_at +FROM phrase_memory_previous; + +-- Only explicitly recorded, supported models with the correct vector size. +-- Unidentified vectors do not acquire a guessed model. Their texts stay archived. +INSERT INTO text_embeddings(revision_id, provider, model, dimensions, profile, embedding, created_at) +SELECT 'source:' || id, 'openai', embedding_model, length(embedding)/4, + 'source-verbatim-v1', embedding, created_at +FROM phrase_memory_previous +WHERE (embedding_model = 'text-embedding-3-small' AND length(embedding) = 1536*4) + OR (embedding_model = 'text-embedding-3-large' AND length(embedding) = 3072*4); + +INSERT OR IGNORE INTO text_unit_tags(unit_id, name, created_at) +SELECT 'unit:' || pm.id, trim(CAST(tag.value AS TEXT)), pm.created_at +FROM phrase_memory_previous pm, + json_each(CASE WHEN json_valid(pm.tags) THEN + CASE WHEN json_type(pm.tags) = 'array' THEN pm.tags ELSE json_array(pm.tags) END + ELSE json_array(pm.tags) END) tag +WHERE tag.type = 'text' AND length(trim(CAST(tag.value AS TEXT))) > 0; + +DROP TABLE phrase_memory_previous; +DROP TABLE source_phrase_embeddings; + +CREATE VIEW IF NOT EXISTS phrase_memory_entries AS +SELECT pm.*, sr.text AS source_phrase, tr.text AS target_phrase, + sr.language AS source_language, tr.language AS target_language, + tu.source_id, tu.source_version_id, tu.provenance, + (SELECT p.workspace_id FROM projects p WHERE p.id = pm.project_id) AS workspace_id +FROM phrase_memory pm +JOIN text_units tu ON tu.id = pm.unit_id +JOIN text_unit_revisions sr ON sr.id = pm.source_revision_id +JOIN text_unit_revisions tr ON tr.id = pm.target_revision_id; + + +CREATE TABLE provenance_events_corpus ( + id TEXT PRIMARY KEY, + occurred_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP, + event_type TEXT NOT NULL, + entity_type TEXT NOT NULL CHECK ( + entity_type IN ( + 'source', 'source_version', 'transcription_document', 'transcription_segment', + 'transcription_revision', 'project', 'translation_chunk', 'artifact', 'job', + 'text_unit', 'text_revision' + ) + ), + entity_id TEXT NOT NULL, + workspace_id TEXT REFERENCES workspaces(id) ON DELETE SET NULL, + actor TEXT NOT NULL DEFAULT 'user' CHECK (actor IN ('user', 'system', 'model')), + job_id TEXT REFERENCES jobs(id) ON DELETE SET NULL, + input_ref TEXT, + output_ref TEXT, + config TEXT, + outcome TEXT, + duration_ms INTEGER, + provider TEXT, + model TEXT, + prompt_version TEXT, + input_tokens INTEGER, + output_tokens INTEGER, + cached_tokens INTEGER, + estimated_cost REAL, + source_language TEXT, + target_language TEXT, + error_kind TEXT, + input_hash TEXT, + output_hash TEXT +); + +INSERT INTO provenance_events_corpus SELECT * FROM provenance_events; +DROP TABLE provenance_events; +ALTER TABLE provenance_events_corpus RENAME TO provenance_events; + +CREATE INDEX IF NOT EXISTS idx_provenance_entity ON provenance_events(entity_type, entity_id, occurred_at); +CREATE INDEX IF NOT EXISTS idx_provenance_workspace ON provenance_events(workspace_id, occurred_at); +CREATE INDEX IF NOT EXISTS idx_provenance_type_time ON provenance_events(event_type, occurred_at); +CREATE INDEX IF NOT EXISTS idx_provenance_model ON provenance_events(model, occurred_at); +CREATE INDEX IF NOT EXISTS idx_provenance_job ON provenance_events(job_id); + diff --git a/src-tauri/migrations/0004_pipeline_work_brief.sql b/src-tauri/migrations/0004_pipeline_work_brief.sql new file mode 100644 index 00000000..2bdc0dc1 --- /dev/null +++ b/src-tauri/migrations/0004_pipeline_work_brief.sql @@ -0,0 +1,2 @@ +-- Optional shared task context. Existing pipelines retain their legacy prompts. +ALTER TABLE pipelines ADD COLUMN work_brief TEXT; diff --git a/src-tauri/migrations/0005_pipeline_shared_context.sql b/src-tauri/migrations/0005_pipeline_shared_context.sql new file mode 100644 index 00000000..5ef10129 --- /dev/null +++ b/src-tauri/migrations/0005_pipeline_shared_context.sql @@ -0,0 +1,3 @@ +ALTER TABLE pipelines DROP COLUMN persona; +ALTER TABLE pipelines DROP COLUMN custom_source_language; +ALTER TABLE pipelines DROP COLUMN custom_target_language; diff --git a/src-tauri/migrations/0006_pipeline_prompt_composition.sql b/src-tauri/migrations/0006_pipeline_prompt_composition.sql new file mode 100644 index 00000000..7c6295e4 --- /dev/null +++ b/src-tauri/migrations/0006_pipeline_prompt_composition.sql @@ -0,0 +1,3 @@ +-- Per-pipeline prompt composition: custom system texts by id and parts switched +-- off, as JSON {"texts": {id: text}, "disabled": [part ids]}. NULL = all defaults. +ALTER TABLE pipelines ADD COLUMN prompt_composition TEXT; diff --git a/src-tauri/migrations/0007_work_languages.sql b/src-tauri/migrations/0007_work_languages.sql new file mode 100644 index 00000000..6e7d0eed --- /dev/null +++ b/src-tauri/migrations/0007_work_languages.sql @@ -0,0 +1,23 @@ +-- Le lingue appartengono all'opera, non alla pipeline (PIANO_PROMPT_PIPELINE §12). +-- projects.source_language / target_language passano dai nomi inglesi a codici +-- ISO 639-3 ('' = non indicata); varietà Glottolog e nota libera accanto. +-- Solo ADD/DROP COLUMN e UPDATE: nessuna tabella ricreata. +ALTER TABLE projects ADD COLUMN source_language_variety TEXT; +ALTER TABLE projects ADD COLUMN source_language_note TEXT NOT NULL DEFAULT ''; +ALTER TABLE projects ADD COLUMN target_language_variety TEXT; +ALTER TABLE projects ADD COLUMN target_language_note TEXT NOT NULL DEFAULT ''; + +UPDATE projects SET + source_language = CASE source_language + WHEN 'English' THEN 'eng' WHEN 'Italian' THEN 'ita' WHEN 'Spanish' THEN 'spa' + WHEN 'French' THEN 'fra' WHEN 'German' THEN 'deu' WHEN 'Portuguese' THEN 'por' + WHEN 'Japanese' THEN 'jpn' WHEN 'Chinese' THEN 'zho' WHEN 'Korean' THEN 'kor' + WHEN 'Russian' THEN 'rus' ELSE '' END, + target_language = CASE target_language + WHEN 'English' THEN 'eng' WHEN 'Italian' THEN 'ita' WHEN 'Spanish' THEN 'spa' + WHEN 'French' THEN 'fra' WHEN 'German' THEN 'deu' WHEN 'Portuguese' THEN 'por' + WHEN 'Japanese' THEN 'jpn' WHEN 'Chinese' THEN 'zho' WHEN 'Korean' THEN 'kor' + WHEN 'Russian' THEN 'rus' ELSE '' END; + +ALTER TABLE pipelines DROP COLUMN source_language; +ALTER TABLE pipelines DROP COLUMN target_language; diff --git a/src-tauri/migrations/0008_text_language_variety.sql b/src-tauri/migrations/0008_text_language_variety.sql new file mode 100644 index 00000000..4b8cd09f --- /dev/null +++ b/src-tauri/migrations/0008_text_language_variety.sql @@ -0,0 +1,5 @@ +-- Varietà Glottolog e nota libera accanto al codice ISO di ogni revisione di +-- testo (PIANO_PROMPT_PIPELINE §12). Le revisioni restano immutabili: una +-- lingua corretta è una revisione nuova con lo stesso testo. +ALTER TABLE text_unit_revisions ADD COLUMN language_variety TEXT; +ALTER TABLE text_unit_revisions ADD COLUMN language_note TEXT NOT NULL DEFAULT ''; diff --git a/src-tauri/src/db.rs b/src-tauri/src/db.rs index 0c9eaf89..694d0e58 100644 --- a/src-tauri/src/db.rs +++ b/src-tauri/src/db.rs @@ -191,7 +191,7 @@ pub fn backup_database_file( /// A JSON array of small non-negative integers is how the frontend represents /// a BLOB column (tauri-plugin-sql/JSON can't carry raw bytes otherwise) — -/// e.g. phrase_memory.embedding, round-tripped through a workspace backup. +/// e.g. text_embeddings.embedding, round-tripped through a workspace backup. /// Without this, such arrays fell through to the generic `JsonValue` bind /// below, which sqlx serializes as JSON *text* instead of raw bytes, silently /// corrupting the embedding for any semantic search that reads it back. @@ -270,6 +270,11 @@ mod tests { "transcription_segments", "transcription_revisions", "translation_origins", + "text_units", + "text_unit_revisions", + "text_embeddings", + "text_unit_tags", + "phrase_memory", "jobs", "artifacts", "provenance_events", diff --git a/src-tauri/src/deepl/client.rs b/src-tauri/src/deepl/client.rs index a370c55c..2144a3c4 100644 --- a/src-tauri/src/deepl/client.rs +++ b/src-tauri/src/deepl/client.rs @@ -26,23 +26,42 @@ fn glossaries_endpoint(api_key: &str) -> String { format!("https://{}/v3/glossaries", deepl_host(api_key)) } +pub(crate) fn build_translate_request( + input: &DeeplStageInput, +) -> Result { + let cfg = &input.deepl_config; + let target_lang = cfg + .target_lang + .as_deref() + .unwrap_or("") + .trim() + .to_uppercase(); + if target_lang.is_empty() { + return Err("Choose a target language in the DeepL settings.".to_string()); + } + Ok(DeeplTranslateRequest { + text: vec![input.text.clone()], + source_lang: cfg + .source_lang + .as_ref() + .filter(|s| !s.trim().is_empty()) + .map(|s| s.trim().to_uppercase()), + target_lang, + model_type: cfg.model_type.clone(), + formality: cfg.formality.clone(), + context: cfg.context.clone(), + preserve_formatting: cfg.preserve_formatting, + glossary_id: cfg.glossary_id.clone(), + show_billed_characters: cfg.show_billed_characters.unwrap_or(true), + }) +} + pub async fn translate( client: &Client, api_key: &str, input: &DeeplStageInput, ) -> Result { - let cfg = input.deepl_config.as_ref(); - let body = DeeplTranslateRequest { - text: vec![input.text.clone()], - source_lang: input.source_lang.clone(), - target_lang: input.target_lang.clone(), - model_type: cfg.and_then(|c| c.model_type.clone()), - formality: cfg.and_then(|c| c.formality.clone()), - context: cfg.and_then(|c| c.context.clone()), - preserve_formatting: cfg.and_then(|c| c.preserve_formatting), - glossary_id: cfg.and_then(|c| c.glossary_id.clone()), - show_billed_characters: cfg.and_then(|c| c.show_billed_characters).unwrap_or(true), - }; + let body = build_translate_request(input)?; let url = translate_endpoint(api_key); let resp = client @@ -278,4 +297,57 @@ mod tests { assert!(url.contains("api.deepl.com"), "got: {url}"); assert!(!url.contains("api-free"), "got: {url}"); } + + fn stage_input(source: Option<&str>, target: Option<&str>) -> DeeplStageInput { + DeeplStageInput { + text: "Lorem ipsum".to_string(), + deepl_config: crate::deepl::types::DeeplConfig { + source_lang: source.map(str::to_string), + target_lang: target.map(str::to_string), + model_type: None, + formality: Some("more".to_string()), + context: None, + preserve_formatting: Some(true), + glossary_id: Some("g-1".to_string()), + show_billed_characters: None, + }, + } + } + + #[test] + fn build_translate_request_rejects_missing_or_blank_target() { + assert!(build_translate_request(&stage_input(Some("la"), None)).is_err()); + assert!(build_translate_request(&stage_input(Some("la"), Some(" "))).is_err()); + } + + #[test] + fn build_translate_request_normalizes_language_codes() { + let request = build_translate_request(&stage_input(Some(" la "), Some(" it "))) + .expect("valid request"); + assert_eq!(request.source_lang.as_deref(), Some("LA")); + assert_eq!(request.target_lang, "IT"); + } + + #[test] + fn build_translate_request_omits_blank_source_for_auto_detection() { + let request = + build_translate_request(&stage_input(Some(" "), Some("en"))).expect("valid request"); + let body = serde_json::to_value(&request).expect("serializable"); + assert!(body.get("source_lang").is_none(), "got: {body}"); + } + + #[test] + fn build_translate_request_serializes_deepl_body_fields() { + let request = + build_translate_request(&stage_input(None, Some("de"))).expect("valid request"); + let body = serde_json::to_value(&request).expect("serializable"); + assert_eq!(body["text"], serde_json::json!(["Lorem ipsum"])); + assert_eq!(body["target_lang"], "DE"); + assert_eq!(body["formality"], "more"); + assert_eq!(body["preserve_formatting"], true); + assert_eq!(body["glossary_id"], "g-1"); + assert_eq!(body["show_billed_characters"], true); + assert!(body.get("model_type").is_none(), "got: {body}"); + assert!(body.get("context").is_none(), "got: {body}"); + } } diff --git a/src-tauri/src/deepl/commands.rs b/src-tauri/src/deepl/commands.rs index 36da4c64..e1fa8802 100644 --- a/src-tauri/src/deepl/commands.rs +++ b/src-tauri/src/deepl/commands.rs @@ -25,6 +25,11 @@ fn get_deepl_api_key(app: &AppHandle) -> Result { }) } +#[tauri::command] +pub fn preview_deepl_stage(input: DeeplStageInput) -> Result { + serde_json::to_value(client::build_translate_request(&input)?).map_err(|err| err.to_string()) +} + #[tauri::command] pub async fn run_deepl_stage( app: AppHandle, diff --git a/src-tauri/src/deepl/types.rs b/src-tauri/src/deepl/types.rs index 6dd6fb96..06b27c08 100644 --- a/src-tauri/src/deepl/types.rs +++ b/src-tauri/src/deepl/types.rs @@ -4,6 +4,8 @@ use serde::{Deserialize, Serialize}; #[derive(Debug, Clone, Serialize, Deserialize)] #[serde(rename_all = "camelCase")] pub struct DeeplConfig { + pub source_lang: Option, + pub target_lang: Option, pub model_type: Option, pub formality: Option, pub context: Option, @@ -17,9 +19,7 @@ pub struct DeeplConfig { #[serde(rename_all = "camelCase")] pub struct DeeplStageInput { pub text: String, - pub source_lang: Option, - pub target_lang: String, - pub deepl_config: Option, + pub deepl_config: DeeplConfig, } // Output di run_deepl_stage diff --git a/src-tauri/src/jobs/engine.rs b/src-tauri/src/jobs/engine.rs index 102f2ff1..a4306e39 100644 --- a/src-tauri/src/jobs/engine.rs +++ b/src-tauri/src/jobs/engine.rs @@ -48,6 +48,8 @@ use super::{ const IDLE_TICK: Duration = Duration::from_millis(500); /// Un tipo di lavoro. La logica lunga sta qui dentro, non nell'interfaccia. +// async_trait adds must_use to boxed futures; Clippy 1.99 flags that generated attribute. +#[allow(clippy::double_must_use)] #[async_trait] pub trait JobHandler: Send + Sync { /// Con quale limite compete. diff --git a/src-tauri/src/languages.rs b/src-tauri/src/languages.rs new file mode 100644 index 00000000..008d54f8 --- /dev/null +++ b/src-tauri/src/languages.rs @@ -0,0 +1,156 @@ +//! Elenco delle lingue scaricato dalle fonti ufficiali. +//! +//! L'app include un elenco ISO 639-3 + Glottolog; «Aggiorna elenco lingue» ne +//! scarica uno nuovo. Qui solo il trasporto: scaricare i due file sorgente (dal +//! backend, fuori dai limiti CORS della webview) e conservare nella cartella dei +//! dati l'elenco già costruito dal frontend, che lo preferisce a quello incluso. + +use std::fs; +use std::path::{Path, PathBuf}; + +use serde::{Deserialize, Serialize}; + +use crate::llm::stream::shared_cloud_http_client; + +const ISO_URL: &str = "https://iso639-3.sil.org/sites/iso639-3/files/downloads/iso-639-3.tab"; +const GLOTTOLOG_URL: &str = + "https://raw.githubusercontent.com/glottolog/glottolog-cldf/master/cldf/languages.csv"; +const LANGUAGES_DIR: &str = "languages"; +const ISO_FILE: &str = "iso639-3.json"; +const VARIETIES_FILE: &str = "glottolog-varieties.json"; + +#[derive(Serialize)] +#[serde(rename_all = "camelCase")] +pub struct LanguageSources { + iso_tab: String, + glottolog_csv: String, +} + +/// I due elenchi nel formato incluso nell'app (JSON come testo). +#[derive(Debug, PartialEq, Serialize, Deserialize)] +pub struct LanguageLists { + iso: String, + varieties: String, +} + +async fn fetch_text(client: &reqwest::Client, url: &str) -> Result { + let response = client + .get(url) + .send() + .await + .map_err(|error| format!("{url}: {error}"))?; + let status = response.status(); + if !status.is_success() { + return Err(format!("{url}: HTTP {status}")); + } + response + .text() + .await + .map_err(|error| format!("{url}: {error}")) +} + +#[tauri::command] +pub async fn languages_fetch_sources() -> Result { + let client = shared_cloud_http_client()?; + let (iso_tab, glottolog_csv) = tokio::try_join!( + fetch_text(&client, ISO_URL), + fetch_text(&client, GLOTTOLOG_URL) + )?; + Ok(LanguageSources { + iso_tab, + glottolog_csv, + }) +} + +fn languages_dir(app: &tauri::AppHandle) -> Result { + Ok(crate::storage_config::resolve_data_dir(app)?.join(LANGUAGES_DIR)) +} + +fn read_lists(dir: &Path) -> Result, String> { + let iso_path = dir.join(ISO_FILE); + let varieties_path = dir.join(VARIETIES_FILE); + if !iso_path.exists() || !varieties_path.exists() { + return Ok(None); + } + let read = |path: &Path| fs::read_to_string(path).map_err(|error| error.to_string()); + Ok(Some(LanguageLists { + iso: read(&iso_path)?, + varieties: read(&varieties_path)?, + })) +} + +/// Ogni file passa da una copia temporanea: un'interruzione lascia l'elenco precedente intero. +fn write_lists(dir: &Path, lists: &LanguageLists) -> Result<(), String> { + for text in [&lists.iso, &lists.varieties] { + serde_json::from_str::>(text) + .map_err(|error| format!("Invalid language list: {error}"))?; + } + fs::create_dir_all(dir).map_err(|error| error.to_string())?; + for (name, text) in [(ISO_FILE, &lists.iso), (VARIETIES_FILE, &lists.varieties)] { + let target = dir.join(name); + let temporary = dir.join(format!("{name}.tmp")); + fs::write(&temporary, text).map_err(|error| error.to_string())?; + fs::rename(&temporary, &target).map_err(|error| error.to_string())?; + } + Ok(()) +} + +#[tauri::command] +pub fn languages_read_saved(app: tauri::AppHandle) -> Result, String> { + read_lists(&languages_dir(&app)?) +} + +#[tauri::command] +pub fn languages_save(app: tauri::AppHandle, lists: LanguageLists) -> Result<(), String> { + write_lists(&languages_dir(&app)?, &lists) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn temp_dir(name: &str) -> PathBuf { + let dir = + std::env::temp_dir().join(format!("glossa-languages-{name}-{}", std::process::id())); + let _ = fs::remove_dir_all(&dir); + dir + } + + fn lists() -> LanguageLists { + LanguageLists { + iso: r#"{"retrievedAt":"2026-10-07","languages":[["ita","Italian","Italiano",0]]}"# + .into(), + varieties: r#"{"retrievedAt":"2026-10-07","varieties":{}}"#.into(), + } + } + + #[test] + fn reading_without_saved_lists_returns_none() -> Result<(), String> { + let dir = temp_dir("empty"); + assert_eq!(read_lists(&dir)?, None); + Ok(()) + } + + #[test] + fn saved_lists_are_read_back_unchanged() -> Result<(), String> { + let dir = temp_dir("roundtrip"); + write_lists(&dir, &lists())?; + assert_eq!(read_lists(&dir)?, Some(lists())); + let _ = fs::remove_dir_all(&dir); + Ok(()) + } + + #[test] + fn invalid_json_is_rejected_and_previous_lists_stay() -> Result<(), String> { + let dir = temp_dir("invalid"); + write_lists(&dir, &lists())?; + let broken = LanguageLists { + iso: "not json".into(), + varieties: lists().varieties, + }; + assert!(write_lists(&dir, &broken).is_err()); + assert_eq!(read_lists(&dir)?, Some(lists())); + let _ = fs::remove_dir_all(&dir); + Ok(()) + } +} diff --git a/src-tauri/src/lib.rs b/src-tauri/src/lib.rs index c3c3dcf4..2eb7a701 100644 --- a/src-tauri/src/lib.rs +++ b/src-tauri/src/lib.rs @@ -10,6 +10,7 @@ mod iiif; mod images; mod jobs; mod keystore; +mod languages; mod llm; mod ocr; mod optimize; @@ -172,10 +173,16 @@ pub fn run() { db::execute_transaction, storage_config::get_data_dir, storage_config::choose_data_dir_folder, + languages::languages_fetch_sources, + languages::languages_read_saved, + languages::languages_save, llm::pipeline::compute_blobs, llm::pipeline::run_stage, llm::pipeline::run_stage_stream, llm::pipeline::preview_stage_prompt, + llm::pipeline::preview_judge_prompt, + llm::pipeline::preview_coherence_prompt, + llm::prompt_texts::prompt_system_texts, llm::pipeline::cancel_stream, llm::pipeline::judge_translation, llm::pipeline::refine_prompt, @@ -227,13 +234,19 @@ pub fn run() { documents::export_markdown_docx, vector::vec_ping, vector::embedding::get_embeddings, - vector::embedding::vec_list_phrase_memory, - vector::embedding::vec_delete_phrase_memory, - vector::embedding::vec_update_phrase_memory, - vector::embedding::vec_search_phrase_memory, - vector::embedding::vec_save_locked_phrases, - vector::embedding::vec_regenerate_all_embeddings, + vector::memory_commands::vec_list_phrase_memory, + vector::memory_commands::vec_get_phrase_memory, + vector::memory_commands::vec_add_phrase_embedding, + vector::memory_commands::vec_set_phrase_tags, + vector::memory_commands::vec_delete_phrase_memory, + vector::memory_commands::vec_update_phrase_memory, + vector::memory_commands::vec_search_phrase_memory, + vector::memory_commands::vec_save_locked_phrases, + vector::memory_commands::vec_count_project_phrase_relabels, + vector::memory_commands::vec_relabel_project_phrases, + vector::memory_commands::vec_regenerate_all_embeddings, deepl::commands::run_deepl_stage, + deepl::commands::preview_deepl_stage, deepl::commands::get_deepl_languages, deepl::commands::list_deepl_glossaries, deepl::commands::create_deepl_glossary, diff --git a/src-tauri/src/llm/composition.rs b/src-tauri/src/llm/composition.rs new file mode 100644 index 00000000..9e09ca33 --- /dev/null +++ b/src-tauri/src/llm/composition.rs @@ -0,0 +1,113 @@ +//! Named pieces of a prompt. Every LLM request is composed from these parts and +//! the preview shows the same parts, so what the user inspects is exactly what +//! is sent: joining the parts of a block reproduces the block byte for byte. + +use crate::llm::types::{PromptBlock, StructuredPrompt}; + +/// One named piece of a prompt. `text` keeps its own separators, so blocks are +/// plain concatenations of their parts. +#[derive(Debug, Clone)] +pub(crate) struct PromptPart { + pub id: &'static str, + /// The editable system text this part is made from, if any. + pub text_id: Option<&'static str>, + pub text: String, +} + +#[derive(Debug, Clone)] +pub(crate) struct ComposedBlock { + pub cacheable: bool, + pub parts: Vec, +} + +#[derive(Debug, Clone)] +pub(crate) struct ComposedPrompt { + pub system: Vec, + pub user: Vec, +} + +/// A part as the preview receives it: which message it belongs to and whether +/// that system block is cached by providers. +#[derive(Debug, Clone, serde::Serialize)] +#[serde(rename_all = "camelCase")] +pub struct PreviewPart { + pub id: &'static str, + pub text_id: Option<&'static str>, + pub message: &'static str, + pub cacheable: bool, + pub text: String, +} + +/// Collects the parts of one block or message, skipping empty ones: an absent +/// optional piece (no glossary, no examples) simply contributes nothing. +#[derive(Default)] +pub(crate) struct Parts(Vec); + +impl Parts { + pub fn push(self, id: &'static str, text: impl Into) -> Self { + self.push_from(id, None, text) + } + + /// A part made from a system text; `text` already carries its separator. + pub fn push_from( + mut self, + id: &'static str, + text_id: Option<&'static str>, + text: impl Into, + ) -> Self { + let text = text.into(); + if !text.is_empty() { + self.0.push(PromptPart { id, text_id, text }); + } + self + } + + pub fn block(self, cacheable: bool) -> ComposedBlock { + ComposedBlock { + cacheable, + parts: self.0, + } + } + + pub fn into_vec(self) -> Vec { + self.0 + } +} + +fn join(parts: &[PromptPart]) -> String { + parts.iter().map(|part| part.text.as_str()).collect() +} + +impl ComposedPrompt { + pub fn into_structured(self) -> StructuredPrompt { + let system = self + .system + .iter() + .map(|block| PromptBlock { + text: join(&block.parts), + cacheable: block.cacheable, + }) + .collect(); + StructuredPrompt::new(system, join(&self.user)) + } + + pub fn preview_parts(&self) -> Vec { + let system = self.system.iter().flat_map(|block| { + block.parts.iter().map(|part| PreviewPart { + id: part.id, + text_id: part.text_id, + message: "system", + cacheable: block.cacheable, + text: part.text.clone(), + }) + }); + let user = self.user.iter().map(|part| PreviewPart { + id: part.id, + text_id: part.text_id, + message: "user", + cacheable: false, + text: part.text.clone(), + }); + system.chain(user).collect() + } +} diff --git a/src-tauri/src/llm/legacy_prompts_test.rs b/src-tauri/src/llm/legacy_prompts_test.rs new file mode 100644 index 00000000..b351f1c4 --- /dev/null +++ b/src-tauri/src/llm/legacy_prompts_test.rs @@ -0,0 +1,431 @@ +//! Copia della composizione dei prompt prima della scomposizione in pezzi +//! (ottobre 2026): serve solo alle prove di equivalenza, che verificano che i +//! testi inviati non siano cambiati di un carattere. +#![allow(dead_code)] + +use crate::llm::types::{ + CoherenceChunkInput, PipelineConfig, PromptBlock, StageConfig, StructuredPrompt, +}; + +fn format_glossary_table(glossary: &[crate::llm::types::GlossaryEntry]) -> String { + super::prompts::format_glossary_table_for_tests(glossary) +} + +fn format_few_shot_block(examples: &[crate::llm::types::FewShotExample]) -> String { + super::prompts::format_few_shot_block_for_tests(examples) +} + +fn work_brief_block(config: &PipelineConfig) -> String { + config + .work_brief + .as_deref() + .map(str::trim) + .filter(|s| !s.is_empty()) + .map(|brief| format!("\n\nTranslation context:\n{brief}")) + .unwrap_or_default() +} + +pub(crate) fn build_stage_prompts( + text: &str, + stage: &StageConfig, + config: &PipelineConfig, + previous_result: Option<&str>, + audit_context: Option<&str>, +) -> StructuredPrompt { + if stage.role.as_deref() == Some("format") { + return build_format_stage_prompts(text, stage); + } + + let glossary_table = format_glossary_table(&config.glossary); + + let markdown_rules = if config.markdown_aware.unwrap_or(false) { + "\n\nMarkdown Preservation Rules:\n\ + - Preserve every Markdown marker exactly as needed (*, **, _, [], (), headings, lists, block quotes, footnotes)\n\ + - Do not remove, reformat, or invent Markdown structure\n\ + - Translate only the human-language content while keeping Markdown syntax valid" + } else { + "" + }; + + let glossary_rules = if glossary_table.is_empty() { + "Glossary Constraints:\n- No glossary entries were provided.".to_string() + } else { + format!( + "Glossary Constraints:\n\ + - Treat every glossary entry as mandatory terminology, not as a suggestion\n\ + - When a source glossary term appears, use the required target term exactly unless the notes explicitly justify a variant\n\ + - Preserve case, product names, abbreviations, and domain terminology consistently across the whole translation\n\ + - Do not omit glossary terms, paraphrase them away, or replace them with near-synonyms\n\ + - If a glossary term appears inside Markdown, links, or footnotes, still apply the glossary while preserving the surrounding syntax\n\ + - Glossary:\n{}", + glossary_table, + ) + }; + + let opener = "You are an expert translator and linguist. Follow the translation context and the instructions for the current stage."; + + let few_shot_block = format_few_shot_block(&config.few_shot_examples); + let work_context = work_brief_block(config); + + // Block 1 (cacheable): static project-level context — role, work brief, constraints, glossary, + // few-shot examples. Identical for every chunk in the run, so caches across the whole + // document. Few-shot examples are folded into this same block (not a separate one) so + // they consume no extra Anthropic cache breakpoint. + let static_block = format!( + "{opener}{work_context}\n\n\ + Structural Preservation Rules:\n\ + - Preserve paragraph boundaries and line breaks unless the source is clearly malformed\n\ + - Do not collapse repeated spaces, tabs, list structure, or footnote placement when they carry formatting meaning\n\n\ + {glossary_rules}{markdown_rules}{few_shot_block}", + ); + + let mut system = vec![PromptBlock { + text: static_block, + cacheable: true, + }]; + + // Blob context (cacheable) comes BEFORE stage instructions so all stable content + // forms a contiguous prefix: [static + blob]. This lets every provider cache the + // longest common prefix — Anthropic via a single breakpoint here, OpenAI/DeepSeek/ + // Gemini via automatic prefix caching — giving cache hits across all stages within + // the same blob, not only within a single stage. + if let Some(blob) = config.blob_context.as_deref().filter(|s| !s.is_empty()) { + system.push(PromptBlock { + text: format!( + "[Reference document block - context only]\n\ + This block may include the current chunk. Use it for terminology, continuity, names, pronouns, formatting, and narrative context.\n\ + Do not translate this block as a whole. Translate only the current chunk identified in the user message.\n\ + {blob}\n\ + [End reference document block]" + ), + cacheable: true, + }); + } + + // Stage-specific instructions come last: they vary per stage but are smaller than + // the static+blob prefix, so non-caching them costs less than before. + let glossary_reminder = ""; + let output_contract = if stage.role.as_deref() == Some("refine") { + "Output the complete refined translation in full. Do not summarize, abbreviate, or output only the changed portions — rewrite the entire chunk from start to finish." + } else { + "Output only the translated text." + }; + system.push(PromptBlock { + text: format!( + "Core Instructions:\n{}{}\n\n{}", + stage.prompt, glossary_reminder, output_contract + ), + cacheable: false, + }); + + let current_chunk_line = config + .blob_current_chunk_id + .as_deref() + .filter(|s| !s.is_empty()) + .map(|id| format!("Current chunk id: {id}\n\n")) + .unwrap_or_default(); + + let user = if stage.role.as_deref() == Some("refine") { + let base = format!( + "{current_chunk_line}Original text for the current chunk:\n{text}\n\n\ + Previous Iteration for the current chunk:\n{}\n\n\ + Refine only the current chunk according to your instructions. Output the complete refined translation in full — every sentence, from start to finish. Do not abbreviate or output only the changed portions.", + previous_result.unwrap_or_default() + ); + if let Some(ctx) = audit_context.filter(|s| !s.trim().is_empty()) { + format!("{base}\n\n---\nPrevious audit findings to address:\n{ctx}\n---") + } else { + base + } + } else { + format!( + "{current_chunk_line}Text to translate from the current chunk:\n{text}\n\n\ + Translate only the current chunk. Output only its translation." + ) + }; + + StructuredPrompt::new(system, user) +} + +fn build_format_stage_prompts(text: &str, stage: &StageConfig) -> StructuredPrompt { + let system = vec![ + PromptBlock { + text: "\ +You are a deterministic text post-processor for already translated text.\n\ +The input is already translated. Do not translate, retranslate, paraphrase, improve style, correct meaning, expand, shorten, or alter wording except where a minimal formatting repair requires it.\n\ +Allowed changes: repair broken Markdown or footnote syntax, and restore clearly corrupted spacing or line breaks.\n\ +Do not add new emphasis, code, link, heading, list, quote, table, or other markup. Change existing Markdown markers only when necessary to restore valid syntax.\n\ +Return the complete text. If no change is needed, return the input exactly.\n\ +Do not return explanations, comments, JSON, diffs, or 'no changes'." + .to_string(), + cacheable: true, + }, + PromptBlock { + text: format!("Core Formatting Instructions:\n{}\n\nOutput only the formatted text.", stage.prompt), + cacheable: false, + }, + ]; + + let user = format!( + "Text to format from the current chunk:\n{text}\n\n\ + Apply only the formatting instructions. Output only the complete formatted text." + ); + + StructuredPrompt::new(system, user) +} + +pub(crate) fn build_judge_prompts( + source_text: &str, + translation: &str, + config: &PipelineConfig, +) -> StructuredPrompt { + let glossary_table = format_glossary_table(&config.glossary); + let opener = + "You are a translation quality judge. Evaluate the translation against the translation context."; + let work_context = work_brief_block(config); + let ui_lang = config + .ui_language + .as_deref() + .filter(|s| !s.is_empty()) + .unwrap_or("English"); + + let glossary_section = if glossary_table.is_empty() { + String::new() + } else { + format!("Glossary to adhere to:\n{glossary_table}\n\n") + }; + + let markdown_rules = if config.markdown_aware.unwrap_or(false) { + "When Markdown is present, verify that the translation preserves markers, footnotes, \ + inline emphasis, and block structure exactly enough to remain valid Markdown.\n\n" + } else { + "" + }; + + // Block 1 (cacheable): static judge context — role, instructions, glossary, format spec. + // The source text and translation are in the user turn so this block is constant for the + // whole project run, enabling near-100% cache hit rate across all chunk judge calls. + let system_block = format!( + "{opener}{work_context}\n\n\ + Specific Audit Instructions:\n{instructions}\n\n\ + {glossary_section}\ + {markdown_rules}\ + Scanning protocol: go through the translation sentence by sentence, checking every \ + sentence against the source for accuracy, every glossary term for adherence, grammar \ + for correctness, and fluency throughout. Complete the full scan before building the issues list. \ + Report EVERY issue you find and EVERY occurrence separately — do not merge, suppress, or \ + limit repeated issues.\n\n\ + You MUST respond with a valid JSON object containing:\n\ + - checkedSentenceIndices: array of 1-based source sentence numbers you verified, in scan order \ + (e.g. [1, 2, 3] for a 3-sentence source) — indices only, never the sentence text itself\n\ + - rating: one of 'critical', 'poor', 'fair', 'good', 'excellent' \ + (semantic translation quality: critical=unusable, poor=weak, fair=usable with revision, \ + good=solid, excellent=publication-ready)\n\ + - issues: array of objects with these fields:\n\ + - type: 'glossary'|'fluency'|'accuracy'|'grammar'\n\ + - severity: 'low'|'medium'|'high'\n\ + - description: string — explanation of the issue in {ui_lang}\n\ + - suggestedFix: string — how to correct it in {ui_lang}\n\ + - phrase: string or null — the exact verbatim substring of the WRONG or problematic text \ + as it appears in the TARGET translation (character-for-character copy from the target text)\n\ + - sourcePhrase: string or null — the exact verbatim substring from the SOURCE text \ + that corresponds to this issue\n\ + - confidence: number or null — your confidence this is a real issue (0.0–1.0)\n\ + Write description and suggestedFix in {ui_lang}. \ + Keep rating and type values as the English literals above.", + instructions = config.judge_prompt, + ); + + let user = format!("Source: {source_text}\nTarget: {translation}\n\nPerform the audit now and return the JSON report."); + + StructuredPrompt { + system: vec![PromptBlock { + text: system_block, + cacheable: true, + }], + user, + images: Vec::new(), + } +} + +pub(crate) fn build_coherence_prompts( + input: &CoherenceChunkInput, + config: &PipelineConfig, +) -> StructuredPrompt { + let glossary_table = format_glossary_table(&config.glossary); + let opener = + "You are a translation coherence auditor. Evaluate consistency against the translation context."; + let work_context = work_brief_block(config); + let ui_lang = config + .ui_language + .as_deref() + .filter(|s| !s.is_empty()) + .unwrap_or("English"); + + let default_instructions = "Evaluate ONLY:\n\ + 1. Terminology consistency — key terms translated differently than in adjacent segments\n\ + 2. Narrative continuity — abrupt breaks in flow at segment boundaries\n\ + 3. Glossary adherence — glossary terms used inconsistently with context\n\ + Do NOT re-evaluate standalone translation quality.\n\ + Be exhaustive: scan ALL dimensions completely before responding. Do not stop after finding \ + the first issue of each type. Only return an empty issues array if you are fully confident \ + — after deliberate review of every dimension — that no problems exist."; + + let instructions = config + .coherence_prompt + .as_deref() + .filter(|s| !s.trim().is_empty()) + .unwrap_or(default_instructions); + + let glossary_section = if glossary_table.is_empty() { + String::new() + } else { + format!("Glossary:\n{glossary_table}\n\n") + }; + + // Block 1 (cacheable): static coherence context — role, instructions, glossary, format spec. + // Constant for the whole project run. + let system_block = format!( + "{opener}{work_context}\n\ + Your task: identify cross-segment inconsistencies between a translated segment and its surrounding context.\n\ + {instructions}\n\ + {glossary_section}\ + Write description and suggestedFix values in {ui_lang}.\n\ + Respond with valid JSON only:\n\ + {{\"issues\": [{{\"type\": \"consistency\"|\"glossary\", \ + \"severity\": \"low\"|\"medium\"|\"high\", \ + \"description\": \"string\", \ + \"suggestedFix\": \"string\", \ + \"phrase\": \"exact verbatim substring of the WRONG text as it appears in the target translation, not the source term nor the correction; first occurrence only\"}}]}}", + ); + + // Block 2 (cacheable): reference document block. Identical for every chunk in the same + // blob, so it's a second cache breakpoint — placed in system, not the user turn, so + // providers actually cache it instead of rebilling it at full price on every chunk. + let context_block = input + .blob_context + .as_deref() + .filter(|s| !s.is_empty()) + .map(|ctx| format!( + "[Reference translated document block - context only]\n\ + This block may include the current chunk. Use it to compare terminology and continuity across the document block.\n\ + The current chunk to audit is identified below.\n\ + {ctx}\n\ + [End reference translated document block]" + )); + + let current_chunk_line = input + .current_chunk_id + .as_deref() + .filter(|s| !s.is_empty()) + .map(|id| format!("Current chunk id: {id}\n\n")) + .unwrap_or_default(); + + let user = format!( + "{current_chunk_line}[Current segment]\nOriginal: {original}\nTranslation: {translation}\n\ + [End of current segment]\n\n\ + Identify cross-segment coherence issues and return the JSON. If no issues, return {{\"issues\": []}}.", + original = input.original, + translation = input.translation, + ); + + let mut system = vec![PromptBlock { + text: system_block, + cacheable: true, + }]; + if let Some(ctx) = context_block { + system.push(PromptBlock { + text: ctx, + cacheable: true, + }); + } + + StructuredPrompt::new(system, user) +} + +mod equivalence { + use crate::llm::prompts; + use crate::llm::types::{ + CoherenceChunkInput, FewShotExample, GlossaryEntry, PipelineConfig, StageConfig, + StructuredPrompt, + }; + + fn same(new: StructuredPrompt, old: StructuredPrompt) { + assert_eq!(new.system.len(), old.system.len()); + for (a, b) in new.system.iter().zip(old.system.iter()) { + assert_eq!(a.text, b.text); + assert_eq!(a.cacheable, b.cacheable); + } + assert_eq!(new.user, old.user); + } + + fn configs() -> Vec { + let bare = PipelineConfig::default(); + let full = PipelineConfig { + work_brief: Some("Venetian, 17th century, into Italian.".into()), + glossary: vec![GlossaryEntry { + term: "arma".into(), + translation: "arms".into(), + notes: Some("heraldic".into()), + }], + few_shot_examples: vec![FewShotExample { + source_text: "uno".into(), + target_text: "one".into(), + label: None, + }], + markdown_aware: Some(true), + coherence_prompt: Some("Check names only.".into()), + ui_language: Some("Italian".into()), + blob_context: Some("x".into()), + blob_current_chunk_id: Some("c1".into()), + judge_prompt: "Be strict.".into(), + ..PipelineConfig::default() + }; + vec![bare, full] + } + + fn stage(role: &str) -> StageConfig { + StageConfig { + id: "s".into(), + name: "S".into(), + role: Some(role.into()), + prompt: "Do it.".into(), + model: "m".into(), + provider: "openai".into(), + enabled: true, + provider_options: None, + custom_provider_id: None, + } + } + + #[test] + fn composed_prompts_are_byte_identical_to_the_previous_builders() { + for config in configs() { + for role in ["translation", "refine", "format"] { + let s = stage(role); + for (prev, audit) in [(None, None), (Some("prev"), Some("fix X"))] { + same( + prompts::build_stage_prompts("Hello", &s, &config, prev, audit), + super::build_stage_prompts("Hello", &s, &config, prev, audit), + ); + } + } + same( + prompts::build_judge_prompts("src", "tgt", &config), + super::build_judge_prompts("src", "tgt", &config), + ); + for blob in [None, Some("".to_string())] { + let input = CoherenceChunkInput { + original: "o".into(), + translation: "t".into(), + blob_context: blob, + current_chunk_id: Some("c1".into()), + }; + same( + prompts::build_coherence_prompts(&input, &config), + super::build_coherence_prompts(&input, &config), + ); + } + } + } +} diff --git a/src-tauri/src/llm/mod.rs b/src-tauri/src/llm/mod.rs index cb1a3f2e..e5f5cf8d 100644 --- a/src-tauri/src/llm/mod.rs +++ b/src-tauri/src/llm/mod.rs @@ -1,6 +1,8 @@ pub mod blobs; +pub(crate) mod composition; pub mod custom_profiles; pub mod pipeline; +pub mod prompt_texts; pub mod prompts; pub mod provider; pub mod providers; @@ -9,5 +11,7 @@ pub mod types; pub use stream::StreamRegistry; +#[cfg(test)] +mod legacy_prompts_test; #[cfg(test)] mod tests; diff --git a/src-tauri/src/llm/pipeline.rs b/src-tauri/src/llm/pipeline.rs index 2bce6321..bbb7fc4f 100644 --- a/src-tauri/src/llm/pipeline.rs +++ b/src-tauri/src/llm/pipeline.rs @@ -3,11 +3,13 @@ use tauri::{AppHandle, Emitter, State}; use crate::keystore::get_api_key; use crate::llm::blobs::{compute_blob_assignments, BlobAssignment, ChunkForBlob}; +use crate::llm::composition::{ComposedPrompt, PreviewPart}; use crate::llm::custom_profiles; use crate::llm::prompts::{ - build_coherence_prompts, build_judge_prompts, build_stage_prompts, escape_prompt_markers, - minimal_pipeline_config, parse_judge_rating, sanitize_llm_json_output, - REFINE_AUDIT_SYSTEM_PROMPT, REFINE_STAGE_SYSTEM_PROMPT, + build_coherence_prompts, build_judge_prompts, build_stage_prompts, compose_coherence_prompts, + compose_judge_prompts, compose_stage_prompts, escape_prompt_markers, minimal_pipeline_config, + parse_judge_rating, sanitize_llm_json_output, REFINE_AUDIT_SYSTEM_PROMPT, + REFINE_STAGE_SYSTEM_PROMPT, }; use crate::llm::provider::{LlmProvider, LlmRequest}; use crate::llm::providers::{get_provider, with_retry_after}; @@ -253,6 +255,20 @@ pub async fn run_stage( pub struct StagePromptPreview { pub system_prompt: String, pub user_prompt: String, + /// The same request as named parts, in sending order. + pub parts: Vec, +} + +impl From for StagePromptPreview { + fn from(composed: ComposedPrompt) -> Self { + let parts = composed.preview_parts(); + let structured = composed.into_structured(); + Self { + system_prompt: structured.flatten_system(), + user_prompt: structured.user, + parts, + } + } } /// Builds the exact prompt a stage would send, without contacting any provider. @@ -265,18 +281,32 @@ pub fn preview_stage_prompt( previous_result: Option, audit_context: Option, ) -> StagePromptPreview { - let structured = build_stage_prompts( + compose_stage_prompts( &text, &stage, &config, previous_result.as_deref(), audit_context.as_deref(), - ); - let system_prompt = structured.flatten_system(); - StagePromptPreview { - system_prompt, - user_prompt: structured.user, - } + ) + .into() +} + +/// Review previews share prompt builders with execution and never contact a provider. +#[tauri::command] +pub fn preview_judge_prompt( + source_text: String, + translation: String, + config: PipelineConfig, +) -> StagePromptPreview { + compose_judge_prompts(&source_text, &translation, &config).into() +} + +#[tauri::command] +pub fn preview_coherence_prompt( + input: CoherenceChunkInput, + config: PipelineConfig, +) -> StagePromptPreview { + compose_coherence_prompts(&input, &config).into() } #[tauri::command] @@ -493,8 +523,12 @@ pub async fn refine_prompt( prov.preflight(&model).await?; let api_key = get_api_key(&app, &provider)?; let client = prov.http_client()?; - let system_text = if context == "audit" { + let system_text = if context == "brief" { + "Rewrite the shared translation context clearly and concisely. Preserve all stated languages, historical varieties, goals, audience and register. Do not invent requirements or add stage-specific commands, evaluation criteria or output formats. Return only the rewritten work brief." + } else if context == "audit" { REFINE_AUDIT_SYSTEM_PROMPT + } else if context == "system" { + "Rewrite this system instruction of a translation pipeline to be clearer and more effective for modern LLMs. Preserve its purpose and every placeholder written as {{NAME}} exactly, in place. Do not add new requirements. Return only the rewritten text." } else { REFINE_STAGE_SYSTEM_PROMPT }; diff --git a/src-tauri/src/llm/prompt_texts.rs b/src-tauri/src/llm/prompt_texts.rs new file mode 100644 index 00000000..98df0936 --- /dev/null +++ b/src-tauri/src/llm/prompt_texts.rs @@ -0,0 +1,387 @@ +//! System texts of the prompts: every piece of wording the program adds around +//! the user's own prompts. Each has a stable id, a default and the placeholders +//! it must keep. A pipeline may override any of them; an override that drops a +//! required placeholder is ignored, so a chunk can never be sent without its text. + +use std::collections::HashMap; + +use crate::llm::types::PipelineConfig; + +pub(crate) struct SystemText { + pub id: &'static str, + pub default: &'static str, + /// Placeholders the text must contain; without them the default is used. + pub required: &'static [&'static str], +} + +const fn text( + id: &'static str, + default: &'static str, + required: &'static [&'static str], +) -> SystemText { + SystemText { + id, + default, + required, + } +} + +pub(crate) const SYSTEM_TEXTS: &[SystemText] = &[ + // Shared by every LLM phase that receives them. + text( + "context-frame", + "Translation context:\n{{TRANSLATION_CONTEXT}}", + &["TRANSLATION_CONTEXT"], + ), + text("chunk-id", "Current chunk id: {{CHUNK_ID}}", &["CHUNK_ID"]), + // Translation and Refine. + text( + "translation.role", + "You are an expert translator and linguist. Follow the translation context and the instructions for the current stage.", + &[], + ), + text( + "translation.structural-rules", + "Structural Preservation Rules:\n\ + - Preserve paragraph boundaries and line breaks unless the source is clearly malformed\n\ + - Do not collapse repeated spaces, tabs, list structure, or footnote placement when they carry formatting meaning", + &[], + ), + text( + "translation.glossary-rules", + "Glossary Constraints:\n\ + - Treat every glossary entry as mandatory terminology, not as a suggestion\n\ + - When a source glossary term appears, use the required target term exactly unless the notes explicitly justify a variant\n\ + - Preserve case, product names, abbreviations, and domain terminology consistently across the whole translation\n\ + - Do not omit glossary terms, paraphrase them away, or replace them with near-synonyms\n\ + - If a glossary term appears inside Markdown, links, or footnotes, still apply the glossary while preserving the surrounding syntax\n\ + - Glossary:\n{{GLOSSARY_TABLE}}", + &["GLOSSARY_TABLE"], + ), + text( + "translation.glossary-empty", + "Glossary Constraints:\n- No glossary entries were provided.", + &[], + ), + text( + "translation.markdown-rules", + "Markdown Preservation Rules:\n\ + - Preserve every Markdown marker exactly as needed (*, **, _, [], (), headings, lists, block quotes, footnotes)\n\ + - Do not remove, reformat, or invent Markdown structure\n\ + - Translate only the human-language content while keeping Markdown syntax valid", + &[], + ), + text( + "translation.examples", + "Example Translations (match this style, register, and tone):\n{{EXAMPLES}}", + &["EXAMPLES"], + ), + text( + "translation.neighbours", + "[Reference document block - context only]\n\ + This block may include the current chunk. Use it for terminology, continuity, names, pronouns, formatting, and narrative context.\n\ + Do not translate this block as a whole. Translate only the current chunk identified in the user message.\n\ + {{NEIGHBOUR_CHUNKS}}\n\ + [End reference document block]", + &["NEIGHBOUR_CHUNKS"], + ), + text( + "translation.stage-frame", + "Core Instructions:\n{{STAGE_PROMPT}}", + &["STAGE_PROMPT"], + ), + text( + "translation.output-contract", + "Output only the translated text.", + &[], + ), + text( + "translation.user-message", + "Text to translate from the current chunk:\n{{TEXT}}\n\nTranslate only the current chunk. Output only its translation.", + &["TEXT"], + ), + text( + "refine.output-contract", + "Output the complete refined translation in full. Do not summarize, abbreviate, or output only the changed portions — rewrite the entire chunk from start to finish.", + &[], + ), + text( + "refine.user-message", + "Original text for the current chunk:\n{{TEXT}}\n\n\ + Previous Iteration for the current chunk:\n{{PREVIOUS_RESULT}}\n\n\ + Refine only the current chunk according to your instructions. Output the complete refined translation in full — every sentence, from start to finish. Do not abbreviate or output only the changed portions.", + &["TEXT", "PREVIOUS_RESULT"], + ), + text( + "refine.audit-findings", + "---\nPrevious audit findings to address:\n{{AUDIT_FINDINGS}}\n---", + &["AUDIT_FINDINGS"], + ), + // Format. + text( + "format.role", + "\ +You are a deterministic text post-processor for already translated text.\n\ +The input is already translated. Do not translate, retranslate, paraphrase, improve style, correct meaning, expand, shorten, or alter wording except where a minimal formatting repair requires it.\n\ +Allowed changes: repair broken Markdown or footnote syntax, and restore clearly corrupted spacing or line breaks.\n\ +Do not add new emphasis, code, link, heading, list, quote, table, or other markup. Change existing Markdown markers only when necessary to restore valid syntax.\n\ +Return the complete text. If no change is needed, return the input exactly.\n\ +Do not return explanations, comments, JSON, diffs, or 'no changes'.", + &[], + ), + text( + "format.stage-frame", + "Core Formatting Instructions:\n{{STAGE_PROMPT}}", + &["STAGE_PROMPT"], + ), + text( + "format.output-contract", + "Output only the formatted text.", + &[], + ), + text( + "format.user-message", + "Text to format from the current chunk:\n{{TEXT}}\n\nApply only the formatting instructions. Output only the complete formatted text.", + &["TEXT"], + ), + // Audit. + text( + "audit.role", + "You are a translation quality judge. Evaluate the translation against the translation context.", + &[], + ), + text( + "audit.stage-frame", + "Specific Audit Instructions:\n{{STAGE_PROMPT}}", + &["STAGE_PROMPT"], + ), + text( + "audit.glossary", + "Glossary to adhere to:\n{{GLOSSARY_TABLE}}", + &["GLOSSARY_TABLE"], + ), + text( + "audit.markdown-rules", + "When Markdown is present, verify that the translation preserves markers, footnotes, \ + inline emphasis, and block structure exactly enough to remain valid Markdown.", + &[], + ), + text( + "audit.review-method", + "Scanning protocol: go through the translation sentence by sentence, checking every \ + sentence against the source for accuracy, every glossary term for adherence, grammar \ + for correctness, and fluency throughout. Complete the full scan before building the issues list. \ + Report EVERY issue you find and EVERY occurrence separately — do not merge, suppress, or \ + limit repeated issues.", + &[], + ), + text( + "audit.response-format", + "You MUST respond with a valid JSON object containing:\n\ + - checkedSentenceIndices: array of 1-based source sentence numbers you verified, in scan order \ + (e.g. [1, 2, 3] for a 3-sentence source) — indices only, never the sentence text itself\n\ + - rating: one of 'critical', 'poor', 'fair', 'good', 'excellent' \ + (semantic translation quality: critical=unusable, poor=weak, fair=usable with revision, \ + good=solid, excellent=publication-ready)\n\ + - issues: array of objects with these fields:\n\ + - type: 'glossary'|'fluency'|'accuracy'|'grammar'\n\ + - severity: 'low'|'medium'|'high'\n\ + - description: string — explanation of the issue in {{UI_LANGUAGE}}\n\ + - suggestedFix: string — how to correct it in {{UI_LANGUAGE}}\n\ + - phrase: string or null — the exact verbatim substring of the WRONG or problematic text \ + as it appears in the TARGET translation (character-for-character copy from the target text)\n\ + - sourcePhrase: string or null — the exact verbatim substring from the SOURCE text \ + that corresponds to this issue\n\ + - confidence: number or null — your confidence this is a real issue (0.0–1.0)\n\ + Write description and suggestedFix in {{UI_LANGUAGE}}. \ + Keep rating and type values as the English literals above.", + &[], + ), + text( + "audit.user-message", + "Source: {{TEXT}}\nTarget: {{TRANSLATION}}\n\nPerform the audit now and return the JSON report.", + &["TEXT", "TRANSLATION"], + ), + // Coherence. + text( + "coherence.role", + "You are a translation coherence auditor. Evaluate consistency against the translation context.", + &[], + ), + text( + "coherence.review-method", + "Your task: identify cross-segment inconsistencies between a translated segment and its surrounding context.", + &[], + ), + text( + "coherence.glossary", + "Glossary:\n{{GLOSSARY_TABLE}}", + &["GLOSSARY_TABLE"], + ), + text( + "coherence.response-format", + "Write description and suggestedFix values in {{UI_LANGUAGE}}.\n\ + Respond with valid JSON only:\n\ + {\"issues\": [{\"type\": \"consistency\"|\"glossary\", \ + \"severity\": \"low\"|\"medium\"|\"high\", \ + \"description\": \"string\", \ + \"suggestedFix\": \"string\", \ + \"phrase\": \"exact verbatim substring of the WRONG text as it appears in the target translation, not the source term nor the correction; first occurrence only\"}]}", + &[], + ), + text( + "coherence.neighbours", + "[Reference translated document block - context only]\n\ + This block may include the current chunk. Use it to compare terminology and continuity across the document block.\n\ + The current chunk to audit is identified below.\n\ + {{NEIGHBOUR_CHUNKS}}\n\ + [End reference translated document block]", + &["NEIGHBOUR_CHUNKS"], + ), + text( + "coherence.user-message", + "[Current segment]\nOriginal: {{TEXT}}\nTranslation: {{TRANSLATION}}\n[End of current segment]\n\n\ + Identify cross-segment coherence issues and return the JSON. If no issues, return {\"issues\": []}.", + &["TEXT", "TRANSLATION"], + ), +]; + +fn spec(id: &str) -> Option<&'static SystemText> { + SYSTEM_TEXTS.iter().find(|entry| entry.id == id) +} + +fn has_placeholder(template: &str, name: &str) -> bool { + template.contains(&format!("{{{{{name}}}}}")) +} + +/// The text a pipeline uses: its override when present and complete, else the default. +pub(crate) fn system_text<'a>(config: &'a PipelineConfig, id: &str) -> &'a str { + // Ids are constants of this module; an unknown one is a programming error. + let Some(entry) = spec(id) else { + debug_assert!(false, "unknown system text id: {id}"); + return ""; + }; + config + .prompt_composition + .texts + .get(id) + .map(String::as_str) + .filter(|custom| !custom.trim().is_empty()) + .filter(|custom| { + entry + .required + .iter() + .all(|name| has_placeholder(custom, name)) + }) + .unwrap_or(entry.default) +} + +/// Replaces `{{NAME}}` placeholders in one pass: inserted values are never +/// scanned again, so a chunk that happens to contain `{{TEXT}}` stays as it is. +pub(crate) fn fill(template: &str, values: &[(&str, &str)]) -> String { + let lookup: HashMap<&str, &str> = values.iter().copied().collect(); + let mut out = String::with_capacity(template.len()); + let mut rest = template; + while let Some(start) = rest.find("{{") { + out.push_str(&rest[..start]); + let after = &rest[start + 2..]; + match after.find("}}") { + Some(end) => match lookup.get(&after[..end]) { + Some(value) => { + out.push_str(value); + rest = &after[end + 2..]; + } + None => { + out.push_str("{{"); + rest = after; + } + }, + None => { + out.push_str(&rest[start..]); + rest = ""; + } + } + } + out.push_str(rest); + out +} + +/// The system text of a pipeline, with its placeholders filled. +pub(crate) fn render(config: &PipelineConfig, id: &str, values: &[(&str, &str)]) -> String { + fill(system_text(config, id), values) +} + +/// Defaults and placeholders for the editor in the prompt preview. +#[derive(Debug, Clone, serde::Serialize)] +#[serde(rename_all = "camelCase")] +pub struct SystemTextInfo { + pub id: &'static str, + pub default_text: &'static str, + pub required: Vec<&'static str>, +} + +#[tauri::command] +pub fn prompt_system_texts() -> Vec { + SYSTEM_TEXTS + .iter() + .map(|entry| SystemTextInfo { + id: entry.id, + default_text: entry.default, + required: entry.required.to_vec(), + }) + .collect() +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::llm::types::PromptComposition; + + #[test] + fn fill_replaces_known_placeholders_once() { + assert_eq!( + fill("A {{TEXT}} B {{OTHER}}", &[("TEXT", "x {{OTHER}}")]), + "A x {{OTHER}} B {{OTHER}}" + ); + } + + #[test] + fn override_without_required_placeholder_falls_back_to_default() { + let mut config = PipelineConfig { + prompt_composition: PromptComposition { + texts: [( + "translation.user-message".to_string(), + "No placeholder".to_string(), + )] + .into_iter() + .collect(), + disabled: vec![], + }, + ..PipelineConfig::default() + }; + assert_eq!( + system_text(&config, "translation.user-message"), + spec("translation.user-message") + .map(|entry| entry.default) + .unwrap_or_default() + ); + config + .prompt_composition + .texts + .insert("translation.role".into(), "Custom role".into()); + assert_eq!(system_text(&config, "translation.role"), "Custom role"); + } + + #[test] + fn every_default_contains_its_required_placeholders() { + for entry in SYSTEM_TEXTS { + for name in entry.required { + assert!( + has_placeholder(entry.default, name), + "{} lacks {name}", + entry.id + ); + } + } + } +} diff --git a/src-tauri/src/llm/prompts.rs b/src-tauri/src/llm/prompts.rs index 949e5775..de891d00 100644 --- a/src-tauri/src/llm/prompts.rs +++ b/src-tauri/src/llm/prompts.rs @@ -1,3 +1,5 @@ +use crate::llm::composition::{ComposedPrompt, Parts}; +use crate::llm::prompt_texts::render; use crate::llm::types::{ CoherenceChunkInput, FewShotExample, ImageAttachment, PipelineConfig, PromptBlock, ProviderRuntimeConfig, StageConfig, StructuredPrompt, @@ -41,40 +43,21 @@ fn format_glossary_table(glossary: &[crate::llm::types::GlossaryEntry]) -> Strin table } -/// Formats hand-picked example translations for the cacheable static block. -/// Returns an empty string when there are none, so the static block is -/// byte-identical to before this feature for pipelines without examples. -fn format_few_shot_block(examples: &[FewShotExample]) -> String { - if examples.is_empty() { - return String::new(); - } - let mut block = - "\n\nExample Translations (match this style, register, and tone):\n".to_string(); - for (i, example) in examples.iter().enumerate() { - block.push_str(&format!( - "\nExample {}:\nSource: {}\nTarget: {}\n", - i + 1, - example.source_text, - example.target_text, - )); - } - block -} - -fn effective_source(config: &PipelineConfig) -> &str { - config - .custom_source_language - .as_deref() - .filter(|s| !s.trim().is_empty()) - .unwrap_or(&config.source_language) -} - -fn effective_target(config: &PipelineConfig) -> &str { - config - .custom_target_language - .as_deref() - .filter(|s| !s.trim().is_empty()) - .unwrap_or(&config.target_language) +/// The hand-picked example translations, as they fill `{{EXAMPLES}}`. Empty +/// when there are none: the examples part is then left out. +fn format_few_shot_list(examples: &[FewShotExample]) -> String { + examples + .iter() + .enumerate() + .map(|(i, example)| { + format!( + "\nExample {}:\nSource: {}\nTarget: {}\n", + i + 1, + example.source_text, + example.target_text, + ) + }) + .collect() } /// Persona, transcription rules and output contract for OCR/HTR (#220). @@ -116,6 +99,30 @@ pub(crate) fn build_ocr_prompt(resolved_prompt: &str, image: ImageAttachment) -> } } +#[cfg(test)] +pub(crate) fn format_glossary_table_for_tests( + glossary: &[crate::llm::types::GlossaryEntry], +) -> String { + format_glossary_table(glossary) +} + +#[cfg(test)] +pub(crate) fn format_few_shot_block_for_tests(examples: &[FewShotExample]) -> String { + let list = format_few_shot_list(examples); + if list.is_empty() { + return list; + } + format!("\n\nExample Translations (match this style, register, and tone):\n{list}") +} + +fn work_brief(config: &PipelineConfig) -> Option<&str> { + config + .work_brief + .as_deref() + .map(str::trim) + .filter(|s| !s.is_empty()) +} + pub(crate) fn build_stage_prompts( text: &str, stage: &StageConfig, @@ -123,159 +130,310 @@ pub(crate) fn build_stage_prompts( previous_result: Option<&str>, audit_context: Option<&str>, ) -> StructuredPrompt { + compose_stage_prompts(text, stage, config, previous_result, audit_context).into_structured() +} + +/// `sep` + the rendered system text, as one part. The separators reproduce the +/// layout the prompts always had; the wording comes from the pipeline. +fn sys(config: &PipelineConfig, sep: &str, id: &'static str, values: &[(&str, &str)]) -> String { + format!("{sep}{}", render(config, id, values)) +} + +/// Whether a switchable part is on for a phase (`phase:part` in the disabled list turns it off). +fn on(config: &PipelineConfig, phase: &str, part: &str) -> bool { + let key = format!("{phase}:{part}"); + !config.prompt_composition.disabled.contains(&key) +} + +/// The text when the part is on, nothing when it is switched off. +fn when_on(config: &PipelineConfig, phase: &str, part: &str, text: String) -> String { + if on(config, phase, part) { + text + } else { + String::new() + } +} + +fn context_part(config: &PipelineConfig) -> String { + work_brief(config) + .map(|brief| { + sys( + config, + "\n\n", + "context-frame", + &[("TRANSLATION_CONTEXT", brief)], + ) + }) + .unwrap_or_default() +} + +/// Translation and refine: [static: role, context, rules, glossary, markdown, +/// examples] → [neighbouring chunks] → [stage instructions]. The order never +/// changes: it is the cacheable prefix every provider relies on. +pub(crate) fn compose_stage_prompts( + text: &str, + stage: &StageConfig, + config: &PipelineConfig, + previous_result: Option<&str>, + audit_context: Option<&str>, +) -> ComposedPrompt { if stage.role.as_deref() == Some("format") { - return build_format_stage_prompts(text, stage); + return compose_format_stage_prompts(text, stage, config); } + let phase = if stage.role.as_deref() == Some("refine") { + "refine" + } else { + "translation" + }; let glossary_table = format_glossary_table(&config.glossary); - - let markdown_rules = if config.markdown_aware.unwrap_or(false) { - "\n\nMarkdown Preservation Rules:\n\ - - Preserve every Markdown marker exactly as needed (*, **, _, [], (), headings, lists, block quotes, footnotes)\n\ - - Do not remove, reformat, or invent Markdown structure\n\ - - Translate only the human-language content while keeping Markdown syntax valid" + let (glossary_id, glossary_text) = if glossary_table.is_empty() { + ( + "translation.glossary-empty", + render(config, "translation.glossary-empty", &[]), + ) } else { - "" + ( + "translation.glossary-rules", + render( + config, + "translation.glossary-rules", + &[("GLOSSARY_TABLE", &glossary_table)], + ), + ) }; - - let glossary_rules = if glossary_table.is_empty() { - "Glossary Constraints:\n- No glossary entries were provided.".to_string() + let markdown = if config.markdown_aware.unwrap_or(false) && on(config, phase, "markdown-rules") + { + sys(config, "\n\n", "translation.markdown-rules", &[]) } else { - format!( - "Glossary Constraints:\n\ - - Treat every glossary entry as mandatory terminology, not as a suggestion\n\ - - When a source glossary term appears, use the required target term exactly unless the notes explicitly justify a variant\n\ - - Preserve case, product names, abbreviations, and domain terminology consistently across the whole translation\n\ - - Do not omit glossary terms, paraphrase them away, or replace them with near-synonyms\n\ - - If a glossary term appears inside Markdown, links, or footnotes, still apply the glossary while preserving the surrounding syntax\n\ - - Glossary:\n{}", - glossary_table, + String::new() + }; + let examples_list = format_few_shot_list(&config.few_shot_examples); + let examples = if examples_list.is_empty() || !on(config, phase, "examples") { + String::new() + } else { + sys( + config, + "\n\n", + "translation.examples", + &[("EXAMPLES", &examples_list)], ) }; - let src = effective_source(config); - let tgt = effective_target(config); - - let default_opener = format!( - "You are an expert translator and linguist specialized in {src} to {tgt} translation.", - ); - let opener = config - .persona - .as_deref() - .filter(|p| !p.trim().is_empty()) - .unwrap_or(&default_opener); - - let few_shot_block = format_few_shot_block(&config.few_shot_examples); - - // Block 1 (cacheable): static project-level context — persona, constraints, glossary, - // few-shot examples. Identical for every chunk in the run, so caches across the whole - // document. Few-shot examples are folded into this same block (not a separate one) so - // they consume no extra Anthropic cache breakpoint. - let static_block = format!( - "{opener}\n\n\ - Structural Preservation Rules:\n\ - - Preserve paragraph boundaries and line breaks unless the source is clearly malformed\n\ - - Do not collapse repeated spaces, tabs, list structure, or footnote placement when they carry formatting meaning\n\n\ - {glossary_rules}{markdown_rules}{few_shot_block}", - ); - - let mut system = vec![PromptBlock { - text: static_block, - cacheable: true, - }]; + let mut system = vec![Parts::default() + .push_from( + "role", + Some("translation.role"), + when_on( + config, + phase, + "role", + render(config, "translation.role", &[]), + ), + ) + .push_from( + "translation-context", + Some("context-frame"), + context_part(config), + ) + .push_from( + "structural-rules", + Some("translation.structural-rules"), + when_on( + config, + phase, + "structural-rules", + sys(config, "\n\n", "translation.structural-rules", &[]), + ), + ) + .push_from( + "glossary-rules", + Some(glossary_id), + when_on( + config, + phase, + "glossary-rules", + format!("\n\n{glossary_text}"), + ), + ) + .push_from( + "markdown-rules", + Some("translation.markdown-rules"), + markdown, + ) + .push_from("examples", Some("translation.examples"), examples) + .block(true)]; - // Blob context (cacheable) comes BEFORE stage instructions so all stable content + // Neighbouring chunks (cacheable) come BEFORE stage instructions so all stable content // forms a contiguous prefix: [static + blob]. This lets every provider cache the // longest common prefix — Anthropic via a single breakpoint here, OpenAI/DeepSeek/ // Gemini via automatic prefix caching — giving cache hits across all stages within // the same blob, not only within a single stage. - if let Some(blob) = config.blob_context.as_deref().filter(|s| !s.is_empty()) { - system.push(PromptBlock { - text: format!( - "[Reference document block - context only]\n\ - This block may include the current chunk. Use it for terminology, continuity, names, pronouns, formatting, and narrative context.\n\ - Do not translate this block as a whole. Translate only the current chunk identified in the user message.\n\ - {blob}\n\ - [End reference document block]" - ), - cacheable: true, - }); + let neighbours_on = on(config, phase, "neighbour-chunks"); + if let Some(blob) = config + .blob_context + .as_deref() + .filter(|s| !s.is_empty() && neighbours_on) + { + system.push( + Parts::default() + .push_from( + "neighbour-chunks", + Some("translation.neighbours"), + render( + config, + "translation.neighbours", + &[("NEIGHBOUR_CHUNKS", blob)], + ), + ) + .block(true), + ); } // Stage-specific instructions come last: they vary per stage but are smaller than - // the static+blob prefix, so non-caching them costs less than before. - let glossary_reminder = if config.glossary.is_empty() { - "" + // the static+blob prefix, so non-caching them costs less. The glossary rules sit + // once in the static block: no reminder is repeated here. + let is_refine = stage.role.as_deref() == Some("refine"); + let contract_id = if is_refine { + "refine.output-contract" } else { - "\n\nGlossary Reminder:\n- Apply the glossary entries specified above when they appear in the source text." + "translation.output-contract" }; - let output_contract = if stage.role.as_deref() == Some("refine") { - "Output the complete refined translation in full. Do not summarize, abbreviate, or output only the changed portions — rewrite the entire chunk from start to finish." + system.push( + Parts::default() + .push_from( + "stage-prompt", + Some("translation.stage-frame"), + render( + config, + "translation.stage-frame", + &[("STAGE_PROMPT", &stage.prompt)], + ), + ) + .push_from( + "output-contract", + Some(contract_id), + when_on( + config, + phase, + "output-contract", + sys(config, "\n\n", contract_id, &[]), + ), + ) + .block(false), + ); + + // The chunk id only points into the neighbouring chunks: it goes with them. + let chunk_id = if neighbours_on { + chunk_id_part(config, config.blob_current_chunk_id.as_deref()) } else { - "Output only the translated text." + String::new() }; - system.push(PromptBlock { - text: format!( - "Core Instructions:\n{}{}\n\n{}", - stage.prompt, glossary_reminder, output_contract - ), - cacheable: false, - }); - - let current_chunk_line = config - .blob_current_chunk_id - .as_deref() - .filter(|s| !s.is_empty()) - .map(|id| format!("Current chunk id: {id}\n\n")) - .unwrap_or_default(); - - let user = if stage.role.as_deref() == Some("refine") { - let base = format!( - "{current_chunk_line}Original text for the current chunk:\n{text}\n\n\ - Previous Iteration for the current chunk:\n{}\n\n\ - Refine only the current chunk according to your instructions. Output the complete refined translation in full — every sentence, from start to finish. Do not abbreviate or output only the changed portions.", - previous_result.unwrap_or_default() - ); - if let Some(ctx) = audit_context.filter(|s| !s.trim().is_empty()) { - format!("{base}\n\n---\nPrevious audit findings to address:\n{ctx}\n---") - } else { - base - } + let user_sep = if chunk_id.is_empty() { "" } else { "\n\n" }; + let user = if is_refine { + Parts::default() + .push_from("chunk-id", Some("chunk-id"), chunk_id) + .push_from( + "user-message", + Some("refine.user-message"), + sys( + config, + user_sep, + "refine.user-message", + &[ + ("TEXT", text), + ("PREVIOUS_RESULT", previous_result.unwrap_or_default()), + ], + ), + ) + .push_from( + "audit-findings", + Some("refine.audit-findings"), + audit_context + .filter(|s| !s.trim().is_empty()) + .map(|ctx| { + sys( + config, + "\n\n", + "refine.audit-findings", + &[("AUDIT_FINDINGS", ctx)], + ) + }) + .unwrap_or_default(), + ) } else { - format!( - "{current_chunk_line}Text to translate from the current chunk:\n{text}\n\n\ - Translate only the current chunk. Output only its translation." - ) + Parts::default() + .push_from("chunk-id", Some("chunk-id"), chunk_id) + .push_from( + "user-message", + Some("translation.user-message"), + sys( + config, + user_sep, + "translation.user-message", + &[("TEXT", text)], + ), + ) }; - StructuredPrompt::new(system, user) + ComposedPrompt { + system, + user: user.into_vec(), + } } -fn build_format_stage_prompts(text: &str, stage: &StageConfig) -> StructuredPrompt { +fn chunk_id_part(config: &PipelineConfig, id: Option<&str>) -> String { + id.filter(|s| !s.is_empty()) + .map(|id| render(config, "chunk-id", &[("CHUNK_ID", id)])) + .unwrap_or_default() +} + +fn compose_format_stage_prompts( + text: &str, + stage: &StageConfig, + config: &PipelineConfig, +) -> ComposedPrompt { let system = vec![ - PromptBlock { - text: "\ -You are a deterministic text post-processor for already translated text.\n\ -The input is already translated. Do not translate, retranslate, paraphrase, improve style, correct meaning, expand, shorten, or alter wording except where a minimal formatting repair requires it.\n\ -Allowed changes: repair broken Markdown or footnote syntax, and restore clearly corrupted spacing or line breaks.\n\ -Do not add new emphasis, code, link, heading, list, quote, table, or other markup. Change existing Markdown markers only when necessary to restore valid syntax.\n\ -Return the complete text. If no change is needed, return the input exactly.\n\ -Do not return explanations, comments, JSON, diffs, or 'no changes'." - .to_string(), - cacheable: true, - }, - PromptBlock { - text: format!("Core Formatting Instructions:\n{}\n\nOutput only the formatted text.", stage.prompt), - cacheable: false, - }, + Parts::default() + .push_from( + "role", + Some("format.role"), + when_on(config, "format", "role", render(config, "format.role", &[])), + ) + .block(true), + Parts::default() + .push_from( + "stage-prompt", + Some("format.stage-frame"), + render( + config, + "format.stage-frame", + &[("STAGE_PROMPT", &stage.prompt)], + ), + ) + .push_from( + "output-contract", + Some("format.output-contract"), + when_on( + config, + "format", + "output-contract", + sys(config, "\n\n", "format.output-contract", &[]), + ), + ) + .block(false), ]; - let user = format!( - "Text to format from the current chunk:\n{text}\n\n\ - Apply only the formatting instructions. Output only the complete formatted text." + let user = Parts::default().push_from( + "user-message", + Some("format.user-message"), + render(config, "format.user-message", &[("TEXT", text)]), ); - StructuredPrompt::new(system, user) + ComposedPrompt { + system, + user: user.into_vec(), + } } pub(crate) fn build_judge_prompts( @@ -283,75 +441,101 @@ pub(crate) fn build_judge_prompts( translation: &str, config: &PipelineConfig, ) -> StructuredPrompt { - let glossary_table = format_glossary_table(&config.glossary); - let src = effective_source(config); - let tgt = effective_target(config); - let ui_lang = config + compose_judge_prompts(source_text, translation, config).into_structured() +} + +fn ui_language(config: &PipelineConfig) -> &str { + config .ui_language .as_deref() .filter(|s| !s.is_empty()) - .unwrap_or(tgt); + .unwrap_or("English") +} - let glossary_section = if glossary_table.is_empty() { +pub(crate) fn compose_judge_prompts( + source_text: &str, + translation: &str, + config: &PipelineConfig, +) -> ComposedPrompt { + let glossary_table = format_glossary_table(&config.glossary); + let glossary = if glossary_table.is_empty() || !on(config, "audit", "glossary-table") { String::new() } else { - format!("Glossary to adhere to:\n{glossary_table}\n\n") - }; - - let markdown_rules = if config.markdown_aware.unwrap_or(false) { - "When Markdown is present, verify that the translation preserves markers, footnotes, \ - inline emphasis, and block structure exactly enough to remain valid Markdown.\n\n" - } else { - "" + sys( + config, + "\n\n", + "audit.glossary", + &[("GLOSSARY_TABLE", &glossary_table)], + ) }; + let markdown = + if config.markdown_aware.unwrap_or(false) && on(config, "audit", "markdown-rules") { + sys(config, "\n\n", "audit.markdown-rules", &[]) + } else { + String::new() + }; - // Block 1 (cacheable): static judge context — role, instructions, glossary, format spec. - // The source text and translation are in the user turn so this block is constant for the - // whole project run, enabling near-100% cache hit rate across all chunk judge calls. - let system_block = format!( - "You are a translation quality judge for {src}→{tgt} translations.\n\n\ - Specific Audit Instructions:\n{instructions}\n\n\ - {glossary_section}\ - {markdown_rules}\ - Scanning protocol: go through the translation sentence by sentence, checking every \ - sentence against the source for accuracy, every glossary term for adherence, grammar \ - for correctness, and fluency throughout. Complete the full scan before building the issues list. \ - Report EVERY issue you find and EVERY occurrence separately — do not merge, suppress, or \ - limit repeated issues.\n\n\ - You MUST respond with a valid JSON object containing:\n\ - - checkedSentenceIndices: array of 1-based source sentence numbers you verified, in scan order \ - (e.g. [1, 2, 3] for a 3-sentence source) — indices only, never the sentence text itself\n\ - - rating: one of 'critical', 'poor', 'fair', 'good', 'excellent' \ - (semantic translation quality: critical=unusable, poor=weak, fair=usable with revision, \ - good=solid, excellent=publication-ready)\n\ - - issues: array of objects with these fields:\n\ - - type: 'glossary'|'fluency'|'accuracy'|'grammar'\n\ - - severity: 'low'|'medium'|'high'\n\ - - description: string — explanation of the issue in {ui_lang}\n\ - - suggestedFix: string — how to correct it in {ui_lang}\n\ - - phrase: string or null — the exact verbatim substring of the WRONG or problematic text \ - as it appears in the TARGET translation (character-for-character copy from the target text)\n\ - - sourcePhrase: string or null — the exact verbatim substring from the SOURCE text \ - that corresponds to this issue\n\ - - confidence: number or null — your confidence this is a real issue (0.0–1.0)\n\ - Write description and suggestedFix in {ui_lang}. \ - Keep rating and type values as the English literals above.", - instructions = config.judge_prompt, - ); - - let user = format!( - "Source ({src}): {source_text}\n\ - Target ({tgt}): {translation}\n\n\ - Perform the audit now and return the JSON report." + // One cacheable block: the source text and translation are in the user turn so this + // block is constant for the whole project run, enabling near-100% cache hit rate + // across all chunk judge calls. + let system = vec![Parts::default() + .push_from( + "role", + Some("audit.role"), + when_on(config, "audit", "role", render(config, "audit.role", &[])), + ) + .push_from( + "translation-context", + Some("context-frame"), + context_part(config), + ) + .push_from( + "stage-prompt", + Some("audit.stage-frame"), + sys( + config, + "\n\n", + "audit.stage-frame", + &[("STAGE_PROMPT", &config.judge_prompt)], + ), + ) + .push_from("glossary-table", Some("audit.glossary"), glossary) + .push_from("markdown-rules", Some("audit.markdown-rules"), markdown) + .push_from( + "review-method", + Some("audit.review-method"), + when_on( + config, + "audit", + "review-method", + sys(config, "\n\n", "audit.review-method", &[]), + ), + ) + .push_from( + "response-format", + Some("audit.response-format"), + sys( + config, + "\n\n", + "audit.response-format", + &[("UI_LANGUAGE", ui_language(config))], + ), + ) + .block(true)]; + + let user = Parts::default().push_from( + "user-message", + Some("audit.user-message"), + render( + config, + "audit.user-message", + &[("TEXT", source_text), ("TRANSLATION", translation)], + ), ); - StructuredPrompt { - system: vec![PromptBlock { - text: system_block, - cacheable: true, - }], - user, - images: Vec::new(), + ComposedPrompt { + system, + user: user.into_vec(), } } @@ -359,94 +543,130 @@ pub(crate) fn build_coherence_prompts( input: &CoherenceChunkInput, config: &PipelineConfig, ) -> StructuredPrompt { - let glossary_table = format_glossary_table(&config.glossary); - let src = effective_source(config); - let tgt = effective_target(config); - let ui_lang = config - .ui_language - .as_deref() - .filter(|s| !s.is_empty()) - .unwrap_or(tgt); - - let default_instructions = "Evaluate ONLY:\n\ - 1. Terminology consistency — key terms translated differently than in adjacent segments\n\ - 2. Narrative continuity — abrupt breaks in flow at segment boundaries\n\ - 3. Glossary adherence — glossary terms used inconsistently with context\n\ - Do NOT re-evaluate standalone translation quality.\n\ - Be exhaustive: scan ALL dimensions completely before responding. Do not stop after finding \ - the first issue of each type. Only return an empty issues array if you are fully confident \ - — after deliberate review of every dimension — that no problems exist."; + compose_coherence_prompts(input, config).into_structured() +} +/// The coherence instructions used when the pipeline has none of its own. +pub(crate) const DEFAULT_COHERENCE_INSTRUCTIONS: &str = "Evaluate ONLY:\n\ + 1. Terminology consistency — key terms translated differently than in adjacent segments\n\ + 2. Narrative continuity — abrupt breaks in flow at segment boundaries\n\ + 3. Glossary adherence — glossary terms used inconsistently with context\n\ + Do NOT re-evaluate standalone translation quality.\n\ + Be exhaustive: scan ALL dimensions completely before responding. Do not stop after finding \ + the first issue of each type. Only return an empty issues array if you are fully confident \ + — after deliberate review of every dimension — that no problems exist."; + +pub(crate) fn compose_coherence_prompts( + input: &CoherenceChunkInput, + config: &PipelineConfig, +) -> ComposedPrompt { + let glossary_table = format_glossary_table(&config.glossary); let instructions = config .coherence_prompt .as_deref() .filter(|s| !s.trim().is_empty()) - .unwrap_or(default_instructions); - - let glossary_section = if glossary_table.is_empty() { + .unwrap_or(DEFAULT_COHERENCE_INSTRUCTIONS); + let glossary = if glossary_table.is_empty() || !on(config, "coherence", "glossary-table") { String::new() } else { - format!("Glossary:\n{glossary_table}\n\n") + sys( + config, + "\n", + "coherence.glossary", + &[("GLOSSARY_TABLE", &glossary_table)], + ) }; + // The glossary table ends with a line break of its own: after it one more blank line. + let response_sep = if glossary.is_empty() { "\n" } else { "\n\n" }; // Block 1 (cacheable): static coherence context — role, instructions, glossary, format spec. // Constant for the whole project run. - let system_block = format!( - "You are a translation coherence auditor for {src}→{tgt} translations.\n\ - Your task: identify cross-segment inconsistencies between a translated segment and its surrounding context.\n\ - {instructions}\n\ - {glossary_section}\ - Write description and suggestedFix values in {ui_lang}.\n\ - Respond with valid JSON only:\n\ - {{\"issues\": [{{\"type\": \"consistency\"|\"glossary\", \ - \"severity\": \"low\"|\"medium\"|\"high\", \ - \"description\": \"string\", \ - \"suggestedFix\": \"string\", \ - \"phrase\": \"exact verbatim substring of the WRONG text as it appears in the target translation, not the source term nor the correction; first occurrence only\"}}]}}", - ); + let mut system = vec![Parts::default() + .push_from( + "role", + Some("coherence.role"), + when_on( + config, + "coherence", + "role", + render(config, "coherence.role", &[]), + ), + ) + .push_from( + "translation-context", + Some("context-frame"), + context_part(config), + ) + .push_from( + "review-method", + Some("coherence.review-method"), + when_on( + config, + "coherence", + "review-method", + sys(config, "\n", "coherence.review-method", &[]), + ), + ) + .push("stage-prompt", format!("\n{instructions}")) + .push_from("glossary-table", Some("coherence.glossary"), glossary) + .push_from( + "response-format", + Some("coherence.response-format"), + sys( + config, + response_sep, + "coherence.response-format", + &[("UI_LANGUAGE", ui_language(config))], + ), + ) + .block(true)]; // Block 2 (cacheable): reference document block. Identical for every chunk in the same // blob, so it's a second cache breakpoint — placed in system, not the user turn, so // providers actually cache it instead of rebilling it at full price on every chunk. - let context_block = input + let neighbours_on = on(config, "coherence", "neighbour-chunks"); + if let Some(ctx) = input .blob_context .as_deref() - .filter(|s| !s.is_empty()) - .map(|ctx| format!( - "[Reference translated document block - context only]\n\ - This block may include the current chunk. Use it to compare terminology and continuity across the document block.\n\ - The current chunk to audit is identified below.\n\ - {ctx}\n\ - [End reference translated document block]" - )); + .filter(|s| !s.is_empty() && neighbours_on) + { + system.push( + Parts::default() + .push_from( + "neighbour-chunks", + Some("coherence.neighbours"), + render(config, "coherence.neighbours", &[("NEIGHBOUR_CHUNKS", ctx)]), + ) + .block(true), + ); + } - let current_chunk_line = input - .current_chunk_id - .as_deref() - .filter(|s| !s.is_empty()) - .map(|id| format!("Current chunk id: {id}\n\n")) - .unwrap_or_default(); - - let user = format!( - "{current_chunk_line}[Current segment]\nOriginal: {original}\nTranslation: {translation}\n\ - [End of current segment]\n\n\ - Identify cross-segment coherence issues and return the JSON. If no issues, return {{\"issues\": []}}.", - original = input.original, - translation = input.translation, - ); + let chunk_id = if neighbours_on { + chunk_id_part(config, input.current_chunk_id.as_deref()) + } else { + String::new() + }; + let user_sep = if chunk_id.is_empty() { "" } else { "\n\n" }; + let user = Parts::default() + .push_from("chunk-id", Some("chunk-id"), chunk_id) + .push_from( + "user-message", + Some("coherence.user-message"), + sys( + config, + user_sep, + "coherence.user-message", + &[ + ("TEXT", &input.original), + ("TRANSLATION", &input.translation), + ], + ), + ); - let mut system = vec![PromptBlock { - text: system_block, - cacheable: true, - }]; - if let Some(ctx) = context_block { - system.push(PromptBlock { - text: ctx, - cacheable: true, - }); + ComposedPrompt { + system, + user: user.into_vec(), } - - StructuredPrompt::new(system, user) } /// Strips markdown code fences and any preamble text that LLMs sometimes wrap around JSON output. @@ -513,8 +733,7 @@ mod tests { fn en_it_config() -> PipelineConfig { PipelineConfig { - source_language: "English".to_string(), - target_language: "Italian".to_string(), + work_brief: Some("Literary translation from English to Italian.".to_string()), ..Default::default() } } @@ -570,9 +789,11 @@ mod tests { // ── system block ────────────────────────────────────────────────── #[test] - fn system_includes_source_and_target_languages() { + fn system_includes_work_brief() { let prompt = build_coherence_prompts(&simple_input(), &en_it_config()); - assert!(prompt.system[0].text.contains("English→Italian")); + assert!(prompt.system[0] + .text + .contains("Translation context:\nLiterary translation from English to Italian.")); } #[test] @@ -699,8 +920,6 @@ mod tests { #[test] fn refine_user_turn_includes_audit_context_when_provided() { let config = PipelineConfig { - source_language: "English".to_string(), - target_language: "Italian".to_string(), ..Default::default() }; let stage = StageConfig { @@ -728,8 +947,6 @@ mod tests { #[test] fn refine_user_turn_omits_audit_section_when_context_is_none() { let config = PipelineConfig { - source_language: "English".to_string(), - target_language: "Italian".to_string(), ..Default::default() }; let stage = StageConfig { diff --git a/src-tauri/src/llm/provider.rs b/src-tauri/src/llm/provider.rs index 35db23d4..9046818c 100644 --- a/src-tauri/src/llm/provider.rs +++ b/src-tauri/src/llm/provider.rs @@ -64,6 +64,8 @@ impl UsageAccumulator { /// /// The factory [`crate::llm::providers::get_provider`] routes provider id strings to concrete /// types. Each provider owns its HTTP client selection, SSE parsing, and error formatting. +// async_trait adds must_use to boxed futures; Clippy 1.99 flags that generated attribute. +#[allow(clippy::double_must_use)] #[async_trait] pub trait LlmProvider: Send + Sync { fn id(&self) -> &'static str; diff --git a/src-tauri/src/llm/tests.rs b/src-tauri/src/llm/tests.rs index 23983f0b..104bef82 100644 --- a/src-tauri/src/llm/tests.rs +++ b/src-tauri/src/llm/tests.rs @@ -160,8 +160,6 @@ impl StreamChunkSource for MockChunkSource { fn make_config() -> PipelineConfig { PipelineConfig { - source_language: "English".into(), - target_language: "Italian".into(), stages: vec![], judge_prompt: "Evaluate translation quality.".into(), judge_model: "gemini-3-flash-preview".into(), @@ -176,10 +174,9 @@ fn make_config() -> PipelineConfig { markdown_aware: None, coherence_prompt: None, review_provider_options: None, - persona: None, + work_brief: Some("Literary translation from English to Italian.".into()), + prompt_composition: Default::default(), ui_language: None, - custom_source_language: None, - custom_target_language: None, blob_context: None, blob_current_chunk_id: None, } @@ -290,7 +287,7 @@ fn stage_prompt_without_previous() { let prompt = build_stage_prompts("Hello world", &stage, &config, None, None); let system = prompt.flatten_system(); - assert!(system.contains("English to Italian")); + assert!(system.contains("Translation context:\nLiterary translation from English to Italian.")); assert!(system.contains("Translate accurately.")); assert!(system.contains("| API | API |")); assert!(prompt.user.contains("Hello world")); @@ -308,7 +305,7 @@ fn stage_prompt_with_blob_context() { let prompt = build_stage_prompts("Hello world", &stage, &config, None, None); let system = prompt.flatten_system(); - assert!(system.contains("English to Italian")); + assert!(system.contains("Translation context:\nLiterary translation from English to Italian.")); assert!(prompt.user.contains("Hello world")); assert!(prompt.user.contains("Current chunk id: chunk-1")); assert!(system.contains("Reference document block")); @@ -317,7 +314,7 @@ fn stage_prompt_with_blob_context() { assert!(prompt.system[0].cacheable); assert!(prompt.system[1].cacheable); assert!(!prompt.system[2].cacheable); - assert!(prompt.system[0].text.contains("English to Italian")); + assert!(prompt.system[0].text.contains("Translation context:")); assert!(prompt.system[1].text.contains("Reference document block")); assert!(prompt.system[2].text.contains("Core Instructions")); } @@ -333,7 +330,7 @@ fn stage_prompt_refine_includes_previous_iteration() { let prompt = build_stage_prompts("Hello world", &stage, &config, prev.as_deref(), None); let system = prompt.flatten_system(); - assert!(system.contains("English to Italian")); + assert!(system.contains("Translation context:\nLiterary translation from English to Italian.")); assert!(prompt.user.contains("Hello world")); assert!(prompt.user.contains("Ciao mondo")); assert!(prompt.user.contains("Previous Iteration")); @@ -373,8 +370,7 @@ fn stage_prompt_multiple_glossary_entries() { assert!(system.contains("| API | API | tech |")); assert!(system.contains("| bug | errore |")); assert!(system.contains("Treat every glossary entry as mandatory terminology")); - assert!(system.contains("Glossary Reminder")); - assert!(system.contains("Apply the glossary entries specified above")); + assert!(!system.contains("Glossary Reminder")); } #[test] @@ -394,7 +390,7 @@ fn format_stage_prompt_omits_glossary_persona_and_source_context() { assert!(system.contains("Output only the formatted text")); assert!(!system.contains("Glossary Constraints")); assert!(!system.contains("| API | API |")); - assert!(!system.contains("English to Italian")); + assert!(!system.contains("Translation context")); assert!(!system.contains("Reference document block")); assert!(!system.contains("Output only the translated text")); assert!(prompt.user.contains("Text to format")); @@ -418,17 +414,39 @@ fn markdown_aware_stage_prompt_preserves_syntax() { assert!(prompt.user.contains("Text with note[^1].")); } +#[test] +fn switched_off_parts_are_left_out_only_in_their_phase() { + let mut config = make_config(); + config.prompt_composition.disabled = vec![ + "refine:structural-rules".into(), + "audit:review-method".into(), + ]; + let translation = make_stage("openai"); + let mut refine = make_stage("openai"); + refine.role = Some("refine".into()); + + let translation_system = + build_stage_prompts("Hello", &translation, &config, None, None).flatten_system(); + let refine_system = + build_stage_prompts("Hello", &refine, &config, Some("Ciao"), None).flatten_system(); + let judge_system = build_judge_prompts("Hello", "Ciao", &config).flatten_system(); + + assert!(translation_system.contains("Structural Preservation Rules")); + assert!(!refine_system.contains("Structural Preservation Rules")); + assert!(!judge_system.contains("Scanning protocol")); + assert!(judge_system.contains("You MUST respond")); +} + // ── build_judge_prompts ────────────────────────────────────────── #[test] -fn judge_prompt_includes_source_and_target() { +fn judge_prompt_includes_work_brief_and_texts() { let config = make_config(); let prompt = build_judge_prompts("Hello", "Ciao", &config); let system = prompt.flatten_system(); // Source/target text now live in the user turn for cacheability - assert!(system.contains("English")); - assert!(system.contains("Italian")); + assert!(system.contains("Translation context:\nLiterary translation from English to Italian.")); assert!(!system.contains("Hello")); assert!(!system.contains("Ciao")); assert!(prompt.user.contains("Hello")); @@ -725,7 +743,7 @@ fn pipeline_config_roundtrip() { let config = make_config(); let json = serde_json::to_string(&config).unwrap(); let parsed: PipelineConfig = serde_json::from_str(&json).unwrap(); - assert_eq!(parsed.source_language, "English"); + assert_eq!(parsed.judge_model, "gemini-3-flash-preview"); assert_eq!(parsed.glossary.len(), 1); } diff --git a/src-tauri/src/llm/types.rs b/src-tauri/src/llm/types.rs index 67a05758..f8c53781 100644 --- a/src-tauri/src/llm/types.rs +++ b/src-tauri/src/llm/types.rs @@ -167,11 +167,19 @@ pub struct StageConfig { pub custom_provider_id: Option, } +/// Pipeline-level overrides of the system texts (by id) and parts switched off. +#[derive(Debug, Clone, Default, Serialize, Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct PromptComposition { + #[serde(default)] + pub texts: std::collections::HashMap, + #[serde(default)] + pub disabled: Vec, +} + #[derive(Debug, Clone, Default, Serialize, Deserialize)] #[serde(rename_all = "camelCase")] pub struct PipelineConfig { - pub source_language: String, - pub target_language: String, pub stages: Vec, pub judge_prompt: String, pub judge_model: String, @@ -186,10 +194,13 @@ pub struct PipelineConfig { pub markdown_aware: Option, pub coherence_prompt: Option, pub review_provider_options: Option, - pub persona: Option, + /// Shared task context for translation, refinement and quality checks. + #[serde(default)] + pub work_brief: Option, + /// Custom system texts and disabled parts of the prompts. + #[serde(default)] + pub prompt_composition: PromptComposition, pub ui_language: Option, - pub custom_source_language: Option, - pub custom_target_language: Option, /// Original text of all chunks in the same blob, injected at call time for context. /// Not persisted — computed from blob assignments before each LLM invocation. pub blob_context: Option, diff --git a/src-tauri/src/provenance.rs b/src-tauri/src/provenance.rs index 99abf3d5..9485242c 100644 --- a/src-tauri/src/provenance.rs +++ b/src-tauri/src/provenance.rs @@ -58,6 +58,8 @@ pub struct Event { /// **allora**. pub input_hash: Option, pub output_hash: Option, + pub provider: Option, + pub model: Option, /// Il resto, in JSON: i dettagli propri di quel tipo di evento. pub config: Option, /// Che cosa rende **distinto** questo fatto dagli altri dello stesso tipo @@ -91,6 +93,8 @@ impl Event { error_kind: None, input_hash: None, output_hash: None, + provider: None, + model: None, config: Some(serde_json::json!({ "jobType": job_type }).to_string()), key_ref: None, } @@ -140,12 +144,13 @@ pub fn fnv1a_hex(text: &str) -> String { pub fn record(conn: &Connection, event: &Event) -> Result<(), String> { conn.execute( "INSERT INTO provenance_events (id, event_type, entity_type, entity_id, workspace_id, \ - actor, job_id, outcome, duration_ms, error_kind, input_hash, output_hash, config) \ - VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12, ?13) \ + actor, job_id, outcome, duration_ms, error_kind, input_hash, output_hash, config, provider, model) \ + VALUES (?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, ?10, ?11, ?12, ?13, ?14, ?15) \ ON CONFLICT(id) DO UPDATE SET occurred_at = CURRENT_TIMESTAMP, \ outcome = excluded.outcome, duration_ms = excluded.duration_ms, \ error_kind = excluded.error_kind, input_hash = excluded.input_hash, \ - output_hash = excluded.output_hash, config = excluded.config", + output_hash = excluded.output_hash, config = excluded.config, \ + provider = excluded.provider, model = excluded.model", params![ event_id(event), event.event_type, @@ -160,6 +165,8 @@ pub fn record(conn: &Connection, event: &Event) -> Result<(), String> { event.input_hash, event.output_hash, event.config, + event.provider, + event.model, ], ) .map_err(|error| format!("registro dei fatti: {error}"))?; @@ -189,7 +196,9 @@ mod tests { duration_ms INTEGER, error_kind TEXT, input_hash TEXT, - output_hash TEXT + output_hash TEXT, + provider TEXT, + model TEXT );", ) .unwrap(); diff --git a/src-tauri/src/vector/embedding.rs b/src-tauri/src/vector/embedding.rs index 648679ad..d9df9e2d 100644 --- a/src-tauri/src/vector/embedding.rs +++ b/src-tauri/src/vector/embedding.rs @@ -1,7 +1,6 @@ use serde::{Deserialize, Serialize}; use std::sync::{Arc, Mutex}; use std::time::{Duration, Instant}; -use tauri::State; const OPENAI_CONNECT_TIMEOUT_SECS: u64 = 10; const OPENAI_REQUEST_TIMEOUT_SECS: u64 = 45; @@ -22,7 +21,7 @@ impl Serialize for EmbeddingError { } } -async fn run_blocking( +pub(super) async fn run_blocking( connection: Arc>, operation: impl FnOnce(&mut rusqlite::Connection) -> Result + Send + 'static, ) -> Result { @@ -36,7 +35,7 @@ async fn run_blocking( .map_err(|error| EmbeddingError::Http(format!("database task failed: {error}")))? } -fn floats_to_blob(v: &[f32]) -> Vec { +pub(super) fn floats_to_blob(v: &[f32]) -> Vec { v.iter().flat_map(|f| f.to_le_bytes()).collect() } @@ -52,6 +51,7 @@ fn openai_client() -> Result { #[derive(Deserialize)] struct EmbeddingObject { + index: usize, embedding: Vec, } @@ -71,6 +71,7 @@ pub async fn get_embeddings( if texts.is_empty() { return Ok(vec![]); } + super::text_units::dimensions(&model)?; let request_started = Instant::now(); let total_chars: usize = texts.iter().map(|text| text.len()).sum(); log::debug!( @@ -135,428 +136,75 @@ pub async fn get_embeddings( parsed.data.len(), request_started.elapsed().as_millis() ); - Ok(parsed.data.into_iter().map(|o| o.embedding).collect()) + ordered_embeddings(parsed.data, texts.len(), &model) } -/// Le frasi che un workspace vede (#213). -/// -/// Una frase nata da una traduzione **segue il suo progetto**: il workspace è -/// quello del progetto, e spostare il progetto porta con sé migliaia di righe -/// senza toccarne una. Una frase importata non ha un progetto, e si collega da -/// sola. Prima la colonna sulla riga diceva entrambe le cose, e si -/// disallineava al primo spostamento. -const IN_WORKSPACE: &str = "(pm.project_id IN (SELECT id FROM projects WHERE workspace_id = :ws) \ - OR EXISTS (SELECT 1 FROM workspace_items wi \ - WHERE wi.item_type = 'phrase' AND wi.item_id = pm.id AND wi.workspace_id = :ws))"; - -#[derive(Debug, Serialize, Deserialize)] -pub struct PhraseMatchResult { - pub phrase_memory_id: String, - pub source_phrase: String, - pub target_phrase: String, - pub distance: f64, - pub confidence: f64, -} - -#[derive(Debug, Serialize, Deserialize)] -pub struct PhraseMemoryEntryResult { - pub id: String, - pub workspace_id: String, - pub source_phrase: String, - pub target_phrase: String, - pub confidence: f64, - pub source_language: String, - pub target_language: String, - pub author: Option, - pub work: Option, - pub domain: Option, - pub tags: Option, - pub notes: Option, - pub chunk_id: Option, - pub project_id: Option, - pub embedding_model: Option, - pub created_at: String, -} - -#[tauri::command] -pub async fn vec_list_phrase_memory( - database: State<'_, crate::vector::VectorDatabase>, - workspace_id: String, -) -> Result, EmbeddingError> { - let connection = database.connection().map_err(EmbeddingError::Http)?; - run_blocking(connection, move |conn| { - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - let query = format!( - "SELECT pm.id, pm.source_phrase, pm.target_phrase, pm.confidence, pm.source_language, \ - pm.target_language, pm.author, pm.work, pm.domain, pm.tags, pm.notes, \ - pm.chunk_id, pm.project_id, pm.embedding_model, pm.created_at \ - FROM phrase_memory pm WHERE {IN_WORKSPACE} \ - ORDER BY datetime(pm.created_at) DESC, pm.id DESC" - ); - let mut statement = conn - .prepare(&query) - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - let entries = statement - .query_map(rusqlite::named_params! { ":ws": workspace_id }, |row| { - Ok(PhraseMemoryEntryResult { - id: row.get(0)?, - workspace_id: workspace_id.clone(), - source_phrase: row.get(1)?, - target_phrase: row.get(2)?, - confidence: row.get(3)?, - source_language: row.get(4)?, - target_language: row.get(5)?, - author: row.get(6)?, - work: row.get(7)?, - domain: row.get(8)?, - tags: row.get(9)?, - notes: row.get(10)?, - chunk_id: row.get(11)?, - project_id: row.get(12)?, - embedding_model: row.get(13)?, - created_at: row.get(14)?, - }) - }) - .map_err(|error| EmbeddingError::Http(error.to_string()))? - .collect::>>() - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - Ok(entries) - }) - .await -} - -#[tauri::command] -pub async fn vec_delete_phrase_memory( - database: State<'_, crate::vector::VectorDatabase>, - write_coordinator: State<'_, crate::db::DbWriteCoordinator>, - workspace_id: String, - phrase_memory_id: String, -) -> Result { - let _write_guard = write_coordinator.lock().await; - let connection = database.connection().map_err(EmbeddingError::Http)?; - run_blocking(connection, move |conn| { - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - // La frase si cancella solo se **quel** workspace la vede: senza il - // controllo, un id indovinato toglierebbe una frase di un altro. - let query = format!( - "DELETE FROM phrase_memory WHERE id = :id \ - AND EXISTS (SELECT 1 FROM phrase_memory pm WHERE pm.id = :id AND {IN_WORKSPACE})" - ); - conn.execute( - &query, - rusqlite::named_params! { ":id": phrase_memory_id, ":ws": workspace_id }, - ) - .map(|count| count as u32) - .map_err(|error| EmbeddingError::Http(error.to_string())) - }) - .await -} - -#[tauri::command] -pub async fn vec_update_phrase_memory( - database: State<'_, crate::vector::VectorDatabase>, - write_coordinator: State<'_, crate::db::DbWriteCoordinator>, - workspace_id: String, - phrase_memory_id: String, - source_phrase: String, - target_phrase: String, - embedding: Vec, -) -> Result { - let _write_guard = write_coordinator.lock().await; - let connection = database.connection().map_err(EmbeddingError::Http)?; - run_blocking(connection, move |conn| { - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - let query = format!( - "UPDATE phrase_memory SET source_phrase = :source, target_phrase = :target, \ - embedding = :embedding \ - WHERE id = :id \ - AND EXISTS (SELECT 1 FROM phrase_memory pm WHERE pm.id = :id AND {IN_WORKSPACE})" - ); - conn.execute( - &query, - rusqlite::named_params! { - ":source": source_phrase, - ":target": target_phrase, - ":embedding": floats_to_blob(&embedding), - ":id": phrase_memory_id, - ":ws": workspace_id, - }, - ) - .map(|count| count as u32) - .map_err(|error| EmbeddingError::Http(error.to_string())) - }) - .await -} - -#[tauri::command] -pub async fn vec_search_phrase_memory( - database: State<'_, crate::vector::VectorDatabase>, - workspace_id: String, - query_embedding: Vec, - threshold: f64, - max_results: u32, - embedding_model: String, -) -> Result, EmbeddingError> { - let blob = floats_to_blob(&query_embedding); - let connection = database.connection().map_err(EmbeddingError::Http)?; - run_blocking(connection, move |conn| { - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - let mut statement = conn - .prepare(&format!( - "WITH ranked AS ( \ - SELECT pm.id, pm.source_phrase, pm.target_phrase, pm.confidence, \ - vec_distance_cosine(pm.embedding, :query) AS distance \ - FROM phrase_memory pm \ - WHERE {IN_WORKSPACE} \ - AND (pm.embedding_model IS NULL OR pm.embedding_model = :model) \ - ) \ - SELECT id, source_phrase, target_phrase, confidence, distance FROM ranked \ - WHERE distance < :threshold ORDER BY distance ASC LIMIT :limit" - )) - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - let matches = statement - .query_map( - rusqlite::named_params! { - ":query": blob, - ":ws": workspace_id, - ":threshold": threshold, - ":limit": max_results, - ":model": embedding_model, - }, - |row| { - Ok(PhraseMatchResult { - phrase_memory_id: row.get(0)?, - source_phrase: row.get(1)?, - target_phrase: row.get(2)?, - confidence: row.get(3)?, - distance: row.get(4)?, - }) - }, - ) - .map_err(|error| EmbeddingError::Http(error.to_string()))? - .collect::>>() - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - Ok(matches) - }) - .await -} - -#[derive(Debug, Deserialize)] -#[serde(rename_all = "camelCase")] -pub struct PhrasePair { - pub source_phrase: String, - pub target_phrase: String, - pub confidence: f64, - pub source_embedding: Vec, -} - -// 2 State injection + 7 parametri di dominio: raggruppabili in una request struct, -// ma cambierebbe la firma del comando Tauri — rimandato a un refactor dedicato. -#[allow(clippy::too_many_arguments)] -#[tauri::command] -pub async fn vec_save_locked_phrases( - database: State<'_, crate::vector::VectorDatabase>, - write_coordinator: State<'_, crate::db::DbWriteCoordinator>, - // Il workspace non si passa più: una frase nata da una traduzione sta dove - // sta il progetto, e chiederlo due volte era il modo di farli divergere. - project_id: String, - chunk_id: String, - pairs: Vec, - source_language: String, - target_language: String, - embedding_model: String, -) -> Result { - let save_started = Instant::now(); - log::debug!( - "phrase_memory.vec_save_locked_phrases.start project_id={project_id} chunk_id={chunk_id} pair_count={}", - pairs.len() - ); - if pairs.is_empty() { - return Ok(0); +fn ordered_embeddings( + mut data: Vec, + count: usize, + model: &str, +) -> Result>, EmbeddingError> { + if data.len() != count { + return Err(EmbeddingError::Parse( + "Incomplete embedding response".into(), + )); } - - let _write_guard = write_coordinator.lock().await; - let connection = database.connection().map_err(EmbeddingError::Http)?; - run_blocking(connection, move |conn| { - - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - - let project_exists: i64 = conn - .query_row( - "SELECT COUNT(*) FROM projects WHERE id = ?1", - rusqlite::params![&project_id], - |row| row.get(0), - ) - .map_err(|e| { - log::warn!( - "phrase_memory.vec_save_locked_phrases.project_check_failed project_id={project_id} error={e}" - ); - EmbeddingError::Http(e.to_string()) - })?; - log::debug!( - "phrase_memory.vec_save_locked_phrases.refs project_id={project_id} project_exists={project_exists}" - ); - if project_exists == 0 { - return Err(EmbeddingError::Http(format!( - "phrase memory references missing: project_id={project_id}" - ))); + data.sort_by_key(|item| item.index); + data.into_iter() + .enumerate() + .map(|(index, item)| { + if item.index != index { + return Err(EmbeddingError::Parse( + "Invalid embedding response indices".into(), + )); + } + super::text_units::validate_embedding(model, &item.embedding)?; + Ok(item.embedding) + }) + .collect() +} + +#[cfg(test)] +mod tests { + use super::*; + + fn object(index: usize, value: f32) -> EmbeddingObject { + EmbeddingObject { + index, + embedding: vec![value; 1536], + } } - let mut saved: u32 = 0; - let mut attempted: u32 = 0; - let tx = conn.transaction().map_err(|e| { - log::warn!("phrase_memory.vec_save_locked_phrases.transaction_failed error={e}"); - EmbeddingError::Http(e.to_string()) - })?; - - let replaced_rows = tx - .execute( - "DELETE FROM phrase_memory WHERE chunk_id = ?1 AND project_id = ?2", - rusqlite::params![&chunk_id, &project_id], + #[test] + fn provider_indices_restore_order_before_pair_alignment() { + let vectors = ordered_embeddings( + vec![object(1, 2.0), object(0, 1.0)], + 2, + "text-embedding-3-small", ) - .map_err(|e| { - log::warn!( - "phrase_memory.vec_save_locked_phrases.replace_failed project_id={project_id} chunk_id={chunk_id} error={e}" - ); - EmbeddingError::Http(e.to_string()) - })?; - log::debug!( - "phrase_memory.vec_save_locked_phrases.replace_done deleted_phrase_memory_rows={replaced_rows}" - ); - - for (index, pair) in pairs.iter().enumerate() { - attempted += 1; - let rows = tx - .execute( - "INSERT OR IGNORE INTO phrase_memory \ - (id, project_id, chunk_id, source_phrase, target_phrase, \ - confidence, source_language, target_language, embedding, embedding_model, created_at) \ - VALUES (lower(hex(randomblob(16))), ?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, ?9, datetime('now'))", - rusqlite::params![ - &project_id, - &chunk_id, - pair.source_phrase, - pair.target_phrase, - pair.confidence.clamp(0.0, 1.0), - &source_language, - &target_language, - floats_to_blob(&pair.source_embedding), - &embedding_model, - ], - ) - .map_err(|e| { - log::warn!( - "phrase_memory.vec_save_locked_phrases.insert_failed project_id={project_id} chunk_id={chunk_id} pair_index={index} source_chars={} target_chars={} embedding_dim={} error={e}", - pair.source_phrase.len(), - pair.target_phrase.len(), - pair.source_embedding.len() - ); - EmbeddingError::Http(format!("phrase memory insert failed at pair {index}: {e}")) - })?; - - saved += rows as u32; + .expect("valid response"); + assert_eq!(vectors[0][0], 1.0); + assert_eq!(vectors[1][0], 2.0); } - log::debug!( - "phrase_memory.vec_save_locked_phrases.insert_loop_done pair_count={} attempted={attempted} saved={saved}", - pairs.len() - ); - let deleted_source_embeddings = tx - .execute( - "DELETE FROM source_phrase_embeddings WHERE chunk_id = ?1 AND project_id = ?2", - rusqlite::params![&chunk_id, &project_id], + #[test] + fn incomplete_duplicate_or_out_of_range_provider_indices_are_rejected() { + assert!(ordered_embeddings(vec![object(0, 1.0)], 2, "text-embedding-3-small").is_err()); + assert!(ordered_embeddings( + vec![object(0, 1.0), object(0, 1.0)], + 2, + "text-embedding-3-small" ) - .map_err(|e| { - log::warn!( - "phrase_memory.vec_save_locked_phrases.cleanup_failed project_id={project_id} chunk_id={chunk_id} error={e}" - ); - EmbeddingError::Http(e.to_string()) - })?; - log::debug!( - "phrase_memory.vec_save_locked_phrases.cleanup_done deleted_source_embeddings={deleted_source_embeddings}" - ); - tx.commit().map_err(|e| { - log::warn!("phrase_memory.vec_save_locked_phrases.commit_failed error={e}"); - EmbeddingError::Http(e.to_string()) - })?; - - log::info!( - "phrase_memory.vec_save_locked_phrases.done project_id={project_id} chunk_id={chunk_id} pair_count={} attempted={attempted} saved={saved} elapsed_ms={}", - pairs.len(), - save_started.elapsed().as_millis() - ); - Ok(saved) - }) - .await -} - -#[tauri::command] -pub async fn vec_regenerate_all_embeddings( - app: tauri::AppHandle, - database: State<'_, crate::vector::VectorDatabase>, - write_coordinator: State<'_, crate::db::DbWriteCoordinator>, - workspace_id: String, - model: String, -) -> Result { - log::debug!( - "phrase_memory.vec_regenerate_all_embeddings.start workspace_id={workspace_id} model={model}" - ); - - // Phase 1: collect entries without blocking the async runtime. - let connection = database.connection().map_err(EmbeddingError::Http)?; - let query_workspace_id = workspace_id.clone(); - let entries: Vec<(String, String)> = run_blocking(Arc::clone(&connection), move |conn| { - crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; - let mut statement = conn - .prepare(&format!( - "SELECT pm.id, pm.source_phrase FROM phrase_memory pm WHERE {IN_WORKSPACE}" - )) - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - let entries = statement - .query_map( - rusqlite::named_params! { ":ws": query_workspace_id }, - |row| Ok((row.get::<_, String>(0)?, row.get::<_, String>(1)?)), - ) - .map_err(|error| EmbeddingError::Http(error.to_string()))? - .collect::>() - .map_err(|error| EmbeddingError::Http(error.to_string()))?; - Ok(entries) - }) - .await?; + .is_err()); + assert!(ordered_embeddings(vec![object(1, 1.0)], 1, "text-embedding-3-small").is_err()); + } - if entries.is_empty() { - log::debug!( - "phrase_memory.vec_regenerate_all_embeddings.empty workspace_id={workspace_id}" + #[test] + fn provider_values_and_dimensions_are_validated_before_any_write() { + assert!( + ordered_embeddings(vec![object(0, f32::INFINITY)], 1, "text-embedding-3-small") + .is_err() ); - return Ok(0); + assert!(ordered_embeddings(vec![object(0, 1.0)], 1, "text-embedding-3-large").is_err()); } - - // Phase 2: embed (async — no conn held across this await) - let phrases: Vec = entries.iter().map(|(_, p)| p.clone()).collect(); - let embeddings = get_embeddings(app.clone(), phrases, model.clone()).await?; - - // Phase 3: update without blocking the async runtime. - let update_model = model.clone(); - let _write_guard = write_coordinator.lock().await; - let updated: u32 = run_blocking(connection, move |conn| { - entries - .iter() - .zip(embeddings.iter()) - .map(|((id, _), embedding)| { - conn.execute( - "UPDATE phrase_memory SET embedding = ?1, embedding_model = ?2 WHERE id = ?3", - rusqlite::params![floats_to_blob(embedding), &update_model, id], - ) - .map(|count| count as u32) - }) - .collect::>>() - .map_err(|error| EmbeddingError::Http(error.to_string())) - .map(|counts| counts.into_iter().sum()) - }) - .await?; - - log::info!( - "phrase_memory.vec_regenerate_all_embeddings.done workspace_id={workspace_id} model={model} updated={updated}" - ); - Ok(updated) } diff --git a/src-tauri/src/vector/memory_commands.rs b/src-tauri/src/vector/memory_commands.rs new file mode 100644 index 00000000..dba70266 --- /dev/null +++ b/src-tauri/src/vector/memory_commands.rs @@ -0,0 +1,258 @@ +use super::embedding::{get_embeddings, run_blocking, EmbeddingError}; +use super::memory_search::PhraseMatchResult; +use super::text_units::{self, EmbeddingInput, MemoryEntry, PhrasePair, UpdateInput}; +use super::VectorDatabase; +use tauri::State; + +fn db(error: rusqlite::Error) -> EmbeddingError { + EmbeddingError::Http(error.to_string()) +} + +#[tauri::command] +pub async fn vec_list_phrase_memory( + database: State<'_, VectorDatabase>, + workspace_id: Option, + chunk_id: Option, +) -> Result, EmbeddingError> { + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| { + crate::vector::verify_phrase_memory_schema(conn).map_err(EmbeddingError::Http)?; + text_units::list(conn, workspace_id.as_deref(), chunk_id.as_deref()) + }, + ) + .await +} + +#[tauri::command] +pub async fn vec_get_phrase_memory( + database: State<'_, VectorDatabase>, + workspace_id: Option, + phrase_memory_id: String, +) -> Result { + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| text_units::get(conn, workspace_id.as_deref(), &phrase_memory_id), + ) + .await +} + +#[tauri::command] +pub async fn vec_update_phrase_memory( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + input: UpdateInput, +) -> Result { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| text_units::update(conn, input), + ) + .await +} + +#[tauri::command] +pub async fn vec_delete_phrase_memory( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + workspace_id: Option, + phrase_memory_id: String, +) -> Result { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| { + let tx = conn.transaction().map_err(db)?; + let entry = text_units::get(&tx, workspace_id.as_deref(), &phrase_memory_id)?; + // The memory owns this unit; remove its link before immutable revisions. + tx.execute("DELETE FROM phrase_memory WHERE id=?1", [&entry.id]) + .map_err(db)?; + tx.execute("DELETE FROM text_units WHERE id=?1", [entry.unit_id]) + .map_err(db)?; + tx.commit().map_err(db)?; + Ok(1) + }, + ) + .await +} + +#[allow(clippy::too_many_arguments)] +#[tauri::command] +pub async fn vec_save_locked_phrases( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + project_id: String, + chunk_id: String, + pairs: Vec, + embedding_model: String, +) -> Result { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| text_units::save_pairs(conn, &project_id, &chunk_id, &embedding_model, pairs), + ) + .await +} + +/// Frasi salvate dall'opera con lingue diverse da quelle attuali dell'opera. +#[tauri::command] +pub async fn vec_count_project_phrase_relabels( + database: State<'_, VectorDatabase>, + project_id: String, +) -> Result { + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| super::text_languages::count_relabels(conn, &project_id), + ) + .await +} + +/// Dà alle frasi salvate dall'opera le lingue attuali dell'opera. +#[tauri::command] +pub async fn vec_relabel_project_phrases( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + project_id: String, +) -> Result { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| super::text_languages::relabel_project(conn, &project_id), + ) + .await +} + +#[allow(clippy::too_many_arguments)] +#[tauri::command] +pub async fn vec_search_phrase_memory( + database: State<'_, VectorDatabase>, + workspace_id: String, + query_embedding: Vec, + threshold: f64, + max_results: u32, + embedding_model: String, + all_workspaces: bool, + source_language: Option, + target_language: Option, +) -> Result, EmbeddingError> { + let input = super::memory_search::SearchInput { + workspace_id, + query_embedding, + threshold, + max_results, + embedding_model, + all_workspaces, + source_language, + target_language, + }; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| super::memory_search::search(conn, input), + ) + .await +} + +#[tauri::command] +pub async fn vec_add_phrase_embedding( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + workspace_id: Option, + phrase_memory_id: String, + source_revision_id: String, + embedding: EmbeddingInput, +) -> Result<(), EmbeddingError> { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| { + let tx = conn.transaction().map_err(db)?; + let entry = text_units::get(&tx, workspace_id.as_deref(), &phrase_memory_id)?; + if entry.source_revision_id != source_revision_id { + return Err(EmbeddingError::Parse( + "Source changed; reload before recalculating".into(), + )); + } + text_units::put_embedding(&tx, &source_revision_id, &embedding)?; + tx.commit().map_err(db) + }, + ) + .await +} + +#[tauri::command] +pub async fn vec_set_phrase_tags( + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + workspace_id: Option, + phrase_memory_id: String, + tags: Vec, +) -> Result<(), EmbeddingError> { + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| { + let tx = conn.transaction().map_err(db)?; + let entry = text_units::get(&tx, workspace_id.as_deref(), &phrase_memory_id)?; + text_units::set_tags(&tx, &entry.unit_id, &tags)?; + tx.commit().map_err(db) + }, + ) + .await +} + +#[tauri::command] +pub async fn vec_regenerate_all_embeddings( + app: tauri::AppHandle, + database: State<'_, VectorDatabase>, + write_coordinator: State<'_, crate::db::DbWriteCoordinator>, + workspace_id: String, + model: String, +) -> Result { + text_units::dimensions(&model)?; + let ws = workspace_id.clone(); + let entries = run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| text_units::list(conn, Some(&ws), None), + ) + .await?; + if entries.is_empty() { + return Ok(0); + } + let vectors = get_embeddings( + app, + entries.iter().map(|e| e.source_phrase.clone()).collect(), + model.clone(), + ) + .await?; + if entries.len() != vectors.len() { + return Err(EmbeddingError::Parse( + "Incomplete embedding response".into(), + )); + } + let _guard = write_coordinator.lock().await; + run_blocking( + database.connection().map_err(EmbeddingError::Http)?, + move |conn| { + let tx = conn.transaction().map_err(db)?; + for (entry, embedding) in entries.iter().zip(vectors) { + let current = text_units::get(&tx, Some(&workspace_id), &entry.id)?; + text_units::ensure_revisions( + ¤t, + &entry.source_revision_id, + &entry.target_revision_id, + )?; + text_units::put_embedding( + &tx, + &entry.source_revision_id, + &EmbeddingInput { + model: model.clone(), + embedding, + }, + )?; + } + tx.commit().map_err(db)?; + Ok(entries.len() as u32) + }, + ) + .await +} diff --git a/src-tauri/src/vector/memory_search.rs b/src-tauri/src/vector/memory_search.rs new file mode 100644 index 00000000..ebdf46d2 --- /dev/null +++ b/src-tauri/src/vector/memory_search.rs @@ -0,0 +1,77 @@ +use super::embedding::{floats_to_blob, EmbeddingError}; +use super::text_units::{self, PROFILE}; +use rusqlite::Connection; +use serde::Serialize; +fn db(error: rusqlite::Error) -> EmbeddingError { + EmbeddingError::Http(error.to_string()) +} +#[derive(Serialize)] +pub struct PhraseMatchResult { + pub phrase_memory_id: String, + pub source_phrase: String, + pub target_phrase: String, + pub distance: f64, + pub confidence: f64, + pub workspace_id: Option, + pub project_id: Option, + pub chunk_id: Option, + pub source_id: Option, + pub source_language: String, + pub target_language: String, + pub source_language_variety: Option, + pub target_language_variety: Option, + pub provenance: serde_json::Value, + pub embedding_model: String, + pub dimensions: usize, +} + +pub struct SearchInput { + pub workspace_id: String, + pub query_embedding: Vec, + pub threshold: f64, + pub max_results: u32, + pub embedding_model: String, + pub all_workspaces: bool, + pub source_language: Option, + pub target_language: Option, +} + +pub fn search( + conn: &Connection, + input: SearchInput, +) -> Result, EmbeddingError> { + let SearchInput { + workspace_id, + query_embedding, + threshold, + max_results, + embedding_model, + all_workspaces, + source_language, + target_language, + } = input; + text_units::validate_embedding(&embedding_model, &query_embedding)?; + if !threshold.is_finite() || !(0.0..=1.0).contains(&threshold) || max_results == 0 { + return Err(EmbeddingError::Parse( + "Invalid search threshold or result limit".into(), + )); + } + let mut query=conn.prepare("WITH compatible AS MATERIALIZED ( + SELECT pm.*, e.embedding, sr.language_variety AS source_language_variety, tr.language_variety AS target_language_variety + FROM phrase_memory_entries pm JOIN text_embeddings e ON e.revision_id=pm.source_revision_id + JOIN text_unit_revisions sr ON sr.id=pm.source_revision_id + JOIN text_unit_revisions tr ON tr.id=pm.target_revision_id + WHERE (:all=1 OR pm.workspace_id=:ws) AND e.provider='openai' AND e.model=:model AND e.dimensions=:dim AND e.profile=:profile + AND (:src IS NULL OR pm.source_language=:src) AND (:tgt IS NULL OR pm.target_language=:tgt) + ), ranked AS ( + SELECT *,vec_distance_cosine(embedding,:query) AS distance FROM compatible + ) SELECT id,source_phrase,target_phrase,distance,confidence,workspace_id,project_id,chunk_id,source_id, + provenance,source_language,target_language,source_language_variety,target_language_variety + FROM ranked WHERE distance<:threshold ORDER BY distance,id LIMIT :limit").map_err(db)?; + let rows=query.query_map(rusqlite::named_params! {":ws":workspace_id,":all":all_workspaces,":model":embedding_model,":dim":query_embedding.len(),":profile":PROFILE, + ":src":source_language,":tgt":target_language,":query":floats_to_blob(&query_embedding),":threshold":threshold,":limit":max_results},|r|Ok(PhraseMatchResult { + phrase_memory_id:r.get(0)?,source_phrase:r.get(1)?,target_phrase:r.get(2)?,distance:r.get(3)?,confidence:r.get(4)?,workspace_id:r.get(5)?,project_id:r.get(6)?,chunk_id:r.get(7)?,source_id:r.get(8)?,source_language:r.get(10)?,target_language:r.get(11)?,source_language_variety:r.get(12)?,target_language_variety:r.get(13)?, + provenance:serde_json::from_str(&r.get::<_,String>(9)?).map_err(|e|rusqlite::Error::FromSqlConversionFailure(9,rusqlite::types::Type::Text,Box::new(e)))?,embedding_model:embedding_model.clone(),dimensions:query_embedding.len(), + })).map_err(db)?.collect::,_>>().map_err(db)?; + Ok(rows) +} diff --git a/src-tauri/src/vector/mod.rs b/src-tauri/src/vector/mod.rs index 9b454bf4..d8c74a4f 100644 --- a/src-tauri/src/vector/mod.rs +++ b/src-tauri/src/vector/mod.rs @@ -1,4 +1,10 @@ pub mod embedding; +pub mod memory_commands; +pub mod memory_search; +pub mod text_languages; +pub mod text_units; +#[cfg(test)] +mod text_units_tests; use rusqlite::{ffi::sqlite3_auto_extension, Connection, Result as RusqliteResult}; use std::path::PathBuf; @@ -50,7 +56,7 @@ pub fn open_vec_connection(db_path: &PathBuf) -> RusqliteResult { Ok(conn) } -/// The frontend owns DDL. Native vector commands only verify the columns they +/// The Rust/sqlx migrator owns DDL. Native vector commands only verify the columns they /// require so a stale or partially initialized database fails clearly instead /// of attempting an ad-hoc migration. pub fn verify_phrase_memory_schema(conn: &Connection) -> Result<(), String> { @@ -66,9 +72,9 @@ pub fn verify_phrase_memory_schema(conn: &Connection) -> Result<(), String> { if columns.is_empty() { return Err("phrase_memory schema is not initialized yet".to_string()); } - if !columns.iter().any(|column| column == "embedding_model") { + if !columns.iter().any(|column| column == "source_revision_id") { return Err( - "phrase_memory.embedding_model is missing; run the frontend schema migration first" + "phrase_memory.source_revision_id is missing; the text corpus migration has not been applied" .to_string(), ); } @@ -124,13 +130,13 @@ mod tests { } #[test] - fn rejects_a_phrase_memory_schema_without_embedding_model() -> RusqliteResult<()> { + fn rejects_a_phrase_memory_schema_without_revisions() -> RusqliteResult<()> { let conn = Connection::open_in_memory()?; conn.execute_batch("CREATE TABLE phrase_memory (id TEXT PRIMARY KEY)")?; let error = verify_phrase_memory_schema(&conn).expect_err("expected missing column error"); - assert!(error.contains("embedding_model")); + assert!(error.contains("source_revision_id")); Ok(()) } } diff --git a/src-tauri/src/vector/text_languages.rs b/src-tauri/src/vector/text_languages.rs new file mode 100644 index 00000000..fc627bd6 --- /dev/null +++ b/src-tauri/src/vector/text_languages.rs @@ -0,0 +1,182 @@ +//! Lingue dei testi salvati: codice ISO 639-3, varietà Glottolog e nota. +//! +//! Le revisioni sono immutabili: dare a una frase le lingue della sua opera vuol +//! dire creare revisioni nuove con lo stesso testo, ricopiare le misure +//! dell'originale e spostare la frase sulle revisioni nuove. Le vecchie restano. + +use super::embedding::EmbeddingError; +use super::text_units::{new_revision, record_fact}; +use rusqlite::{params, Connection, OptionalExtension}; + +/// ISO 639 «undetermined»: un testo salvato ha sempre una lingua, questa dice che non è stata indicata. +pub const UNDETERMINED_LANGUAGE: &str = "und"; + +#[derive(Clone, Debug, PartialEq, Eq)] +pub struct LanguageLabel { + pub code: String, + pub variety: Option, + pub note: String, +} + +fn db(error: rusqlite::Error) -> EmbeddingError { + EmbeddingError::Http(error.to_string()) +} + +fn label(code: Option, variety: Option, note: Option) -> LanguageLabel { + let code = code.map(|c| c.trim().to_string()).filter(|c| !c.is_empty()); + LanguageLabel { + variety: code.as_ref().and( + variety + .map(|v| v.trim().to_string()) + .filter(|v| !v.is_empty()), + ), + code: code.unwrap_or_else(|| UNDETERMINED_LANGUAGE.to_string()), + note: note.unwrap_or_default().trim().to_string(), + } +} + +/// Lingue di partenza e arrivo dell'opera; «und» dove non sono indicate. +pub fn project_languages( + conn: &Connection, + project: &str, +) -> Result<(LanguageLabel, LanguageLabel), EmbeddingError> { + conn.query_row( + "SELECT source_language, source_language_variety, source_language_note, + target_language, target_language_variety, target_language_note + FROM projects WHERE id=?1", + [project], + |r| { + Ok(( + label(r.get(0)?, r.get(1)?, r.get(2)?), + label(r.get(3)?, r.get(4)?, r.get(5)?), + )) + }, + ) + .optional() + .map_err(db)? + .ok_or_else(|| EmbeddingError::Parse("Translation not found".into())) +} + +pub fn revision_language( + conn: &Connection, + revision: &str, +) -> Result { + conn.query_row( + "SELECT language, language_variety, language_note FROM text_unit_revisions WHERE id=?1", + [revision], + |r| Ok(label(r.get(0)?, r.get(1)?, r.get(2)?)), + ) + .map_err(db) +} + +struct SavedPair { + id: String, + unit: String, + source_revision: String, + target_revision: String, +} + +/// Frasi dell'opera la cui lingua (di partenza o di arrivo) non è quella dell'opera. +fn pairs_to_relabel( + conn: &Connection, + project: &str, + source: &LanguageLabel, + target: &LanguageLabel, +) -> Result, EmbeddingError> { + let mut query = conn + .prepare( + "SELECT p.id, p.unit_id, p.source_revision_id, p.target_revision_id, + s.language, s.language_variety, s.language_note, + t.language, t.language_variety, t.language_note + FROM phrase_memory p + JOIN text_unit_revisions s ON s.id = p.source_revision_id + JOIN text_unit_revisions t ON t.id = p.target_revision_id + WHERE p.project_id=?1 ORDER BY p.created_at, p.id", + ) + .map_err(db)?; + let rows = query + .query_map([project], |r| { + Ok(( + SavedPair { + id: r.get(0)?, + unit: r.get(1)?, + source_revision: r.get(2)?, + target_revision: r.get(3)?, + }, + label(r.get(4)?, r.get(5)?, r.get(6)?), + label(r.get(7)?, r.get(8)?, r.get(9)?), + )) + }) + .map_err(db)? + .collect::, _>>() + .map_err(db)?; + Ok(rows + .into_iter() + .filter(|(_, saved_source, saved_target)| saved_source != source || saved_target != target) + .map(|(pair, _, _)| pair) + .collect()) +} + +pub fn count_relabels(conn: &Connection, project: &str) -> Result { + let (source, target) = project_languages(conn, project)?; + let count = pairs_to_relabel(conn, project, &source, &target)?.len(); + u32::try_from(count).map_err(|_| EmbeddingError::Parse("Too many phrases".into())) +} + +/// Revisione con lo stesso testo e la lingua nuova; le misure dell'originale vengono ricopiate. +fn relabel_revision( + conn: &Connection, + unit: &str, + revision: &str, + language: &LanguageLabel, + copy_embeddings: bool, +) -> Result { + if revision_language(conn, revision)? == *language { + return Ok(revision.to_string()); + } + let (role, text): (String, String) = conn + .query_row( + "SELECT role, text FROM text_unit_revisions WHERE id=?1", + [revision], + |r| Ok((r.get(0)?, r.get(1)?)), + ) + .map_err(db)?; + let relabelled = new_revision(conn, unit, &role, language, &text)?; + if copy_embeddings { + conn.execute( + "INSERT INTO text_embeddings(revision_id, provider, model, dimensions, profile, embedding) + SELECT ?1, provider, model, dimensions, profile, embedding FROM text_embeddings WHERE revision_id=?2", + params![relabelled, revision], + ) + .map_err(db)?; + } + Ok(relabelled) +} + +/// Dà alle frasi salvate dall'opera le lingue attuali dell'opera, tutte insieme o nessuna. +pub fn relabel_project(conn: &mut Connection, project: &str) -> Result { + let tx = conn.transaction().map_err(db)?; + let (source, target) = project_languages(&tx, project)?; + let pairs = pairs_to_relabel(&tx, project, &source, &target)?; + for pair in &pairs { + let source_revision = + relabel_revision(&tx, &pair.unit, &pair.source_revision, &source, true)?; + let target_revision = + relabel_revision(&tx, &pair.unit, &pair.target_revision, &target, false)?; + tx.execute( + "UPDATE phrase_memory SET source_revision_id=?1, target_revision_id=?2 WHERE id=?3", + params![source_revision, target_revision, pair.id], + ) + .map_err(db)?; + record_fact( + &tx, + "text.language.changed", + "text_unit", + &pair.unit, + None, + None, + )?; + } + tx.commit().map_err(db)?; + u32::try_from(pairs.len()).map_err(|_| EmbeddingError::Parse("Too many phrases".into())) +} diff --git a/src-tauri/src/vector/text_units.rs b/src-tauri/src/vector/text_units.rs new file mode 100644 index 00000000..a8c3e1d2 --- /dev/null +++ b/src-tauri/src/vector/text_units.rs @@ -0,0 +1,501 @@ +use super::embedding::{floats_to_blob, EmbeddingError}; +use crate::provenance::fnv1a_hex; +use rusqlite::{params, Connection, OptionalExtension, Transaction}; +use serde::{Deserialize, Serialize}; + +pub const PROFILE: &str = "source-verbatim-v1"; +const SCOPE: &str = "(:ws IS NULL OR pm.workspace_id = :ws)"; + +#[derive(Clone, Debug, Serialize, Deserialize, PartialEq)] +pub struct EmbeddingInfo { + pub provider: String, + pub model: String, + pub dimensions: usize, + pub profile: String, +} + +#[derive(Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct EmbeddingInput { + pub model: String, + pub embedding: Vec, +} + +#[derive(Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct PhrasePair { + pub source_phrase: String, + pub target_phrase: String, + pub confidence: f64, + pub source_embedding: Vec, +} + +#[derive(Debug, Serialize)] +pub struct MemoryEntry { + pub id: String, + pub unit_id: String, + pub source_revision_id: String, + pub target_revision_id: String, + pub workspace_id: Option, + pub source_phrase: String, + pub target_phrase: String, + pub confidence: f64, + pub source_language: String, + pub target_language: String, + pub source_language_variety: Option, + pub target_language_variety: Option, + pub author: Option, + pub work: Option, + pub domain: Option, + pub tags: Vec, + pub notes: Option, + pub chunk_id: Option, + pub project_id: Option, + pub source_id: Option, + pub source_version_id: Option, + pub provenance: serde_json::Value, + pub embeddings: Vec, + pub created_at: String, +} + +fn db(error: rusqlite::Error) -> EmbeddingError { + EmbeddingError::Http(error.to_string()) +} + +pub use super::text_languages::{project_languages, revision_language, LanguageLabel}; + +pub fn dimensions(model: &str) -> Result { + match model { + "text-embedding-3-small" => Ok(1536), + "text-embedding-3-large" => Ok(3072), + _ => Err(EmbeddingError::Parse("Unsupported embedding model".into())), + } +} + +pub fn validate_embedding(model: &str, vector: &[f32]) -> Result<(), EmbeddingError> { + if vector.len() != dimensions(model)? + || vector.iter().any(|v| !v.is_finite()) + || vector.iter().all(|v| *v == 0.0) + { + return Err(EmbeddingError::Parse( + "Invalid embedding dimensions or values".into(), + )); + } + Ok(()) +} + +pub fn embedding_infos( + conn: &Connection, + revision_id: &str, +) -> Result, EmbeddingError> { + let mut query = conn.prepare("SELECT provider, model, dimensions, profile FROM text_embeddings WHERE revision_id=?1 ORDER BY model, profile").map_err(db)?; + let rows = query + .query_map([revision_id], |row| { + Ok(EmbeddingInfo { + provider: row.get(0)?, + model: row.get(1)?, + dimensions: row.get(2)?, + profile: row.get(3)?, + }) + }) + .map_err(db)? + .collect::, _>>() + .map_err(db)?; + Ok(rows) +} + +pub fn put_embedding( + conn: &Connection, + revision_id: &str, + input: &EmbeddingInput, +) -> Result<(), EmbeddingError> { + validate_embedding(&input.model, &input.embedding)?; + conn.execute("INSERT INTO text_embeddings (revision_id, provider, model, dimensions, profile, embedding) VALUES (?1, 'openai', ?2, ?3, ?4, ?5) + ON CONFLICT(revision_id, provider, model, dimensions, profile) DO UPDATE SET embedding=excluded.embedding, created_at=CURRENT_TIMESTAMP", + params![revision_id, input.model, input.embedding.len(), PROFILE, floats_to_blob(&input.embedding)]).map_err(db)?; + let hash: String = conn + .query_row( + "SELECT content_hash FROM text_unit_revisions WHERE id=?1", + [revision_id], + |r| r.get(0), + ) + .map_err(db)?; + record_fact( + conn, + "text.embedding.saved", + "text_revision", + revision_id, + Some(&input.model), + Some(&hash), + )?; + Ok(()) +} + +pub(crate) fn record_fact( + conn: &Connection, + event_type: &str, + entity_type: &str, + id: &str, + model: Option<&str>, + hash: Option<&str>, +) -> Result<(), EmbeddingError> { + let key: String = conn + .query_row("SELECT lower(hex(randomblob(16)))", [], |r| r.get(0)) + .map_err(db)?; + let workspace: Option=conn.query_row("SELECT CASE WHEN pm.id IS NOT NULL THEN pm.workspace_id ELSE tu.workspace_id END FROM text_units tu + LEFT JOIN phrase_memory_entries pm ON pm.unit_id=tu.id WHERE tu.id=?1 OR tu.id=(SELECT unit_id FROM text_unit_revisions WHERE id=?1)", + [id],|r|r.get(0)).optional().map_err(db)?.flatten(); + let event = crate::provenance::Event { + event_type: event_type.into(), + entity_type: entity_type.into(), + entity_id: id.into(), + workspace_id: workspace, + actor: "user", + job_id: None, + outcome: Some("completed".into()), + duration_ms: None, + error_kind: None, + input_hash: hash.map(str::to_string), + output_hash: hash.map(str::to_string), + provider: model.map(|_| "openai".into()), + model: model.map(str::to_string), + config: Some( + serde_json::json!({"model":model,"profile":PROFILE,"entityId":id}).to_string(), + ), + key_ref: Some(key), + }; + crate::provenance::record(conn, &event).map_err(EmbeddingError::Http) +} + +pub fn new_revision( + conn: &Connection, + unit: &str, + role: &str, + language: &LanguageLabel, + text: &str, +) -> Result { + if text.trim().is_empty() || language.code.trim().is_empty() { + return Err(EmbeddingError::Parse( + "Text and language are required".into(), + )); + } + let id: String = conn + .query_row("SELECT lower(hex(randomblob(16)))", [], |r| r.get(0)) + .map_err(db)?; + conn.execute("INSERT INTO text_unit_revisions (id, unit_id, role, language, language_variety, language_note, text, content_hash, revision_number) + SELECT ?1, ?2, ?3, ?4, ?5, ?6, ?7, ?8, COALESCE(MAX(revision_number),0)+1 FROM text_unit_revisions WHERE unit_id=?2 AND role=?3 AND language=?4", + params![id, unit, role, language.code, language.variety, language.note, text, fnv1a_hex(text)]).map_err(db)?; + record_fact( + conn, + "text.revision.created", + "text_revision", + &id, + None, + Some(&fnv1a_hex(text)), + )?; + Ok(id) +} + +pub fn list( + conn: &Connection, + workspace: Option<&str>, + chunk: Option<&str>, +) -> Result, EmbeddingError> { + list_with_id(conn, workspace, chunk, None) +} + +fn list_with_id( + conn: &Connection, + workspace: Option<&str>, + chunk: Option<&str>, + id: Option<&str>, +) -> Result, EmbeddingError> { + let mut query=conn.prepare(&format!("SELECT pm.id, pm.unit_id, pm.source_revision_id, pm.target_revision_id, pm.workspace_id, + pm.source_phrase, pm.target_phrase, pm.confidence, pm.source_language, pm.target_language, pm.author, pm.work, pm.domain, + pm.notes, pm.chunk_id, pm.project_id, pm.source_id, pm.source_version_id, pm.provenance, pm.created_at, + sr.language_variety, tr.language_variety + FROM phrase_memory_entries pm + JOIN text_unit_revisions sr ON sr.id=pm.source_revision_id + JOIN text_unit_revisions tr ON tr.id=pm.target_revision_id + WHERE {SCOPE} AND (:chunk IS NULL OR pm.chunk_id=:chunk) AND (:id IS NULL OR pm.id=:id) ORDER BY pm.created_at DESC, pm.id")).map_err(db)?; + let mut entries = query + .query_map( + rusqlite::named_params! {":ws":workspace, ":chunk":chunk, ":id":id}, + |row| { + let provenance: String = row.get(18)?; + Ok(MemoryEntry { + id: row.get(0)?, + unit_id: row.get(1)?, + source_revision_id: row.get(2)?, + target_revision_id: row.get(3)?, + workspace_id: row.get(4)?, + source_phrase: row.get(5)?, + target_phrase: row.get(6)?, + confidence: row.get(7)?, + source_language: row.get(8)?, + target_language: row.get(9)?, + source_language_variety: row.get(20)?, + target_language_variety: row.get(21)?, + author: row.get(10)?, + work: row.get(11)?, + domain: row.get(12)?, + notes: row.get(13)?, + chunk_id: row.get(14)?, + project_id: row.get(15)?, + source_id: row.get(16)?, + source_version_id: row.get(17)?, + provenance: serde_json::from_str(&provenance).map_err(|e| { + rusqlite::Error::FromSqlConversionFailure( + 18, + rusqlite::types::Type::Text, + Box::new(e), + ) + })?, + tags: vec![], + embeddings: vec![], + created_at: row.get(19)?, + }) + }, + ) + .map_err(db)? + .collect::, _>>() + .map_err(db)?; + for entry in &mut entries { + entry.embeddings = embedding_infos(conn, &entry.source_revision_id)?; + let mut tags = conn + .prepare("SELECT name FROM text_unit_tags WHERE unit_id=?1 ORDER BY name") + .map_err(db)?; + entry.tags = tags + .query_map([&entry.unit_id], |r| r.get(0)) + .map_err(db)? + .collect::, _>>() + .map_err(db)?; + } + Ok(entries) +} + +pub fn get( + conn: &Connection, + workspace: Option<&str>, + id: &str, +) -> Result { + // Filter before reads of revisions/misures: changing scope cannot write an unrelated pair. + let allowed = conn + .query_row( + &format!( + "SELECT EXISTS(SELECT 1 FROM phrase_memory_entries pm WHERE pm.id=:id AND {SCOPE})" + ), + rusqlite::named_params! {":id":id, ":ws":workspace}, + |r| r.get::<_, bool>(0), + ) + .map_err(db)?; + if !allowed { + return Err(EmbeddingError::Parse( + "Memory entry is unavailable in this workspace".into(), + )); + } + list_with_id(conn, workspace, None, Some(id))? + .into_iter() + .find(|e| e.id == id) + .ok_or_else(|| EmbeddingError::Parse("Memory entry not found".into())) +} + +#[derive(Deserialize)] +#[serde(rename_all = "camelCase")] +pub struct UpdateInput { + pub workspace_id: Option, + pub phrase_memory_id: String, + pub source_revision_id: String, + pub target_revision_id: String, + pub source_phrase: String, + pub target_phrase: String, + pub embeddings: Vec, +} + +pub fn update(conn: &mut Connection, input: UpdateInput) -> Result { + let tx = conn.transaction().map_err(db)?; + let entry = get(&tx, input.workspace_id.as_deref(), &input.phrase_memory_id)?; + ensure_revisions(&entry, &input.source_revision_id, &input.target_revision_id)?; + let source = input.source_phrase.trim(); + let target = input.target_phrase.trim(); + if source.is_empty() || target.is_empty() { + return Err(EmbeddingError::Parse("Both texts are required".into())); + } + let source_id = if source != entry.source_phrase { + let mut models = entry + .embeddings + .iter() + .map(|e| e.model.as_str()) + .collect::>(); + models.sort(); + models.dedup(); + let mut supplied = input + .embeddings + .iter() + .map(|e| e.model.as_str()) + .collect::>(); + supplied.sort(); + if models != supplied { + return Err(EmbeddingError::Parse( + "All source embeddings must be recalculated".into(), + )); + } + let revision = new_revision( + &tx, + &entry.unit_id, + "source", + &revision_language(&tx, &entry.source_revision_id)?, + source, + )?; + for embedding in &input.embeddings { + put_embedding(&tx, &revision, embedding)?; + } + revision + } else { + if !input.embeddings.is_empty() { + return Err(EmbeddingError::Parse( + "Unchanged source requires no new embeddings".into(), + )); + } + entry.source_revision_id + }; + let target_id = if target != entry.target_phrase { + new_revision( + &tx, + &entry.unit_id, + "translation", + &revision_language(&tx, &entry.target_revision_id)?, + target, + )? + } else { + entry.target_revision_id + }; + tx.execute( + "UPDATE phrase_memory SET source_revision_id=?1,target_revision_id=?2 WHERE id=?3", + params![source_id, target_id, entry.id], + ) + .map_err(db)?; + tx.commit().map_err(db)?; + Ok(1) +} + +pub fn ensure_revisions( + entry: &MemoryEntry, + source: &str, + target: &str, +) -> Result<(), EmbeddingError> { + if entry.source_revision_id != source || entry.target_revision_id != target { + return Err(EmbeddingError::Parse( + "The text changed during this operation; reload before saving".into(), + )); + } + Ok(()) +} + +pub fn save_pairs( + conn: &mut Connection, + project: &str, + chunk: &str, + model: &str, + pairs: Vec, +) -> Result { + dimensions(model)?; + let tx = conn.transaction().map_err(db)?; + // Le lingue sono quelle dell'opera da cui nasce la frase, lette qui: una sola fonte. + let (source_language, target_language) = project_languages(&tx, project)?; + let (workspace,project_name,chunk_source,chunk_target,position,approved): (Option,String,String,String,Option,Option)=tx.query_row( + "SELECT p.workspace_id,p.name,t.source_processing_text,t.translation_processing_text,t.position,t.approved_revision_id + FROM projects p JOIN translations t ON t.project_id=p.id WHERE p.id=?1 AND t.id=?2 AND t.translation_locked=1", + params![project,chunk],|r|Ok((r.get(0)?,r.get(1)?,r.get(2)?,r.get(3)?,r.get(4)?,r.get(5)?))).map_err(db)?; + let workspace_name: Option = if let Some(ws) = workspace.as_deref() { + tx.query_row("SELECT name FROM workspaces WHERE id=?1", [ws], |r| { + r.get(0) + }) + .optional() + .map_err(db)? + } else { + None + }; + let book: Option<(String, String, String, String)> = tx + .query_row( + "SELECT s.id,v.id,s.title,v.label FROM translation_origins o + LEFT JOIN transcription_documents d ON d.id=o.transcription_document_id + JOIN source_versions v ON v.id=COALESCE(o.source_version_id,d.source_version_id) + JOIN sources s ON s.id=v.source_id WHERE o.project_id=?1", + [project], + |r| Ok((r.get(0)?, r.get(1)?, r.get(2)?, r.get(3)?)), + ) + .optional() + .map_err(db)?; + let mut added = 0; + for pair in pairs { + let source = pair.source_phrase.trim(); + let target = pair.target_phrase.trim(); + if source.is_empty() || target.is_empty() || !pair.confidence.is_finite() { + return Err(EmbeddingError::Parse("Invalid text pair".into())); + } + validate_embedding(model, &pair.source_embedding)?; + let existing: Option<(String,String)>=tx.query_row("SELECT id,source_revision_id FROM phrase_memory_entries WHERE project_id=?1 AND chunk_id=?2 AND source_phrase=?3 AND target_phrase=?4 AND source_language=?5 AND target_language=?6", + params![project,chunk,source,target,source_language.code,target_language.code],|r|Ok((r.get(0)?,r.get(1)?))).optional().map_err(db)?; + if let Some((_, revision)) = existing { + put_embedding( + &tx, + &revision, + &EmbeddingInput { + model: model.into(), + embedding: pair.source_embedding, + }, + )?; + continue; + } + let unit: String = tx + .query_row("SELECT lower(hex(randomblob(16)))", [], |r| r.get(0)) + .map_err(db)?; + let position_start = if chunk_source.match_indices(source).count() == 1 { + chunk_source + .find(source) + .map(|start| chunk_source[..start].chars().count()) + } else { + None + }; + let provenance = serde_json::json!({"projectId":project,"projectName":project_name,"chunkId":chunk,"chunkPosition":position, + "workspaceId":workspace,"workspaceName":workspace_name,"approvedTranslationRevisionId":approved,"sourceHash":fnv1a_hex(&chunk_source),"targetHash":fnv1a_hex(&chunk_target), + "sourceId":book.as_ref().map(|b|&b.0),"sourceVersionId":book.as_ref().map(|b|&b.1), + "sourceTitle":book.as_ref().map(|b|&b.2),"sourceVersionLabel":book.as_ref().map(|b|&b.3),"selection":{"exact":source,"start":position_start,"end":position_start.map(|start|start+source.chars().count())}}); + tx.execute("INSERT INTO text_units(id,source_id,source_version_id,workspace_id,provenance) VALUES(?1,?2,?3,?4,?5)", + params![unit,book.as_ref().map(|b|&b.0),book.as_ref().map(|b|&b.1),workspace,provenance.to_string()]).map_err(db)?; + let sr = new_revision(&tx, &unit, "source", &source_language, source)?; + let tr = new_revision(&tx, &unit, "translation", &target_language, target)?; + tx.execute("INSERT INTO phrase_memory(id,unit_id,source_revision_id,target_revision_id,project_id,chunk_id,confidence,work) VALUES(?1,?1,?2,?3,?4,?5,?6,?7)", + params![unit,sr,tr,project,chunk,pair.confidence.clamp(0.0,1.0),book.as_ref().map(|b|&b.2)]).map_err(db)?; + put_embedding( + &tx, + &sr, + &EmbeddingInput { + model: model.into(), + embedding: pair.source_embedding, + }, + )?; + added += 1; + } + tx.commit().map_err(db)?; + Ok(added) +} + +pub fn set_tags(tx: &Transaction<'_>, unit: &str, tags: &[String]) -> Result<(), EmbeddingError> { + tx.execute("DELETE FROM text_unit_tags WHERE unit_id=?1", [unit]) + .map_err(db)?; + for name in tags { + let name = name.trim(); + if name.is_empty() { + return Err(EmbeddingError::Parse("Empty tag".into())); + } + tx.execute( + "INSERT OR IGNORE INTO text_unit_tags(unit_id,name) VALUES(?1,?2)", + params![unit, name], + ) + .map_err(db)?; + } + record_fact(tx, "text.tags.changed", "text_unit", unit, None, None)?; + Ok(()) +} diff --git a/src-tauri/src/vector/text_units_tests.rs b/src-tauri/src/vector/text_units_tests.rs new file mode 100644 index 00000000..e091a173 --- /dev/null +++ b/src-tauri/src/vector/text_units_tests.rs @@ -0,0 +1,435 @@ +use super::memory_search::{search, SearchInput}; +use super::text_units::*; +use rusqlite::{params, Connection}; + +type TestResult = Result<(), Box>; +const SMALL: &str = "text-embedding-3-small"; +const LARGE: &str = "text-embedding-3-large"; + +#[test] +fn incremental_upgrade_preserves_texts_and_only_identified_embeddings() -> TestResult { + let conn = Connection::open_in_memory()?; + conn.execute_batch(include_str!("../../migrations/0001_baseline_2_0.sql"))?; + for (id, model) in [("known", Some(SMALL)), ("unknown", None)] { + conn.execute( + "INSERT INTO phrase_memory(id,source_phrase,target_phrase,confidence, + source_language,target_language,tags,embedding,embedding_model,created_at) + VALUES(?1,'È già: spada ⚔','A sword',0.9,'it','en','[\"arma\"]',?2,?3,CURRENT_TIMESTAMP)", + params![id, super::embedding::floats_to_blob(&vector(SMALL)), model], + )?; + } + conn.execute_batch(include_str!( + "../../migrations/0002_workspace_memory_scope.sql" + ))?; + conn.execute_batch(include_str!("../../migrations/0003_text_corpus.sql"))?; + conn.execute_batch(include_str!( + "../../migrations/0004_pipeline_work_brief.sql" + ))?; + conn.execute_batch(include_str!( + "../../migrations/0005_pipeline_shared_context.sql" + ))?; + conn.execute_batch(include_str!( + "../../migrations/0006_pipeline_prompt_composition.sql" + ))?; + conn.execute_batch(include_str!("../../migrations/0007_work_languages.sql"))?; + conn.execute_batch(include_str!( + "../../migrations/0008_text_language_variety.sql" + ))?; + let entries = list(&conn, None, None)?; + assert_eq!(entries.len(), 2); + for entry in &entries { + assert_eq!(entry.source_phrase, "È già: spada ⚔"); + assert_eq!(entry.target_phrase, "A sword"); + assert_eq!(entry.tags, vec!["arma"]); + assert_eq!(entry.embeddings.len(), usize::from(entry.id == "known")); + let hash: String = conn.query_row( + "SELECT content_hash FROM text_unit_revisions WHERE id=?1", + [&entry.source_revision_id], + |row| row.get(0), + )?; + assert_eq!(hash, crate::provenance::fnv1a_hex(&entry.source_phrase)); + } + assert!(!conn.prepare("PRAGMA foreign_key_check")?.exists([])?); + Ok(()) +} + +fn connection() -> Result> { + super::register_vec_extension(); + let conn = Connection::open_in_memory()?; + conn.execute_batch(include_str!("../../migrations/0001_baseline_2_0.sql"))?; + conn.execute_batch(include_str!( + "../../migrations/0002_workspace_memory_scope.sql" + ))?; + conn.execute_batch(include_str!("../../migrations/0003_text_corpus.sql"))?; + conn.execute_batch(include_str!( + "../../migrations/0004_pipeline_work_brief.sql" + ))?; + conn.execute_batch(include_str!( + "../../migrations/0005_pipeline_shared_context.sql" + ))?; + conn.execute_batch(include_str!( + "../../migrations/0006_pipeline_prompt_composition.sql" + ))?; + conn.execute_batch(include_str!("../../migrations/0007_work_languages.sql"))?; + conn.execute_batch(include_str!( + "../../migrations/0008_text_language_variety.sql" + ))?; + conn.execute_batch("INSERT INTO workspaces(id,name,created_at) VALUES('ws-a','Archivio',CURRENT_TIMESTAMP),('ws-b','Studio',CURRENT_TIMESTAMP); + INSERT INTO projects(id,name,workspace_id,source_language,target_language) VALUES('project','Traduzione','ws-a','la','en'); + INSERT INTO translations(id,project_id,source_processing_text,translation_processing_text,position,translation_locked) + VALUES('chunk','project','Salve amice','Hello friend',2,1); + INSERT INTO sources(id,title,kind) VALUES('book','Libro medievale','manuscript'); + INSERT INTO source_versions(id,source_id,label,version_kind) VALUES('version','book','Testimone A','edition'); + INSERT INTO translation_origins(project_id,origin_type,source_version_id) VALUES('project','source_level','version');")?; + Ok(conn) +} + +fn vector(model: &str) -> Vec { + vec![1.0; if model == SMALL { 1536 } else { 3072 }] +} +fn save(conn: &mut Connection, model: &str) -> Result { + save_pairs( + conn, + "project", + "chunk", + model, + vec![PhrasePair { + source_phrase: "Salve".into(), + target_phrase: "Hello".into(), + confidence: 0.9, + source_embedding: vector(model), + }], + ) +} +fn change(entry: &MemoryEntry, source: &str, target: &str, models: &[&str]) -> UpdateInput { + UpdateInput { + workspace_id: entry.workspace_id.clone(), + phrase_memory_id: entry.id.clone(), + source_revision_id: entry.source_revision_id.clone(), + target_revision_id: entry.target_revision_id.clone(), + source_phrase: source.into(), + target_phrase: target.into(), + embeddings: models + .iter() + .map(|model| EmbeddingInput { + model: (*model).into(), + embedding: vector(model), + }) + .collect(), + } +} +fn find( + conn: &Connection, + model: &str, + workspace: &str, + global: bool, + target: &str, +) -> Result, super::embedding::EmbeddingError> { + search( + conn, + SearchInput { + workspace_id: workspace.into(), + query_embedding: vector(model), + threshold: 0.5, + max_results: 10, + embedding_model: model.into(), + all_workspaces: global, + source_language: Some("la".into()), + target_language: Some(target.into()), + }, + ) +} + +#[test] +fn adding_and_recalculating_models_preserves_unit_and_other_embeddings() -> TestResult { + let mut conn = connection()?; + assert_eq!(save(&mut conn, SMALL)?, 1); + assert_eq!(save(&mut conn, LARGE)?, 0); + assert_eq!(save(&mut conn, SMALL)?, 0); + let entries = list(&conn, None, None)?; + assert_eq!(entries.len(), 1); + assert_eq!(entries[0].embeddings.len(), 2); + assert_eq!(entries[0].source_id.as_deref(), Some("book")); + assert_eq!(entries[0].source_version_id.as_deref(), Some("version")); + assert_eq!(entries[0].provenance["workspaceName"], "Archivio"); + assert_eq!(entries[0].provenance["selection"]["start"], 0); + assert_eq!(find(&conn, SMALL, "ws-a", false, "en")?.len(), 1); + assert_eq!(find(&conn, LARGE, "ws-a", false, "en")?.len(), 1); + let facts:i64=conn.query_row("SELECT count(*) FROM provenance_events WHERE event_type='text.embedding.saved' AND provider='openai' AND model IS NOT NULL",[],|row|row.get(0))?; + assert_eq!(facts, 3); + Ok(()) +} + +#[test] +fn compatible_search_excludes_other_models_profiles_languages_and_workspaces() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + let entry = list(&conn, None, None)?.remove(0); + assert!(find(&conn, LARGE, "ws-a", false, "en")?.is_empty()); + assert!(find(&conn, SMALL, "ws-a", false, "it")?.is_empty()); + assert!(find(&conn, SMALL, "ws-b", false, "en")?.is_empty()); + assert_eq!(find(&conn, SMALL, "ws-b", true, "en")?.len(), 1); + // Even equal-size vectors from another input profile must never be compared. + conn.execute( + "UPDATE text_embeddings SET profile='normalized-v1' WHERE revision_id=?1", + [&entry.source_revision_id], + )?; + assert!(find(&conn, SMALL, "ws-a", false, "en")?.is_empty()); + conn.execute( + "UPDATE text_embeddings SET profile=?1,model='different-model'", + [PROFILE], + )?; + assert!(find(&conn, SMALL, "ws-a", false, "en")?.is_empty()); + Ok(()) +} + +#[test] +fn target_edits_preserve_source_revision_and_vectors_and_source_edits_archive_all() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + save(&mut conn, LARGE)?; + let original = list(&conn, None, None)?.remove(0); + update(&mut conn, change(&original, "Salve", "Greetings", &[]))?; + let translated = get(&conn, None, &original.id)?; + assert_eq!(translated.source_revision_id, original.source_revision_id); + assert_ne!(translated.target_revision_id, original.target_revision_id); + assert_eq!(translated.embeddings.len(), 2); + update( + &mut conn, + change(&translated, "Salve amice", "Greetings", &[SMALL, LARGE]), + )?; + let revised = get(&conn, None, &original.id)?; + assert_ne!(revised.source_revision_id, original.source_revision_id); + assert_eq!(revised.embeddings.len(), 2); + assert_eq!( + embedding_infos(&conn, &original.source_revision_id)?.len(), + 2 + ); + assert_eq!( + find(&conn, SMALL, "ws-a", false, "en")?[0].source_phrase, + "Salve amice" + ); + assert!(conn + .execute( + "UPDATE text_unit_revisions SET text='changed' WHERE id=?1", + [&original.source_revision_id] + ) + .is_err()); + Ok(()) +} + +#[test] +fn incomplete_or_invalid_recalculation_rolls_back_every_revision_and_embedding() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + save(&mut conn, LARGE)?; + let original = list(&conn, None, None)?.remove(0); + assert!(update( + &mut conn, + change(&original, "New source", "Hello", &[SMALL]) + ) + .is_err()); + let mut invalid = change(&original, "New source", "Hello", &[SMALL, LARGE]); + invalid.embeddings[1].embedding = vec![1.0; 2]; + assert!(update(&mut conn, invalid).is_err()); + assert_eq!( + get(&conn, None, &original.id)?.source_revision_id, + original.source_revision_id + ); + let revisions: i64 = conn.query_row("SELECT count(*) FROM text_unit_revisions", [], |row| { + row.get(0) + })?; + assert_eq!(revisions, 2); + Ok(()) +} + +#[test] +fn stale_edits_and_wrong_workspace_writes_are_rejected() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + let entry = list(&conn, None, None)?.remove(0); + update(&mut conn, change(&entry, "Salve", "Greetings", &[]))?; + assert!(update(&mut conn, change(&entry, "Salve", "Obsolete", &[])).is_err()); + let current = get(&conn, None, &entry.id)?; + let mut wrong_scope = change(¤t, "Salve", "Other", &[]); + wrong_scope.workspace_id = Some("ws-b".into()); + assert!(update(&mut conn, wrong_scope).is_err()); + assert_eq!(get(&conn, None, &entry.id)?.target_phrase, "Greetings"); + Ok(()) +} + +#[test] +fn workspace_move_changes_scope_without_changing_historical_provenance() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + conn.execute( + "UPDATE projects SET workspace_id='ws-b' WHERE id='project'", + [], + )?; + assert!(find(&conn, SMALL, "ws-a", false, "en")?.is_empty()); + assert_eq!(find(&conn, SMALL, "ws-b", false, "en")?.len(), 1); + let entry = list(&conn, Some("ws-b"), None)?.remove(0); + assert_eq!(entry.provenance["workspaceId"], "ws-a"); + let tx = conn.transaction()?; + set_tags( + &tx, + &entry.unit_id, + &["arma".into(), "Arma".into(), "linguistica".into()], + )?; + tx.commit()?; + assert_eq!( + get(&conn, None, &entry.id)?.tags, + vec!["arma", "linguistica"] + ); + Ok(()) +} + +#[test] +fn invalid_vector_and_model_metadata_cannot_be_stored() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + let entry = list(&conn, None, None)?.remove(0); + assert!(validate_embedding(SMALL, &vec![f32::NAN; 1536]).is_err()); + assert!(validate_embedding(SMALL, &vec![0.0; 1536]).is_err()); + assert!(conn + .execute( + "INSERT INTO text_embeddings(revision_id,provider,model,dimensions,profile,embedding) + VALUES(?1,'openai','',1536,?2,?3)", + params![entry.source_revision_id, PROFILE, vec![0u8; 6144]] + ) + .is_err()); + assert!(conn + .execute( + "INSERT INTO text_embeddings(revision_id,provider,model,dimensions,profile,embedding) + VALUES(?1,'openai',?2,3072,?3,?4)", + params![entry.source_revision_id, LARGE, PROFILE, vec![0u8; 6144]] + ) + .is_err()); + Ok(()) +} + +#[test] +fn unapproved_chunk_cannot_be_archived() -> TestResult { + let mut conn = connection()?; + conn.execute("UPDATE translations SET translation_locked=0", [])?; + assert!(save(&mut conn, SMALL).is_err()); + assert!(list(&conn, None, None)?.is_empty()); + Ok(()) +} + +#[test] +fn equal_texts_from_different_chunks_remain_separate_units() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + conn.execute("INSERT INTO translations(id,project_id,source_processing_text,translation_processing_text,translation_locked) + VALUES('chunk-two','project','Salve amice','Hello friend',1)",[])?; + save_pairs( + &mut conn, + "project", + "chunk-two", + SMALL, + vec![PhrasePair { + source_phrase: "Salve".into(), + target_phrase: "Hello".into(), + confidence: 1.0, + source_embedding: vector(SMALL), + }], + )?; + assert_eq!(list(&conn, None, None)?.len(), 2); + Ok(()) +} + +#[test] +fn removing_translation_preserves_archived_texts_models_and_book_provenance() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + save(&mut conn, LARGE)?; + let original = list(&conn, None, None)?.remove(0); + let tx = conn.transaction()?; + tx.execute( + "UPDATE phrase_memory SET project_id=NULL,chunk_id=NULL WHERE project_id='project'", + [], + )?; + tx.execute("DELETE FROM projects WHERE id='project'", [])?; + tx.commit()?; + let archived = get(&conn, None, &original.id)?; + assert!(archived.workspace_id.is_none()); + assert!(archived.project_id.is_none()); + assert!(archived.chunk_id.is_none()); + assert_eq!(archived.embeddings.len(), 2); + assert_eq!(archived.provenance["sourceTitle"], "Libro medievale"); + assert_eq!(archived.provenance["projectName"], "Traduzione"); + assert_eq!(find(&conn, SMALL, "ws-b", true, "en")?.len(), 1); + assert!(find(&conn, SMALL, "ws-a", false, "en")?.is_empty()); + Ok(()) +} + +#[test] +fn relabelling_gives_saved_phrases_the_work_languages_keeping_text_and_measures() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + conn.execute( + "UPDATE projects SET source_language='lat', source_language_variety='medi1250', + source_language_note='sec. XV', target_language='ita' WHERE id='project'", + [], + )?; + assert_eq!(super::text_languages::count_relabels(&conn, "project")?, 1); + + let before = list(&conn, Some("ws-a"), None)?.remove(0); + assert_eq!( + super::text_languages::relabel_project(&mut conn, "project")?, + 1 + ); + let after = list(&conn, Some("ws-a"), None)?.remove(0); + + assert_eq!( + ( + after.source_language.as_str(), + after.target_language.as_str() + ), + ("lat", "ita") + ); + assert_eq!( + (after.source_phrase, after.target_phrase), + (before.source_phrase, before.target_phrase) + ); + assert_ne!(after.source_revision_id, before.source_revision_id); + assert_eq!(after.embeddings, before.embeddings); + let (variety, note): (Option, String) = conn.query_row( + "SELECT language_variety, language_note FROM text_unit_revisions WHERE id=?1", + [&after.source_revision_id], + |row| Ok((row.get(0)?, row.get(1)?)), + )?; + assert_eq!( + (variety.as_deref(), note.as_str()), + (Some("medi1250"), "sec. XV") + ); + // Le revisioni precedenti restano nello storico; un secondo passaggio non trova nulla. + let revisions: i64 = + conn.query_row("SELECT COUNT(*) FROM text_unit_revisions", [], |r| r.get(0))?; + assert_eq!(revisions, 4); + assert_eq!(super::text_languages::count_relabels(&conn, "project")?, 0); + Ok(()) +} + +#[test] +fn relabelling_only_the_target_keeps_the_source_revision_and_its_measures() -> TestResult { + let mut conn = connection()?; + save(&mut conn, SMALL)?; + conn.execute( + "UPDATE projects SET source_language='la', target_language='ita' WHERE id='project'", + [], + )?; + let before = list(&conn, Some("ws-a"), None)?.remove(0); + assert_eq!( + super::text_languages::relabel_project(&mut conn, "project")?, + 1 + ); + let after = list(&conn, Some("ws-a"), None)?.remove(0); + + assert_eq!(after.source_revision_id, before.source_revision_id); + assert_ne!(after.target_revision_id, before.target_revision_id); + assert_eq!(after.target_language, "ita"); + assert_eq!(after.embeddings, before.embeddings); + Ok(()) +} diff --git a/src/App.tsx b/src/App.tsx index 6c512719..95923639 100644 --- a/src/App.tsx +++ b/src/App.tsx @@ -2,7 +2,7 @@ import { lazy, Suspense, useCallback, useEffect, useRef, useState } from 'react' import { useTranslation } from 'react-i18next'; import { initLogger } from './utils/logger'; import { Header } from './components/layout'; -import { ShellNext } from './components/layout/shell-next/ShellNext'; +import { TranslationStudio } from './components/translation/TranslationStudio'; import { WorkspaceShellNext } from './components/layout/shell-next/WorkspaceShellNext'; import { AppStatusBar } from './components/layout/AppStatusBar'; import { ErrorBoundary, ConfirmDialog, PreflightDialog, RunResumeBanner, PanelTransitionVeil } from './components/common'; @@ -14,8 +14,9 @@ import { useKeyboardShortcuts } from './hooks/useKeyboardShortcuts'; import { useJobsFeed } from './hooks/useJobsFeed'; import { useRestoreFollowUp } from './hooks/useRestoreFollowUp'; import { useCloseGuard } from './hooks/useCloseGuard'; +import { useIsDarkTheme } from './hooks/useIsDarkTheme'; import { useUiStore } from './stores/uiStore'; -import type { UiFont, DocumentLineHeight, ColorScheme } from './stores/uiStore'; +import type { UiFont, DocumentLineHeight } from './stores/uiStore'; import { DOC_FONT_SIZE_CSS } from './stores/uiStore'; import { useConfigStore } from './stores/configStore'; import { useProjectStore } from './stores/projectStore'; @@ -30,7 +31,7 @@ import { TranslationsArea } from './components/workspace/TranslationsArea'; import { LibraryCatalogArea } from './components/workspace/LibraryCatalogArea'; import { TranscriptionsCatalogArea } from './components/workspace/TranscriptionsCatalogArea'; import { AnalysisArea } from './components/workspace/AnalysisArea'; -import { importTextFile } from './services/fileService'; +import { importErrorMessageKey, importTextFile, type ImportedTextFile } from './services/fileService'; import { ollamaService } from './services/llmService'; import { savePipelineConfig } from './services/pipelineService'; import { extractFootnotes } from './utils/footnoteExtractor'; @@ -43,11 +44,9 @@ import { toast } from 'sonner'; import { HL_COLORS_LIGHT, HL_COLORS_DARK } from './stores/uiStore'; function HighlightColorSync() { - const colorScheme = useUiStore((s) => s.colorScheme); + const isDark = useIsDarkTheme(); const highlightColors = useUiStore((s) => s.highlightColors); useEffect(() => { - const prefersDark = window.matchMedia('(prefers-color-scheme: dark)').matches; - const isDark = colorScheme === 'dark' || (colorScheme === 'system' && prefersDark); const fallback = isDark ? HL_COLORS_DARK : HL_COLORS_LIGHT; // Merge chiave per chiave (non solo a livello di oggetto): uno stato persistito // incompleto (chiavi mancanti da una migrazione precedente) non deve scrivere @@ -60,19 +59,17 @@ function HighlightColorSync() { root.style.setProperty('--hl-search-bg', colors.search); root.style.setProperty('--hl-audit-bg', colors.auditPhrase); root.style.setProperty('--hl-annot-bg', colors.annotation); - }, [colorScheme, highlightColors]); + }, [isDark, highlightColors]); return null; } function AccentColorSync() { - const colorScheme = useUiStore((s) => s.colorScheme); + const isDark = useIsDarkTheme(); const editorialAccentColor = useUiStore((s) => s.editorialAccentColor); useEffect(() => { - const prefersDark = window.matchMedia('(prefers-color-scheme: dark)').matches; - const isDark = colorScheme === 'dark' || (colorScheme === 'system' && prefersDark); const color = isDark ? editorialAccentColor.dark : editorialAccentColor.light; document.documentElement.style.setProperty('--color-editorial-accent', color); - }, [colorScheme, editorialAccentColor]); + }, [isDark, editorialAccentColor]); return null; } @@ -110,20 +107,10 @@ function DocTypographySync() { } function ThemeSync() { - const colorScheme = useUiStore((s) => s.colorScheme); + const isDark = useIsDarkTheme(); useEffect(() => { - const root = document.documentElement; - const mq = window.matchMedia('(prefers-color-scheme: dark)'); - const apply = (scheme: ColorScheme, prefersDark: boolean) => { - const dark = scheme === 'dark' || (scheme === 'system' && prefersDark); - root.classList.toggle('dark', dark); - }; - apply(colorScheme, mq.matches); - if (colorScheme !== 'system') return; - const handler = (e: MediaQueryListEvent) => apply('system', e.matches); - mq.addEventListener('change', handler); - return () => mq.removeEventListener('change', handler); - }, [colorScheme]); + document.documentElement.classList.toggle('dark', isDark); + }, [isDark]); return null; } @@ -193,7 +180,6 @@ function EditorView() { const { t } = useTranslation(); const { runPipeline, - runAuditOnly, runSingleChunk, auditSingleChunk, runCoherenceAudit, @@ -222,36 +208,53 @@ function EditorView() { if (showLibraryPanel) libraryPanelLoaded.current = true; const [pendingImport, setPendingImport] = useState(null); + const leaveProject = useProjectStore((state) => state.leaveProject); + const updateWorkLanguages = useProjectStore((state) => state.updateWorkLanguages); + const navigate = useUiStore((state) => state.navigate); + const leaveTranslation = useCallback(async () => { + if (await leaveProject()) navigate({ area: 'translations' }); + }, [leaveProject, navigate]); const editorContentKey = `editor-panel-${currentProjectId ?? 'none'}`; + /** Apre l'anteprima dell'import con un file già letto: dal comando + * dell'editor o dalla finestra che crea una traduzione con il suo file. */ + const startImport = useCallback((imported: ImportedTextFile) => { + const isMarkdown = imported.format === 'markdown'; + const cleanText = isMarkdown ? extractFootnotes(imported.text).cleanText : imported.text; + setPendingImport({ + fileName: imported.name, + text: cleanText, + rawText: imported.text, + useChunking: config.useChunking !== false, + wordsPerChunk: config.wordsPerChunk ?? chunkPresetMedium, + headingAware: config.headingAware ?? true, + carryTrailingShortBlocks: config.carryTrailingShortBlocks ?? true, + format: imported.format, + experimental: imported.experimental, + }); + }, [chunkPresetMedium, config]); + const handleImportDocument = useCallback(async () => { try { const imported = await importTextFile(); - if (!imported) return; - const isMarkdown = imported.format === 'markdown'; - const cleanText = isMarkdown ? extractFootnotes(imported.text).cleanText : imported.text; - setPendingImport({ - fileName: imported.name, - text: cleanText, - rawText: imported.text, - useChunking: config.useChunking !== false, - wordsPerChunk: config.wordsPerChunk ?? chunkPresetMedium, - headingAware: config.headingAware ?? true, - carryTrailingShortBlocks: config.carryTrailingShortBlocks ?? true, - format: imported.format, - experimental: imported.experimental, - }); + if (imported) startImport(imported); } catch (err: unknown) { const msg = err instanceof Error ? err.message : String(err); - if (msg === 'pdf_no_text_layer') { - toast.error(t('files.pdfScannedError')); - } else if (msg === 'text_not_utf8') { - toast.error(t('files.textEncodingError')); - } else { - toast.error(t('files.importError'), { description: msg }); - } + const key = importErrorMessageKey(msg); + if (key === 'files.importError') toast.error(t(key), { description: msg }); + else toast.error(t(key)); } - }, [chunkPresetMedium, config, t]); + }, [startImport, t]); + + // Il file scelto nella finestra che crea la traduzione arriva qui quando + // l'editor della traduzione appena creata è montato. + const pendingImportFile = useUiStore((state) => state.pendingImportFile); + const setPendingImportFile = useUiStore((state) => state.setPendingImportFile); + useEffect(() => { + if (!pendingImportFile) return; + setPendingImportFile(null); + startImport(pendingImportFile); + }, [pendingImportFile, setPendingImportFile, startImport]); const handleConfirmImport = useCallback(async ( manualChunks?: string[], @@ -280,8 +283,6 @@ function EditorView() { : config.stages; const updatedConfig = { ...config, - sourceLanguage: pipelineConfig?.sourceLanguage ?? config.sourceLanguage, - targetLanguage: pipelineConfig?.targetLanguage ?? config.targetLanguage, stages: updatedStages, useChunking: pendingImport.useChunking, wordsPerChunk, @@ -298,6 +299,16 @@ function EditorView() { chunkedWithContextWindow: contextWindow, }; setConfig(() => updatedConfig); + if (pipelineConfig) { + try { + await updateWorkLanguages(pipelineConfig.languages); + } catch (err: unknown) { + logger.error('saveWorkLanguages after import failed', { + error: err instanceof Error ? err.message : String(err), + }); + toast.warning(t('files.languagesSaveAfterImportFailed')); + } + } loadDocument( pendingImport.rawText, { @@ -336,25 +347,23 @@ function EditorView() { setConfig, setShowConfigDrawer, t, + updateWorkLanguages, ]); return ( <>
- - +
-
+
@@ -497,15 +506,17 @@ export default function App() {
- {isShellView ? ( -
- -
- {/* Cambiando area il contenuto entra con una dissolvenza e - pochi pixel di scivolo: il salto secco fra due schermate - piene non dice se si è arrivati o se qualcosa è saltato. - Niente uscita animata — l'area vecchia se ne va subito, - tenerne due montate insieme costa letture doppie. */} + {/* La barra principale resta anche dentro una traduzione, come nelle + altre aree: da lì si cambia area senza passare dal catalogo. */} +
+ +
+ {/* Cambiando area il contenuto entra con una dissolvenza e + pochi pixel di scivolo: il salto secco fra due schermate + piene non dice se si è arrivati o se qualcosa è saltato. + Niente uscita animata — l'area vecchia se ne va subito, + tenerne due montate insieme costa letture doppie. */} + {isShellView ? ( )} -
-
-
- ) : ( -
- -
- )} + ) : ( + + )} +
+
+
{isShellView ? ( ) : null} - {/* In vista progetto la barra di stato vive dentro ShellNext (solo sotto rail+documento, - non sotto l'ispettore destro); qui resta solo per la vista workspace/home. */} - {isShellView && } + diff --git a/src/components/common/ErrorBoundary.tsx b/src/components/common/ErrorBoundary.tsx index d127c07e..698b713d 100644 --- a/src/components/common/ErrorBoundary.tsx +++ b/src/components/common/ErrorBoundary.tsx @@ -41,7 +41,7 @@ export class ErrorBoundary extends Component diff --git a/src/components/common/MarkdownEditor.tsx b/src/components/common/MarkdownEditor.tsx index 8a1bf9b0..5b245c18 100644 --- a/src/components/common/MarkdownEditor.tsx +++ b/src/components/common/MarkdownEditor.tsx @@ -494,7 +494,7 @@ export function MarkdownEditor({ // Menu esterno (shell nuova): pannello a scomparsa unico con tutti i controlli testo. // Il pulsante che lo apre vive nell'header della pagina, non qui. const textMenuPanel = ( -
+
{t('editor.viewLabel')} {modeControls} @@ -506,7 +506,7 @@ export function MarkdownEditor({ {helpButton}
{markdownEnabled ? ( -
+
{formattingControls}
) : null} @@ -524,8 +524,8 @@ export function MarkdownEditor({ ) : (
{markdownEnabled && ( @@ -539,14 +539,14 @@ export function MarkdownEditor({ {toolbarOpen ? : } )} - {markdownEnabled &&