Digitize the whole archive
LiveScanned magazines, documents and out-of-print books turned into clean, searchable text — layout and Spanish diacritics preserved. Decades of print that were only images become readable by both people and the model.
Case Study — Applied AI · Cultural Archive
Rialta is an editorial house and magazine of Cuban and Latin American culture — thousands of essays, interviews, books and archival dossiers built over years by hundreds of writers. Xari built the AI layer that reads all of it: it digitizes the archive, makes it understand itself, and answers questions across the whole of it. Everything runs on local, open-weight models on Rialta's own infrastructure — so the corpus never leaves, nothing trains someone else's model, and every answer cites the piece it came from.
Built with Rialta (rialta.org) and running on Rialta's infrastructure. No archive text or reader query is sent to a third-party AI service.
The Archive
Rialta's public archive alone holds documents, authors' own recordings, dossiers, scanned magazines, videos and books — and the Magazine and Rialta Ediciones add thousands more essays and titles. This is the corpus Xari's AI layer reads, on Rialta's own machines.
Visit rialta.org →
The Archivo alone spans six collections — documents, authors' recordings, dossiers, magazines, videos and books. Every one of them is text, or becomes text, that the local models can read.
The Question
A corpus like Rialta's is its own asset and its writers' labor — decades of criticism, interviews and out-of-print books. The quick way to make it "AI-powered" is to hand it to a commercial chatbot API. But that means the archive becomes training fuel for a model Rialta will never own, and the intelligence it produces belongs to a vendor. For an independent, Spanish-language cultural publisher, that is the whole archive walking out the door. So we took the other path.
Upload the archive to a hosted model. It answers — but every essay is logged on someone else's servers, may feed someone else's training set, and the "smart archive" evaporates the day the contract, the pricing or the model changes. The value drains outward, to a company that did none of the writing.
Run open-weight models on Rialta's own machines. The archive is read where it already lives; the embeddings, index, knowledge graph and tuned weights — the part that actually learned the archive — are Rialta's property, portable and vendor-free. The value stays with the people whose work created it.
The System
Four stages, each on local models: get every piece into clean text, make the archive understand itself, let anyone ask it in plain language, and give editors tools built on top. Nothing in the pipeline calls out to a third-party AI.
Scanned magazines, documents and out-of-print books turned into clean, searchable text — layout and Spanish diacritics preserved. Decades of print that were only images become readable by both people and the model.
‘En voz del autor’, podcasts and event recordings transcribed and time-stamped, so spoken archives are searchable too. The same models read essays aloud in a local voice for listen-anywhere access.
Every piece placed against Rialta's own taxonomy — genre, theme, author, period — so the backlog is organized the way the editors already think, not by whatever metadata a CMS happened to keep.
A knowledge graph of people, works, movements and places, drawn automatically across the whole corpus: every essay that touches an author, every thread between figures and their circles — the shape of a literary history, made navigable.
An editor-reviewed abstract for each piece, and near-duplicate detection that catches reprints and versions scattered across years of the archive — so the catalog is clean and every entry earns its place.
Grounded conversational search: ask a question in plain language and get an answer that cites the exact essays it drew from — with a clear ‘not in the archive’ when the corpus doesn't cover it. No invented facts under Rialta's name.
‘If you read this, read that’ across the entire corpus, computed on-site from the text itself — turning a deep back-catalog into something readers can wander through. No reader behavior is shipped to a third party.
Tools built on the same local models: drafting in Rialta's voice (blurbs, headlines, newsletter), human-in-the-loop translation that keeps the corpus in-house, and anthology assembly that surfaces thematic threads for new Rialta Ediciones titles.
The Boundary
The archive, the models, the index and the graph all sit on Rialta's own infrastructure. A question from an editor or a reader is answered by a local model reading local text; the only thing that ever crosses to the public internet is the finished, cited answer on Rialta's site. No essay, no page, no query reaches a third-party AI — so nothing can be logged, mined, or used to train a model Rialta doesn't own.
Who Owns What It Learns
An archive like Rialta's is decades of writing by hundreds of contributors. When a model learns from it, the value that comes out shouldn't drain to a technology vendor. Three principles kept it with Rialta.
Every page is processed on Rialta's machines. Nothing is uploaded to a third-party model, and nothing is added to anyone else's training set. The corpus is used to serve Rialta — and only Rialta.
The embeddings, the index, the knowledge graph, the tuned weights — the part of the system that actually learned the archive — are Rialta's property. Portable, exportable, and free of vendor lock. Change hardware or provider, and the intelligence comes along.
Every answer cites the essay it came from and links back to it. The model points readers to the author's work; it never dissolves that work into anonymous output with no trail home.
Why Local Models
Local isn't only about privacy. It's about ownership, cost and longevity — the reasons a cultural institution should hold its own intelligence rather than rent it.
What It Unlocks
Our archive is decades of thinking by hundreds of writers — not something we were ever willing to hand to someone else's model. Xari built an AI layer that runs on our own machines: the whole archive is finally searchable, every answer points back to the original essay, and it still belongs to us.— Carlos Aníbal Alonso, Rialta
What Xari Does
Because we build the models, the data pipeline and the product on top, the AI isn't a feature bolted on from outside — it's one system, owned by the client who paid for it.
Let's talk
Thousands of documents, decades of work, scattered formats — Xari turns it into an archive that answers, on models you own and keep.
Get in touchBuilt by Xari for Rialta (rialta.org). The system runs on Rialta's own infrastructure with open-weight models; no archive content or reader query is sent to third-party AI services. Author and work names are examples from Rialta's public archive.