The problem
A large body of Soviet and Russian pharmacology on actoprotectors and bioregulators lives in Russian full-text. Institute collections, scanned series, open repositories such as CyberLeninka. Much of it is untranslated. English often gets a thin stub: a title, an abstract fragment, a secondary review that cites the primary work without carrying it.
For these compound families the densest experimental and historical record is in that Russian trail. English secondary digests rearrange the story for a different audience. Readers who only see the digest lose the archive.
The evidence of the research programmes exists. It is scattered across full-text PDFs, institute house journals, and bibliographic islands that do not talk to each other. No single public system was ingesting that corpus, structuring it, and keeping it attached to the compounds it describes. We built one.
What sovietrx is
sovietrx is a reading-room archive and a biocomp pipeline. Visitors see compound dossiers that map literature, two research-line essays, and a papers index led by Russian sources. The pipeline grows the room: intake of Russian full-text, extraction into structured paper records, translation of missing English fields from the Russian primary, attachment of those records to compound pages.
Compound pages are archive maps. A reader can see which papers sit behind a name, in which language, and from which tradition. Models invent no findings here. Pages sell no protocols. History and research literature only. No shop. No medical advice.
Two research lines
The featured shelf follows two traditions. That is how the literature itself clustered.
Actoprotectors come from Soviet military and sports pharmacology associated with V. M. Vinogradov and the Kirov Military Medical Academy tradition. The programme treated work capacity, hypoxia, and load stress as experimental endpoints. Molecules in this line are small compounds studied for physical work under stress and low oxygen. Bemitil is an older benzimidazole reference. Bromantane sits in the adamantane class. Ademol is a later Ukrainian registration, now discontinued.
Bioregulators come from the St. Petersburg peptide school under V. Kh. Khavinson: short peptides and tissue extracts studied as organ-targeted bioregulators. Sequence-locked peptides (epitalon, thymogen, vilon, and related sequences) sit beside tissue fractions such as thymalin. The densest trail for this line lives in St. Petersburg school papers and same-circle reviews.
A first-time visitor needs only the split to start. One shelf of military-sports small-molecule literature. One shelf of St. Petersburg peptide literature. The dossiers and the Two lines essays carry the detail.
How the pipeline works
The biocomp stack is a staged path from Russian full-text to a page that is allowed to ship. Each stage has a job. Later stages may only use what earlier stages found.
| Stage | What happens |
|---|---|
| 1. Intake | CyberLeninka and mapped sources land in a corpus keyed to compounds and paper ids. |
| 2. Extract | rurx-extract-v4 turns unstructured paper text into paper JSON tied to a compound slug. |
| 3. RU↔EN | rurx-ruen-v1 fills missing English fields from the Russian primary. It leaves the science alone. |
| 4. Stopline | Dose, route, schedule, and medical-claim language are stripped or blocked before a page can ship. |
| 5. Dossier | Structured records attach to compound pages as literature maps (archive sheets). |
Intake feeds extract. Extract feeds translation. Translation feeds the stopline. The dossier is where the visitor meets the result: a compound page that points at papers.
Models and their roles
Two public product ids sit on the path. Vendor base-model names stay off this page and off reader cards. What matters here is the job each specialist performs.
| Public id | Role |
|---|---|
rurx-extract-v4 |
Paper text → structured paper JSON attachable to a compound slug |
rurx-ruen-v1 |
RU↔EN fill for missing English fields from Russian primary |
LoRA adapters and training scaffolds exist for these roles. This page publishes no eval scores and claims no production accuracy numbers we have not measured for visitors. Batch extract runs on vault GPU when a corpus pass needs them. The extract path is offline between batches. Models are instruments inside the pipeline. They invent no findings. They never override the stopline.
Ambition
We want this literature open and usable. Russian primary papers on actoprotectors and bioregulators should be readable, citable, and linked from the compounds they study. Language barriers and thin digests should stop blocking the trail back to the source.
Open access here also means structured records. Titles, abstracts, findings, provenance, compound links. Machine-readable enough for search, comparison, graphing, and audit. The pipeline extracts what the paper states, fills missing English fields from the Russian text, keeps the citation path intact, and refuses dosing or medical claims at the stopline.
If the work succeeds, a researcher can move from a compound name to the Russian experimental record behind it, in open form, with enough structure to reuse as evidence. Thin spots in the literature stay visible. The reading room grows as that open evidence layer grows.
How to read the archive
Start with the featured compounds on the home page (nine dossiers chosen as a clear entry shelf). Open a dossier for the literature map attached to that compound. Use Papers for the CyberLeninka-led index, and Two lines for short essays on actoprotectors and bioregulators. Read the trail back to the sources. Treat no page here as a protocol.