Cross-Lingual Information Diffusion and Temporal Alignment: Russian State Media and Polish Outlets
An empirical forensic study of whether a structured, directional transmission channel links Russian state media to monitored Polish-language outlets. Over a 52-day window (24 Jul – 14 Sep 2026), 20,010 embedding vectors and 6,487 temporal pairings are audited with exact cross-lingual cosine retrieval, sub-day latency analysis of 18 verified propagation chains, and residualized time-series cross-correlation under Monte Carlo permutation and Benjamini–Hochberg FDR control.
Long read. The complete paper — every figure, table and reference — is best read as the PDF.
Read the PDF instead ↓Peer-review working paper · Computational Social Science & Information Forensics · Snapshot 14 September 2026. ProjektLustro.eu Open Research Initiative. Primary window 24 July – 14 September 2026 (52 days). Corpus scale: 34,266 verified publications; 66,795 embedding vectors. Replication bundle released under MIT / CC-BY-4.0.
An Empirical Forensic Investigation of Semantic Replication, Transmission Latency, and Directional Synchrony in Computational Propaganda.
Structured Abstract#
Background & Motivation. Cross-border Foreign Information Manipulation and Interference (FIMI) campaigns increasingly rely on decentralized “grey” media amplifiers, linguistic proxies, and unvetted syndication networks to bypass platform countermeasures and conceal original attribution. In the Polish information environment, identifying structured translation and narrative-relay pipelines from Russian state media has historically relied on anecdotal case collections rather than calibrated statistical baselines.
Objectives. We perform an exhaustive, reproducible computational investigation into whether a structured, directional transmission channel operates between Russian state news organs and a monitored tier of Polish-language digital outlets, across three dimensions: cross-lingual semantic replication fidelity, temporal arrival latency, and macroscopic time-series pacing.
Data & Methods. We analyse 34,266 retained posts and 34,265 detection records covering Russian state media,
monitored Polish-language outlets, and an admitted baseline of six mainstream Polish control domains. Using
1024-dimensional dense multilingual embeddings (intfloat/multilingual-e5-large), we execute an exact metric
retrieval audit over 20,010 eligible vectors and 6,487 temporal pairings. Directional lead-lag relationships are
tested through 6-hour binned time-series cross-correlations residualized against diurnal harmonics, longitudinal
trends, and day-of-week fixed effects, validated against 500 Monte Carlo circular-shift permutations under
Benjamini–Hochberg False Discovery Rate control.
Principal Findings. Three empirical results establish a high-velocity relay architecture. First,
poland.news-pravda.com functions quantitatively as a dedicated localization node of Russian state apparatuses
rather than an autonomous newsroom: 77.3% of its matched articles exhibit cosine similarity ≥ 0.85 against
Russian-source texts (median 0.859), contrasted with 2.5%–22.1% for mainstream Polish control hosts, robust under
Benjamini–Hochberg adjustment (q < 0.001). Second, cross-lingual claim diffusion is rapid: across 18 verified
transmission events over 10 Russian origin chains, median lag is 4.594 hours (range 0.975–32.098 h), with 55.6%
(10/18) within 6 hours and 88.9% (16/18) within 24 hours. Third, aggregate volume on ria.ru strongly paces
monitored Polish volume (ρ = 0.954, p ≈ 0.002), whereas domestic Polish outlets show no leading cross-correlation.
Network telemetry corroborates provider convergence, including hosting of dziennik-polityczny.com on Russian
infrastructure (AS48282 / VDSINA).
Conclusions & Impact. We formalize five falsifiable hypotheses governing cross-lingual relay mechanisms, outline reproducible verification protocols, and present a four-tier cryptographic provenance framework (immutable archived snapshots, signed lineage decisions, temporal checks, and graph edges) for digital disinformation forensics.
Keywords. Computational Propaganda; Foreign Information Manipulation and Interference (FIMI); Cross-Lingual Semantic Retrieval; Multilingual Embeddings; Information Diffusion Latency; Digital Forensics; OSINT.
Primary research thesis. A quantifiable, statistically divergent, and temporally rapid information-diffusion channel links Russian state media to the Polish digital sphere across two structurally distinct operational tiers: (1) automated, high-volume localization mirrors (e.g.
poland.news-pravda.com) characterized by bulk near-duplicate translation, and (2) a secondary tier of domestic ideological selective amplifiers (e.g. Kresy.pl, wMeritum) that ingest and recontextualize specific Russian claims with a median transmission latency of 4.594 hours despite lower full-text semantic duplication than mainstream foreign wire reporting.
1. Introduction & Problem Formulation#
In contemporary computational propaganda and hybrid state competition, information operations seldom rely exclusively on direct, overt broadcasts from state-owned outlets. Western regulatory sanctions, audience skepticism toward primary state mouthpieces, and algorithmic de-amplification on major platforms have compelled foreign threat actors to evolve multi-tiered dissemination architectures. These deploy intermediate “relay belts” of automated mirror portals, partisan alternative newsrooms, and ideologically compliant aggregators that republish, translate, and recontextualize source material for targeted foreign populations (Ferrara et al., 2020; Starbird, 2019; EEAS, 2024).
In the Polish information space, identifying such transmission channels is compounded by language boundaries. A Russian-language publication on RIA Novosti, TASS, or RT must cross both a linguistic and an institutional boundary before reaching domestic readers. While journalistic inquiries and intelligence advisories have repeatedly flagged pro-Kremlin bias in selected Polish web domains (Mierzyńska, 2018; VIGINUM, 2024), empirical social computing has lacked a unified, reproducible quantitative framework that separates natural same-day coverage of legitimate breaking news from systematic, high-fidelity narrative laundering.
Research questions#
- RQ1 — Cross-lingual semantic fidelity & node typology. Do monitored Polish-language outlets exhibit similarity distributions against Russian state publications that are statistically indistinguishable from mainstream Polish baselines, or do specific nodes display near-duplicate localization rates indicative of automated syndication versus domestic editorial rewriting?
- RQ2 — Temporal transmission latency. What is the empirical latency distribution governing transit of claims from Russian origin nodes to Polish amplifiers, and does the arrival curve support spontaneous domestic discovery or high-velocity coordinated uptake?
- RQ3 — Directional macro-synchrony vs. common shocks. Do publication volumes exhibit directional lead-lag correlation, and how do we disentangle directional pacing from common exogenous geopolitical shocks?
- RQ4 — Thematic narrative clustering. Is amplified content uniformly distributed across world affairs, or does it concentrate within a finite, recurring typology of geopolitical and historical wedge narratives?
- RQ5 — Cryptographic provenance & forensic verifiability. Can transmission paths be corroborated through multi-dimensional stance coding, cryptographic content hashing, immutable web archiving, and network telemetry?
Core contributions#
- Calibrated cross-lingual retrieval baseline & control-inversion resolution. An exact, blockwise cosine retrieval audit over 20,010 eligible vectors, benchmarked against six mainstream Polish control newsrooms, eliminating subjective thresholds and establishing an odds-ratio separation over 22:1 for dedicated localization nodes.
- Empirical diffusion-latency curve. A forensic case series of 18 verified events over 10 Russian origin chains, demonstrating 55.6% of amplifications within 6 hours and 88.9% within 24 hours (median 4.594 h).
- Deseasonalized cross-correlation framework. A protocol residualizing 6-hour binned volumes against longitudinal trend, diurnal cycles, and day-of-week seasonality, tested with 500 Monte Carlo circular-shift permutations under FDR control.
- Multi-dimensional provenance standard & annotator reliability. A four-stage forensic standard combining immutable web-archive snapshots, signed lineage decisions, blinded dual-reviewer stance coding (κ = 0.84), and network infrastructure telemetry.
2. Related Work & Theoretical Foundations#
Computational propaganda & FIMI. The FIMI framing formulated by the EEAS (2024) emphasizes pattern-based, intentional, resource-backed operations. Foundational work in computational propaganda (Ferrara et al., 2016, 2020) focused on automated accounts; subsequent research revealed “information laundering” (Pacheco et al., 2021; Starbird, 2019), where deceptive claims originate in fringe state outlets, transit through proxy portals, and are injected into mainstream discourse.
Cross-lingual diffusion & embeddings. Tracking narrative diffusion across languages historically required keyword matching or lossy machine translation (Reimers & Gurevych, 2020). Dense multilingual representations — the E5 family (Wang et al., 2024) — project multi-language text into a unified space where cosine distance mirrors cross-lingual equivalence. Few prior studies calibrated dense thresholds against real adversarial operations.
Stance detection & source attribution. A critical vulnerability is conflating semantic similarity with endorsement (ALDayel & Magdy, 2021): an article reporting that “Russian state media claims X” may score high while taking a critical stance. We separate source citation, stance adoption, and truth assessment into distinct dimensions.
Investigative attribution & Portal Kombat. VIGINUM’s 2024 report detailed “Portal Kombat,” a structured
network of pro-Russian portals with localized subdomains across European languages. We independently corroborate and
extend these findings with article-level semantic and temporal telemetry for the Polish node
(poland.news-pravda.com).
3. Dataset Architecture & Sampling Frame#
The corpus derives from LUSTRO, an operational PostgreSQL research store monitoring digital news, political commentary
and social channels across Central and Eastern Europe. It maintains immutable scraped contents, token hashes,
normalized timestamps, entity tokens, and dense vectors. Publications are classified by a rule-governed taxonomy:
Russian origin nodes (ru_origin: RIA, TASS, RT, Sputnik), monitored Polish-language outlets (polish_monitored),
a mainstream Polish control baseline (mainstream_control: Onet, RMF24, Interia, Gazeta, TVN24, Wprost), and a
residual tier.

Sampling frame & ethics. The corpus is an observational convenience sample conditioned on prior regulatory and investigative concern; absolute counts do not represent population-level prevalence across the Polish internet. All data are stored in a restricted, read-only PostgreSQL instance. In compliance with privacy-by-design (GDPR Art. 89), user comments and private author identities are excluded from automated extraction; records are indexed via SHA-256 hash digests.
4. Quantitative & Forensic Methodology#
Dense cross-lingual representation. Article bodies and headlines are embedded with intfloat/multilingual-e5-large
(24-layer transformer, output dimension D = 1024), normalized, prefixed with "passage: ", truncated to 512 subword
tokens, and L2-normalized. Cross-lingual similarity between an origin article and a candidate Polish article is their
inner product (cosine similarity). Parallel evaluation across 32,602 Polish-specific ModernBERT vectors
(PKOBP/embed-modernbert-395m) corroborates the E5 separation on Polish-Polish cohesion.
Exact metric retrieval audit. To avoid ANN quantization artifacts, a blockwise exact matrix multiplier runs over all eligible vectors (active model version, verified unit norm, valid timestamp, strict temporal precedence). For the 52-day window, exact cross-evaluation covered 20,010 eligible vectors, generating 6,487 valid temporal pairings. The near-duplicate threshold is frozen at τ = 0.85: random cross-lingual pairs score 0.60–0.75, same-event independent wording 0.78–0.83, and scores above 0.85 indicate direct paraphrase, machine translation, or verbatim replication.
Time-series residualization & permutations. Volumes are aggregated into uniform 6-hour bins (T = 210 across the window) and residualized against a secular trend, hour-of-day, and day-of-week fixed effects. Cross-correlation is computed on the residualized series and evaluated against B = 500 circular shifts, with Benjamini–Hochberg FDR adjustment across hosts.

Multi-dimensional stance & provenance framework. Each propagation path is decomposed across five orthogonal dimensions: claim relationship, source relation, stance toward claim, publication form, and factuality evaluation. Annotations were executed under a blinded dual-reviewer protocol without access to model scores. Inter-rater reliability (Cohen’s κ) reached 0.84 (95% CI [0.74, 0.94]) for claim relationship and 0.79 (95% CI [0.68, 0.90]) for stance adoption; discrepancies were resolved by a third senior reviewer.
5. Empirical Findings & Quantitative Proof#
All quantitative extractions were recomputed read-only from the live database on 14 September 2026 under strict
transaction isolation (manifest run run-20260914-1302-live; script SHA-256 74669e68…457b; results SHA-256
abbe91d1…c690).
Finding 1 — Extreme asymmetry in semantic replication (F1)#

For mainstream newsrooms, elevated similarity (0.80–0.84) reflects professional foreign-desk reporting on
high-visibility wire stories; their near-duplicate share stays bounded between 2.5% and 22.1% (pooled baseline 13.3%,
101/759). By contrast, poland.news-pravda.com maintains a median similarity of 0.859 across 3,688 articles with 77.3%
(2,852/3,688) meeting the near-duplicate criterion — a level unattainable through independent journalism.
Resolution of the control-inversion paradox. Mainstream controls (Interia 22.1% ≥ 0.85, Onet 21.2%) exhibit higher near-duplicate rates than domestic alternative portals (wMeritum 13.2%, Kresy.pl 11.6%, Zmiany na Ziemi 8.4%), and every monitored domestic outlet yields an adjusted q of 1.000 on cosine alone. This resolves once production practice is examined: mainstream desks translate agency wire dispatches nearly verbatim (high cosine), while domestic partisan outlets selectively cherry-pick and rewrite specific claims with partisan framing, depressing whole-document similarity into the 0.78–0.83 band. Dense similarity is therefore a definitive detector for automated mirrors, but is insufficient on its own to identify domestic amplifiers — those require claim-level provenance and latency tracing.
Finding 2 — High-velocity sub-day diffusion latency (F2)#
We evaluated a curated manifest of 18 verified Polish observations from 10 distinct Russian origin chains (5 TASS, 4 RT, 1 RIA Novosti), with origin and amplifier timestamps verified against cryptographic web-archive captures. The 18 observations measure latency conditioned on successful cross-lingual penetration — a high-confidence forensic sample of minimum arrival boundaries, not a population-wide transmission probability.

The arrival distribution exhibits pronounced positive skew, characteristic of rapid viral diffusion. The observed minimum is 0.975 hours (58 minutes) for a military drone claim replicated by Pravda PL from TASS; the maximum is 32.098 hours for an elaborate rewrite in News Front. In all 18 chains Δt > 0 — exactly zero negative lags, confirming that Russian media consistently preceded Polish publication for these claims.


Finding 3 — Macroscopic volume synchrony & directional pacing (F3)#
Across the 210-bin (52-day) series, publication counts were residualized against secular trend, hour of day, and day
of week, then tested against 500 circular-shift permutations under Benjamini–Hochberg correction. The Russian flagship
ria.ru exhibits a dominant, statistically significant positive cross-correlation with the aggregate monitored Polish
series (ρ = 0.954, empirical permutation p = 0.001996), with zero-lag and positive-lag dominance: surges in RIA output
pace corresponding surges across the Polish ecosystem. Domestic outlets show no pacing — e.g. dziennik-polityczny.com
yields ρ = 0.162 (p = 0.319, adjusted q = 1.000). Domestic portals do not lead the network; they respond to incoming
waves.

Finding 4 — Narrative clustering & thematic typology (F4)#
Assigned cluster centroids reveal that amplified claims are not uniformly distributed: 86.4% of cross-lingually amplified volume concentrates into six discrete narrative families — (1) anti-Ukrainian historical grievances (Wołyń / UPA); (2) institutional defense corruption & Western aid diversion; (3) refugee criminality & social friction; (4) Polish–Ukrainian border & economic friction; (5) Western diplomatic coercion & partition rumors; and (6) sensationalist engagement bait (cosmic & energy phenomena).
Finding 5 — Network infrastructure & cryptographic provenance (F5)#
Text metrics are corroborated by external network forensics: poland.news-pravda.com is cataloged in the French
SGDSN/VIGINUM domain registry (row dated 06/11/2024, immutable commit SHA-256 3ddeb417…76d89), identified as a
coordinated Russian operation administered by a Crimea-based front company. DNS A-record queries on 11 September 2026
show dziennik-polityczny.com resolving to 94.103.86.245, registered under RIPE NCC to VDSINA-NET in the Russian
Federation (AS48282). Ten complete origin→amplifier paths are verified with immutable HTML snapshots, SHA-256 hashes,
signed lineage reviews, and accepted graph edges.

6. Formal Scientific Hypotheses for Peer Verification#
- H1 (Localization architecture) — confirmed (q < 0.001).
poland.news-pravda.comoperates as a dedicated, automated localization node exhibiting near-duplicate replication at rates significantly exceeding independent controls. Fisher’s exact odds ratio 22.22 (p < 10⁻¹⁵, BH q < 0.001); the null is decisively rejected. - H2 (Ideological selective amplification) — supported. A secondary tier of domestic Polish alternative portals selectively absorbs and amplifies Russian claims through editorial rewriting, achieving latencies under six hours despite lower full-text duplication than mainstream wire reporting. Median latency 4.594 h; 10/18 arrive ≤ 6 h.
- H3 (Macro-volume alignment & directional pacing) — confirmed (p < 0.005). RIA Novosti paces monitored Polish volume at non-negative lags (ρ = 0.954, k ≥ 0), whereas domestic outlets show no leading correlation (ρ ≤ 0.162, q = 1.000).
- H4 (Infrastructure persistence & evasion) — corroborated by telemetry.
dziennik-polityczny.comrelies on Russian-domiciled hosting (AS48282 / VDSINA; DNS → 94.103.86.245; BGP → AS48282; RIPE RDAP corroboration), reflecting takedown evasion beyond conventional European hosting. - H5 (Thematic selectivity) — confirmed (86.4% concentration). Amplification selectively targets a finite taxonomy of six narrative families focused on undermining Polish–Ukrainian relations and NATO cohesion. Chi-square goodness-of-fit χ² = 142.6 (p < 10⁻¹²); the null is rejected.
7. Mechanistic Discussion & Policy Implications#
The Polish pro-Kremlin ecosystem is not monolithic. Tier 1 — dedicated localization nodes (Pravda PL): extreme volume, automated scraping, direct machine translation, near-duplicate similarity (≥ 0.85 for 77.3% of output). Tier 2 — ideological relay outlets (Kresy.pl, Myśl Polska): curated, rewritten claims framed within domestic debates; moderate similarity (0.80–0.84) and rapid latencies (2–6 h). Tier 3 — commercial engagement / astroturfing portals (wMeritum, Zmiany na Ziemi): ad-driven republication with superficial distancing while propagating the core narrative.
Information laundering & editorial distancing. Domestic portals often cite a Russian source explicitly (“jak podaje TASS…”) in neutral or skeptical language, yet the claim’s reach is achieved: an unverified allegation is injected into Polish discourse within hours. Moderation systems relying on sentiment or keyword blacklists fail here because the framing appears formally journalistic.
Policy implications. (1) Cross-lingual real-time indexing by national CERTs/CSIRTs to detect new proxy subdomains within 24 hours; (2) infrastructure-level KYC and sanctions compliance on hosting resellers such as AS48282 / VDSINA; (3) algorithmic de-amplification of domains with high near-duplicate syndication rates against state-attributed networks.
8. Threats to Validity & Limitations#
- Convenience-sampling bias. The monitored tier was selected on prior scrutiny; measured rates cannot be extrapolated as baselines across all Polish media.
- Confounding by global breaking news. Major events generate simultaneous independent coverage; isolated claims still require multi-annotator lineage verification to exclude common-cause reporting.
- Control overlap window. The control baseline overlaps for eight continuous days within the window; a 90-day continuous control baseline is committed for future iterations.
- Transformer context truncation. E5 truncates at 512 subword tokens; late-essay commentary may be underrepresented.
- Semantic similarity vs. causal intent. High cosine proves lexical/semantic correspondence, not legal coordination, remuneration, or subjective intent — which require external OSINT and financial forensics.
9. Reproducibility & Open Science Protocols#
The live database run is committed with cryptographic script and results hashes (Section 5). The pipeline is deterministically re-executable:
# Re-run quantitative verification and consistency audit
env/bin/python scripts/research/regenerate_paper.py --check
# Execute exact retrieval and time-series cross-correlation runner
env/bin/python scripts/research/live_analysis.py --manifest data/manifests/reviewed_propagation_chains.yaml
The 10 accepted origin→amplifier paths and the evidence_ready_v2 pilot cluster are backed by immutable WARC/HTML
snapshots and digitally signed reviewer decisions in the project’s authorized research repository.
References & Evidence Sources#
- ALDayel, A., & Magdy, W. (2021). Stance detection on social media: State of the art and trends. Information Processing & Management, 58(4), 102597.
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate. J. R. Stat. Soc. B, 57(1), 289–300.
- European External Action Service (EEAS). (2024). 4th EEAS Annual Report on FIMI Threats. Brussels.
- EUvsDisinfo. (2024). The Kremlin’s Disaster Disinformation Exploits in Poland. East StratCom Task Force.
- Ferrara, E., Varol, O., Davis, C., Menczer, F., & Flammini, A. (2016). The rise of social bots. Comm. ACM, 59(7), 96–104.
- Ferrara, E., Cresci, S., & Luceri, L. (2020). Misinformation, manipulation, and weapons of mass deception. IEEE Access, 8, 208889–208899.
- Freelon, D., & Wells, C. (2020). Disinformation as a form of social engineering. Political Communication, 37(2), 145–156.
- LUSTRO. (2026a). Curated Propagation Chains Manifest.
data/manifests/reviewed_propagation_chains.yaml. Commit2bf56c4. - LUSTRO. (2026b). White Paper: Evidence and Pareto Improvement Plan. Internal Research Directive, 11 September 2026.
- LUSTRO. (2026c). Attributed Origin Domains Registry.
data/manifests/attributed_origin_domains.yaml. Commit2bf56c4. - Mierzyńska, A. (2018). Rosyjską propagandę szerzą polskie portale. OKO.press.
- Miranda, S., Artetxe, M., & Zubiaga, A. (2018). Multilingual cross-document coreference resolution. EMNLP, 2412–2422.
- Pacheco, D., et al. (2021). Uncovering coordinated networks on social media. ICWSM, 15(1), 455–466.
- Reimers, N., & Gurevych, I. (2020). Making monolingual sentence embeddings multilingual. EMNLP, 4512–4525.
- RIPE NCC. (2026). RDAP Registration Data for 94.103.86.245 (AS48282, VDSINA-NET). Observed 11 September 2026.
- SGDSN / VIGINUM. (2024). Portal Kombat: A Structured and Coordinated Pro-Russian Propaganda Network. Paris. Registry commit SHA-256
3ddeb417…76d89. - Starbird, K. (2019). Disinformation’s spread: Bots, trolls and all of us. Nature, 571(7766), 449.
- Wang, L., et al. (2024). Multilingual E5: Text embeddings for 100+ languages. arXiv:2402.05672.
ProjektLustro.eu · Open Science Research Monograph v2.1 · Licensed CC BY 4.0 · Production verification hash abbe91d1…c690. This on-page reading version reproduces the paper’s structure and figures; the linked PDF is the canonical artifact.