2026 - 2026 · Biomedical Informatics

Release Sensitivity in Biomedical Knowledgebases

Researcher[s]: Guan, T.; Feng, R.; Li, W.‡,⊤; Guan, L.‡,⊤; Li, C.‡,⊤; Liu, F.‡,⊤; Equal Contribution No Precedence

About

This project examines how analytical conclusions change when the same biomedical knowledgebase is updated across releases. Using Open Targets Platform releases 25.12, 26.03, and 26.06, the study separates four distinct forms of release sensitivity: score revision among persistent disease–target pairs, reranking on fixed target rosters, candidate turnover on release-native rosters, and changes associated with entity mapping or metadata evolution. The analysis includes 2,313,492 disease–target pairs persistent across all three releases and 8,491 and 7,542 eligible fixed disease panels for the two adjacent transitions. Release sensitivity is evaluated using score-change distributions, Kendall tau-b, tie-aware leading-target and top-k membership, Jaccard overlap, leader margins, tolerance analyses, stable-entity controls, matched-support comparisons, genetic-association endpoints, and therapeutic-area stratification. The overall finding is that release identity can materially affect downstream biomedical rankings even when aggregate agreement remains high: leading-target sets changed in 6.21% and 17.38% of fixed panels across the two transitions, with most changes reflecting replacement of the unique leading target rather than tie-related variation.

Keywords

Open Targets Platform; biomedical knowledgebases; release sensitivity; database versioning; disease–target associations; target prioritisation; ranking stability; score revision; candidate turnover; entity mapping; biomedical informatics; reproducibility; data provenance; data reuse; computational robustness;

Study Design

Cross-release analyses are separated into support-specific estimands before any metric is calculated. Three-release fixed support is used to quantify score revision among identical persistent disease–target pairs; adjacent-release fixed rosters are used to evaluate reranking without candidate turnover; release-native rosters are used to quantify entrants, exits, and top-set turnover; and mapping and entity controls are used to distinguish identifier or metadata evolution from changes among stable entities. Disease harmonisation follows explicit ontology and cross-reference rules, while targets are harmonised by exact Ensembl gene identifiers. Primary fixed ranking panels require at least 30 eligible targets and at least two distinct scores in each release. Exact ties are retained using average ranks and tie-aware top-set definitions. Robustness analyses include score tolerances, leader margins, alternative panel-size thresholds, staged entity-stability restrictions, a broader disease-universe sensitivity, independently constructed genetic-association supports, and matched disease–target rosters. Validation is fail-closed and includes schema, support, digest, configuration, and acceptance checks, repeated clean builds, independent rank and set recomputation, and source-traced case verification.

Data and Code Availability

The public version 2.0.0 derived-data record is archived on Zenodo at doi:10.5281/zenodo.22110813 and contains harmonised associations, entity crosswalks, support membership, score and ranking records, robustness outputs, source-traced cases, manifests, dictionaries, and SHA-256 checksums. The version 2.0.0 source-to-record workflow, validation software, tests, and figure-generation code are publicly available through the OpenTargets-Release-Sensitivity GitHub repository and archived on Zenodo at doi:10.5281/zenodo.22121355. Raw Open Targets Platform source objects are not redistributed; release-specific source identities, formats, licences, row counts, and verified digests are retained in the public source manifest.

← Back to Research