Osteomyelitis is common, costly, and clinically well recognized, yet it remains one of the few serious infectious diseases without a drug approved specifically for its treatment or universally accepted national or international treatment guidelines.
For sponsors, that combination is both an opportunity and a trap. The opportunity is obvious: a genuine unmet need with room for a first labeled therapy. The trap is subtler, and it has derailed more than one program. Before a molecule can succeed, the trial has to be able to measure success, and in osteomyelitis, measuring success is remarkably hard.
The core difficulty is that “cure” in osteomyelitis is not a simple, self-evident concept. Unlike many acute infections, where the pathogen is cleared and the patient recovers on a predictable timeline, bone infection lacks a consensus, operational definition of cure versus remission. That single gap propagates through everything downstream: it undermines protocol design, complicates regulatory alignment, and erodes the interpretability of trial results. The consequences show up as failed or inconclusive studies, underpowered designs, and even when a drug works, difficulty securing a differentiated label and a defensible value story. Getting the endpoint right is therefore not a statistical afterthought. It is the strategic center of any osteomyelitis program.
The Endpoint Problem: Cure Versus Remission in Osteomyelitis Clinical Trials
Ask ten clinicians what “cure” means in osteomyelitis and you will get a range of answers spanning resolution of symptoms, normalization of inflammatory markers, absence of radiologic progression, negative cultures, no further surgery, and no further antibiotics. Each captures a real dimension of recovery, but none is sufficient on their own, and they do not always move together. A patient can be symptom-free while imaging still shows activity; inflammatory markers can normalize while a smoldering focus persists in devascularized bone.
For this reason, many bone-infection specialists prefer the language of remission over cure, an acknowledgment that the goal is durable, sustained control of infection rather than a binary, provable eradication event. That distinction matters enormously for trial design. If the true outcome of interest is durable remission, then a short horizon “clinical response” endpoint measured at the end of therapy is answering the wrong question and answering it optimistically.
The Late-Failure Pattern
The single most consequential feature of Osteomyelitis for endpoint design is its late failure pattern. Relapses frequently emerge well after therapy ends, and often after 60 days and continuing out to 12 months or beyond. This means that any endpoint assessed at or near the end of treatment will systematically overestimate success and obscure the relapses that matter most to patients and payers.
The published evidence bears this out. Recurrence rates in the range of 20–30% are reported even when patients receive both appropriate antibiotics and surgical debridement. A trial that declares victory at end-of-therapy will miss a large share of these events entirely.
The design tension this creates is real and unavoidable. Meaningful outcome assessment requires long follow-up, but long follow-up drives cost, lengthens timelines, and increases attrition, and attrition in a chronic, comorbid population is itself a threat to interpretability. The answer is not to shorten follow-up to make the trial convenient but to design the follow-up architecture, retention strategy, and analysis plan around the late-failure biology from the outset.
Surgical Confounding
Osteomyelitis is rarely treated with a drug alone. Aggressive debridement, resection, hardware removal, and reconstruction can independently produce good outcomes and are sometimes largely irrespective of which antibiotic the patient received. Surgical practice, moreover, varies substantially from center to center and even surgeon to surgeon.
The implication for a drug trial is uncomfortable: when surgery is doing much of the therapeutic work and doing it inconsistently, it becomes difficult to attribute an observed cure to the drug, to the surgery, or to their combination. Unmodeled surgical variability inflates noise and can either mask a genuinely effective therapy or flatter a weak one. Any credible osteomyelitis protocol must treat surgery not as background but as a co-intervention that must be standardized, documented, and accounted for in the analysis.
The Field Is Relitigating Fundamentals: What Recent Trials Are Telling Us
The good news is that the evidence base is maturing, and recent trials offer both a warning and a template.
OVIVA (Oral versus Intravenous Antibiotics for bone and joint infection; Li et al., New England Journal of Medicine, 2019) randomized 1,054 adults across 26 UK centers to oral or intravenous antibiotics for the first six weeks of definitive therapy. Its primary endpoint was definite treatment failure within one year, not a convenient end-of-therapy snapshot, and oral therapy proved non-inferior to intravenous therapy, with failure rates of roughly 13.2% versus 14.6% against a 7.5% non-inferiority margin. Two features of OVIVA are worth studying as much as its headline result. First, it chose a clinically meaningful one-year horizon, respecting the late-failure biology. Second, because the trial was open label, it relied on a blinded endpoint adjudication committee applying pre-defined microbiological, histological, and clinical criteria to classify every outcome. In other words, OVIVA’s credibility rested on exactly the endpoint infrastructure that so many earlier osteomyelitis studies lacked.
SALATIO (Short Against Long Antibiotic Therapy for Infected Orthopaedic Sites; Uçkay and colleagues) extends the question from route to duration. Its two parallel randomized trials compare six versus twelve weeks of therapy for implant-related infection and three versus six weeks for implant-free infection, with remission, clinical failure, and microbiologically identical recurrence as primary outcomes and a minimum twelve-month follow-up. The second interim analysis, published in 2026, delivered a finding that should reshape how sponsors think about stratification: diabetes and the number of surgical debridements and not antibiotic duration, emerged as the independent predictors of clinical failure, while shorter courses were associated with fewer adverse events. Formal non-inferiority has not yet been reached, owing to limited statistical power at interim, but the direction of travel is clear.
Alongside these trials, 2025 guideline activity (e.g. COA/AMMI Canada 2025 Joint Statement and SPIRIT 2025 Statement, Updated Guideline for Protocols of Randomized Trials 2025) has continued to strengthen the evidence for oral therapy and earlier intravenous-to-oral switch, moving the field further from legacy dogma toward shorter, more patient-friendly regimens.
Read together, these results carry a consistent message. When outcomes are defined rigorously, adjudicated blindly, and followed long enough, they are stable and interpretable, and they reveal that patient and surgical factors often matter more than the antibiotic variable a trial was designed to test. That is simultaneously a caution and an opportunity: the programs that build rigorous, standardized endpoint frameworks into their protocols are the ones positioned to detect a true drug effect against that noisy background.
Re-Thinking Protocol and Endpoint Design in Osteomyelitis Clinical Trials
A Precise, Operational Definition of Cure
The starting point is a multi-component, prospectively defined primary endpoint: not a single surrogate, and not a definition improvised at the analysis stage. A defensible composite for osteomyelitis typically requires, at a pre-specified time point, all the following:
- Clinical resolution or clinically meaningful improvement of the index infection
- No unplanned surgical intervention for infection at the index site
- No need for additional systemic antibiotics directed at osteomyelitis
- No radiologic progression on standardized, centrally read imaging
Equally important is the follow-up window. Assessment should extend through a clinically meaningful horizon (on the order of six to twelve months) with an explicit, documented justification tied to the disease’s late-failure pattern rather than to operational convenience.
A composite endpoint is only as good as its adjudication rules, and those rules must be written before the first patient is enrolled. How are partial responses treated? How are patients who require salvage therapy classified? Are those lost to follow-up handled as failures, as censored, or through a pre-specified imputation approach, with sensitivity analyses to test the assumption? Ambiguity on any of these points is where trials quietly lose their power and their interpretability.
Aligning Primary and Key Secondary Endpoints with Stakeholders
A strong endpoint package does more than satisfy a statistician; it has to persuade regulators and speak to the people who make treatment and reimbursement decisions.
On the regulatory side, the practical challenge is balancing feasibility against robustness. One workable structure anchors the primary endpoint at a mid-term point, for example, three to six months, while carrying long-term durability, at twelve months, as a key secondary. This gives regulators a feasible primary while preserving the durability evidence that reflects the disease’s real behavior. These are precisely the design trade-offs worth raising early in scientific advice, rather than defending after the fact.
From the clinical and sponsor perspective, endpoints should map onto what actually matters in practice: freedom from repeat surgery, functional outcomes and quality of life, and the avoidance of long-term disability and limb loss. Endpoints chosen with this lens do double duty and they strengthen the regulatory dossier and, at the same time, build the value and differentiation story that supports the eventual label and market position.
Blinded Adjudication Committees: Making “Cure” and “Failure” Credible
If there is one structural lesson from OVIVA, it is that adjudication is not optional in this indication. Osteomyelitis outcomes are shot through with subjectivity. Clinical thresholds for declaring “success” vary, surgical philosophies differ, and imaging interpretation is notoriously reader dependent. Left unmanaged, that subjectivity becomes bias, and bias in an open-label or pragmatically run trial can swamp a real treatment effect.
A blinded endpoint adjudication committee addresses this directly. By applying a single standardized definition of cure and failure, reviewing clinical, surgical, microbiological, and imaging data without knowledge of treatment assignment, the committee removes site- and investigator-level variation from the primary outcome. It converts a soft, contestable judgment into a consistent, defensible one, which is exactly what regulators, and later payers, will want to see.
Standardizing Surgery and Oral-Switch Criteria Across Sites in Osteomyelitis Clinical Trials
Why Standardization is Non-Negotiable
The same variability that confounds attribution at the patient level compounds across a multicenter trial. Debridement aggressiveness, the decision to remove or retain hardware, the timing of reconstruction, and the use of local antibiotic carriers such as beads or spacers all differ by site. So do intravenous-to-oral switch decisions, which in routine practice hinge on clinician comfort, local antibiograms, and patient comorbidities.
Every one of these unstandardized choices adds variance. In aggregate, they inflate noise, reduce statistical power, and jeopardize the trial’s ability to detect a true difference between arms. The SALATIO interim finding that a number of debridements independently predicted failure is a concrete reminder that surgical variables are not nuisance parameters to be ignored; they are among the strongest drivers of the outcome.
Operationalizing Surgical and Switch Standards
The remedy for this variability is to bring the co-interventions under protocol control. That means study-specific surgical algorithms that define, as far as clinical judgment allows, when and how debridement, hardware management, and reconstruction are performed, are reinforced by documentation requirements and site training so the algorithm is followed in practice, not just on paper.
It also means pre-specified, objective criteria for the intravenous-to-oral switch, rather than leaving the timing to discretion. Those criteria can then be integrated into stratification and into endpoint interpretation, so that switch timing becomes a controlled, analyzable feature of the design rather than a hidden source of variability. Grounding these standards in the current evidence base OVIVA, SALATIO, and the 2025 guideline updates are signals to regulators and investigators that the protocol is built on contemporary data rather than legacy assumption.
Imaging Core Labs: Making MRI and PET Into Reliable Diagnostic and Outcome Tools
Imaging sits at the heart of osteomyelitis diagnosis and response assessment. MRI and PET are central both to baseline confirmation and characterization of disease and to distinguishing genuine response from persistent or recurrent infection over time. Yet imaging is also one of the least standardized elements of routine care.
An imaging core lab resolves this. By standardizing acquisition protocols across sites, implementing central and blinded reads, and harmonizing the imaging read-out with the clinical endpoint definition, the core lab turns a subjective, site-dependent modality into a reproducible outcome measure. In an indication where “radiologic progression” is one leg of the cure composite, that reproducibility is not a luxury, but it is a prerequisite for the endpoint to mean anything.
How Sponsors Can De-Risk Osteomyelitis Clinical Trials With Smarter Design
The through-line of everything above is that osteomyelitis programs succeed or fail on design decisions made early. A few principles are worth adopting from the first protocol draft:
- Define cure precisely and justify the time horizon. State explicitly what “cure” means, on what timeline, and why, anchored to the disease’s late-failure biology rather than to trial convenience.
- Decide in advance how borderline and discordant cases are adjudicated. Write the rules for partial responses, salvage therapy, and loss to follow-up before enrollment, and pre-specify sensitivity analyses for alternative cure definitions and time points.
- Build operational structures around the endpoint. A blinded adjudication committee, an imaging core lab, and standardized surgical and oral-switch criteria are not add-ons; they are what make the primary endpoint credible.
- Plan for follow-up and retention as first-class problems. In a chronic, comorbid population followed for a year, retention strategy is part of the scientific design, not an operational detail.
Where a Specialty CRO Partner Adds Value
Executing this well requires a partner fluent in the specific demands of complex infectious-disease indications, integrating medical, surgical, imaging, and statistical perspectives into a single coherent endpoint framework rather than treating them as separate workstreams. In practice, that spans complex-indication protocol design and regulatory strategy; the standing up and operational management of blinded endpoint adjudication committees and imaging core labs; and the design of standardized surgical and switch criteria that hold up across a heterogeneous site network.
Just as important is the working relationship. The programs that avoid the endpoint trap tend to engage in scientific consulting early, pressure-test their cure definition against feasibility and regulatory feedback and refine the endpoint package iteratively before committing to a pivotal design.
Conclusion
Osteomyelitis offers a rare combination of high unmet need and open regulatory territory. The barrier to capturing it is not primarily pharmacological, but methodological. Trials keep floundering on the same reef: an ambiguous definition of cure, follow-up too short for a disease that fails late, and surgical and imaging variability that drowns out the drug effect the trial was built to detect. The recent evidence, from OVIVA’s adjudicated one-year endpoint to SALATIO’s finding that patient and surgical factors outweigh antibiotic choices, points to a clear path forward. Define cure rigorously, adjudicate it blindly, standardize the co-interventions, and measure over a horizon that reflects the biology. Sponsors who solve the endpoint problem first are the ones most likely to bring the first approved osteomyelitis therapy to market, and to defend its value once they do.