Part 3. Nova's modeling and simulation approach
TL;DR
The clinical outcome is the ultimate focus of the M&S process.
The process has six steps:
- Problem formulation (project plan)
- Systemic knowledge review (Knowledge Model)
- Computational Model
- VPop
- Validation
- Simulation
The approach is transparent. You can trace a simulation output back to the knowledge or data that fed it, and see how the output was obtained, through clear documentation of the M&S process that the Jinkō platform generates automatically. Traceability is a property of a knowledge-driven model, and it is what a statistical or algorithmic model is generally not built to offer.
The context of use (CoU) is where the model's predictions apply: the time and space domain of the patient condition and environment of interest, standard of care included. The research question the model was designed for defines the CoU, not the model.
Risk assessment in M&S means assessing what role M&S plays in the alternative decisions and what harm those decisions could do to patients.
VPop and calibration are two intertwined processes. A VPop is defined by parameters, VP descriptors, and variables.
Validation is what makes simulation outputs credible, which matters for regulatory acceptance of in silico work. It never really ends; it stops when the model is credible enough in the context of use that the research question implies.
Simulation is the main output of the in silico approach: running the model to answer the question.
Nova built Jinkō, a unified M&S and in silico clinical trial simulation platform, to run the six steps and to meet the demands of transparency and continuous model updating.
Overview of the model
Figure 2 gives an overview of the whole model. The functional key parts are shown schematically and dynamically, following administration of a drug, whether a chemical or a biological entity:
- Pharmacokinetics (PK): what the body does to the drug, that is, the journey of the molecule from the intake site (mouth, skin, and so on) to its site or sites of action or to elimination, whether by chemical transformation or excretion unchanged.
- Pharmacodynamics (PD): what the drug does to the body, that is, how and where the molecule interacts with it. At the intended site of action the drug interacts with its target or targets, and the way that interaction takes place is the drug's Mode of Action (MoA).
- Drug effects (DE): how far and by what route the drug alters the course of the disease. The Mechanism of Action names the pathways whose functioning the drug changes. The changes the drug induces are much broader than the system components involved in the Mechanism of Action; DE covers, for instance, how the drug changes the occurrence of clinical outcomes. All these alterations depend on system properties, written into the components of the system and their functions. If you know the drug's target, you do not need its Mechanism of Action: the drug effects you observe by simulation follow from the properties you accounted for. Including the relevant properties of the system is necessary and sufficient, which is why the balance between false positive and false negative components, discussed below, carries so much weight in the model's predictive value.
- The Virtual Population (VP): it carries the inter- and intra-patient variability seen in real life at every level, genomic, phenotypic, environmental, and aging.
- Clinical outcome: the ultimate focus of the M&S process. Even when the question comes at the very start of R&D, such as "which is the best target for this condition?", answering it means ranking effects on the clinical outcome. The EM is what makes that possible. See boxes 4 and 5.
Figure 2: Overview of the whole model

A six-step process
Figures 3, 4, and 5 show the six-step Nova M&S approach, and Box 5 illustrates it. The steps:
- Problem formulation. Delineate the scope of the model so it answers the scientific questions the problem raises. To avoid missing a component of the living system of interest, accept false positive components rather than discard relevant ones, that is, set the perimeter a little wider than necessary. Then split the scope into a series of connected but largely independent submodels. One submodel is mandatory: the clinical submodel that links to the clinical outcomes of interest.
- Systemic knowledge review. Consider for inclusion and curation all knowledge within that perimeter that can be extracted from the scientific literature. The main question in curating a piece of knowledge is its strength of evidence. The piece is then rewritten as an assertion: two or more biologicals associated with, or connected by, a function. Two things at this step determine the relevance of the whole M&S process. First, the review should be broader than necessary so that no relevant knowledge is missed, even if the computational model is reduced later, for instance after sensitivity analysis. Second, normal physiology and biology must go into the model before pathophysiology. This step is delicate and often the most time consuming. It produces the Knowledge Model (KM), a mixture of figures and text arranging assertions in structured order (see Box 5 and Figures 4 and 5).
- Computational Model. Represent the assertions as mathematical equations, following the known laws of biochemistry, physics, and biophysics, then translate those equations into computational code in Jinkō, the Nova platform.
- Virtual Population (VP). Design the virtual population to capture the variability relevant to the question within its context of use (see below). Model calibration happens alongside.
- Validation. Validate the model against the question and the CoU.
- Simulation. Run the simulations designed to answer the question.
Figure 3: six-step Nova M&S approach. Jinkō is the platform designed by Nova to implement the process.

Figure 4: A section of the text of a Knowledge Model
The GALT lies throughout the intestine and represents the largest immune system in the body[1soe1]. The gut epithelium has evolved structurally to form a passive barrier between the antigen-rich environment of the lumen and the sterile core of the organism [2soe1]. Specific entry ports, formed by microfold (M) cells, are present within the epithelial belt overlaying the inductive Peyer’s patches (PPs) and lymphoid follicles [2soe4][3soe1]. Peyer’s patches are small masses of lymphatic tissue that resemble lymph nodes, with a central B cell follicle and surrounding intervening T cell areas [4soe1]. Following antigenic exposure, activated lymphocytes migrate out through lymphatic channels to regional mesenteric lymph nodes (mLNs) [4soe4] where they further differentiate, mature and proliferate [4soe4]. The activated lymphocytes then migrate out through the thoracic duct, into the systemic circulation, and then “home” back to the effector sites of GALT, the lamina propria (LP) [4soe4]; an amorphous collection of vascular and lymphatic-rich connective tissue beneath the gut epithelium [4soe4]. A portion of the circulating cells, originally processed at the level of the gut, migrate out to distant sites within the common mucosa-associated lymphoid tissue (MALT) system (adding to the mass of lymphoid tissue at the mucosal surfaces of such organs as the lung, liver, and urogenital tract) [4soe4].
Legend: Each sentence is an assertion, and the tag at the end of it carries the assertion's strength of evidence. Clicking the SoE button opens the assertion file, with its annotations, the original piece of knowledge, and a link to that piece in the PDF of the full article.
Figure 5: Graphic representation of a Knowledge Model

Box 5: From knowledge to simulation: illustration
a Scientific articles on the biological system to be modeled are identified and analyzed. To apply the process detailed in Box 4, all known and relevant processes from genes to clinical events need to be accounted for, covering both the physiology and the pathology. Relevant knowledge pieces are extracted, curated, annotated, and translated in assertion, function(s) linking biological entities (e.g. “IkB kinase phosphorylates IkB, resulting in a dissociation of NF-kappaB from the complex with its inhibitor”). The curation and annotation process includes the assessment of the strength of evidence (e.g. "Strong: The assertion is unequivocally confirmed by several other experiments or observations; weak: contradictory findings'').
b The assertions are assembled in a Knowledge Model which takes both a textual and graphical form. A number of appendices to the Knowledge Model are documented (knowledge gaps, simplifying hypotheses, etc.).
c Then, the entities and the relationships between them, contained in the Knowledge Model form the components and the functional relationships, respectively, of the final model. This network represents the system of interest. Functional relationships are translated into functional equations (e.g. Michaelis-Menten equation, ionic exchanges through a channel, etc.). These functional equations are then translated into a series of differential equations in order to reproduce the dynamics of the system of interest. Parameters of these equations will become the patient descriptors of the Virtual Population (see Box 2). The resulting system of equations forms the Formal Model, which is the mathematical representation of the Knowledge Model.
d Equations forming the Formal Model are then translated into computer code to generate the Computational Model.
e The Computational Model is used to run in silico experiments (Figure e) and to simulate a clinical trial and predict the clinical benefit of a drug candidate (see Box 4).
A transparent process
Transparency of an M&S process means every step is traceable. In practice, someone who was not involved can trace a simulation output back to the knowledge or data that fed it, and understand how the output was obtained. It serves two tightly linked aims, credibility and understanding. Documentation is the first brick: clear, honest, and current documentation of the M&S process is mandatory. The Jinkō platform generates it automatically.
Transparency is one of the properties a knowledge-driven M&S approach can offer that a purely algorithmic one is not built to provide. Where an algorithmic model answers what the data show, this one answers why, and the two are most useful read together.
Context of use of a model
The context of use (CoU) is where the model's predictions apply. It is the time and space domain of the patient condition and environment of interest, standard of care included.
Kuemmel et al.[2] published a definition of a model's CoU for mechanistic modeling in drug development, adapted from the V&V40 [1] framework for Computational Model validation and originally framed for physiologically based pharmacokinetic (PBPK) models. The CoU, they write, describes "how the model will be used to address the question of interest, i.e. the specific role and scope of the model". They add that "ambiguity in the question of interest and COU can result in (i) reluctance to accept modeling and simulation in a given drug development or regulatory review scenario or (ii) undesirably protracted dialogue between drug developers and regulators on the data requirements needed to establish model credibility. It is, therefore, critical to unambiguously and explicitly state the question of interest and how the proposed modeling and simulation approach will address it".
Both the definition and the comment tie the CoU tightly to the questions the model is designed to answer. In a phase 2 trial, the question may be: what is the optimal dose? The CoU is then the choice of dose for a phase 3 trial comparing the treatment at that dose to its control, together with the patients who will be included in that trial.
In a phase 3 trial, if the question is the effectiveness of the treatment, the CoU covers the eligible patients and, by extension, the patients defined as the target population, since that population is imagined from the findings of this trial and previous ones, and in MIDD from M&S findings as well.
A post-marketing study should include patients corresponding to the target population as defined by the clinical trials and the other data collected during development, simulation outputs included. If the model is used to answer a question about that study, the CoU stays the same, unless the study aims at questions beyond verifying that the phase 3 results hold in real life and that efficacy is sustained over an extended period in a chronic condition without new side effects.
Contrary to a common belief, the CoU is defined by the research question the model was designed for, not by the model.
Just as the CoU is attached to a question, the decision attached to that question is attached to the CoU.
Decision attached to a question and risk assessment
An M&S process is a nonlinear series of operations, some of which appear in Figure 6. Its main output is meant to support a decision. Directly connected to the decision box is risk assessment, which means assessing what role M&S plays in the alternative decisions and what harm those decisions could do to patients. Risk assessment therefore links to M&S validation, through the credibility of the prediction.
Figure 6: Connection between decision and other key elements of the M&S process. Model prediction validation step is not shown in the diagram, although it is linked with almost all the operations shown here.

Virtual Population and calibration
Two intertwined operations
Calibrating a model and creating a Virtual Population (VP) for it are two intertwined operations. Calibration means selecting relevant values for the model parameters, and those values are by construction the VP descriptor values (see Box 2). The descriptors carry the inter- and intra-patient variability, and when a simulation runs, their values are the model inputs. Relevant parameter and descriptor values come from constraining the model dynamics to be biologically plausible, that is, to reproduce what is known or has been observed.
Quantitatively correct and credible model outputs therefore require identifying plausible parameter value ranges. Parameter identification then makes it possible to explore, identify, and assign the variability associated with each parameter.
Parameters, descriptors, variables
Several components of the model-VP combination can vary between virtual patients, over time, or both. All can be either input or output of a simulation, depending on the moment.
Parameters attach to model components, whether representations of biological entities or the functions connecting them. They take their values from VP descriptor values.
VP descriptors are mostly images of the parameters, giving them values. Others are values of variables, or of models specific to those descriptors, such as compliance.
Variables come in several types:
- output of the model, readable from a simulation run, such as the degree of liver fibrosis after a few months of running a non-alcoholic steatohepatitis model of liver disease
- output from a construct, either a specific adjunct to the VP descriptors that makes the VP more realistic, or the result of allometric scaling of parameter values; both are special VP descriptors with no one-to-one parameter connection
- modeled descriptors: in a dynamic VP, some patient descriptors are added and made variable over time through a specific model, such as compliance. They are inputs when a simulation runs, and comorbidities can be handled this way.
Descriptors, in other words, are all the scalar values you could extract from the model, from its dynamics observed at a given time or from its inputs. Some are observable in real life, depending on what can be measured, but most are not.
Descriptor valuation and calibration
Valuation implements the process shown in Box 2, and it is far from automated. During the literature review, published data from text, graphs, and tables about the expected biological constraints on molecular, cellular, and tissue variables and their dynamic behavior is extracted. It then serves as initial control checkpoints, or as constraints for automatic parameter identification, such as biological variables having to be positive. Initial parameter values are generally chosen to obey those constraints.
To narrow the ranges, different techniques are applied, including sensitivity analysis of the model. A multi-objective optimization then determines the parameter value distribution that satisfies the biological constraints under reference conditions.
The list of constraints implemented as scoring, and the record of their non-violation, is generated automatically for each Virtual Patient run as part of the model documentation.
Initial parameter identification infers values more or less directly through hypotheses such as analogy or scaling. What matters is that the value ranges are plausible against what is known. Calibration partly overlaps with this, but uses numerical fitting to refine parameter values, and so descriptors, and to inform the parameters that remain unknown. It accounts for system behavior and constraints and brings predictions closer to experimental outputs.
Given the complexity of the model, not all parameter values can be informed or inferred accurately from the literature or from preclinical experiments, so clinical data refines the chosen values as well. Calibration is an automatic or systematic manual procedure that changes model parameters one at a time and several at a time to fit a desired individual output better than the uncalibrated model does. Calibrating the moments of the VP descriptor distributions, that is, population-level calibration, or calibrating a set of discrete patient realizations in the VP and inferring parameter distributions, are both tasks that clinical datasets typically support. The data used should be as close as possible to the context of use.
Data-driven calibration, generalizability, and the CoU
A note on vocabulary
Generalizability, external generalizability, robustness, external validity: these terms sit close together, and intuition about what they mean is not enough to put them to work in knowledge-based M&S. Some, such as external validity, are well defined for statistical models, that is, models fitted to a dataset. For models that stay close to the data they are fitted to, the domain of generalizability does not extend far beyond the domain of external validity [3].
Without settling the vocabulary here, each term needs a working meaning in knowledge-based M&S. The external validity of a Nova model is the domain of validity described elsewhere in this section, and it relates closely, if not entirely, to validation. The robustness of the model is its ability to keep predictions reliable when the CoU shifts somewhat, which is close to its reusability.
Generalizability is the model's ability to predict reliably beyond the checked domain of validity. It is linked to robustness, and it depends heavily on the material the model is built from. Knowledge-based modeling generalizes further than data-driven modeling, because knowledge is more stable than data, as described above. To a first approximation, robustness is a factor of generalizability. Virtual populations are what let you verify both properties.
Finding the compromise
There are two complementary ways to value model parameters, and so the VP descriptors, and so build the VP. The first derives values from knowledge. The second uses a scoring process, an algorithm, to find values consistent with knowledge, usually knowledge about the behavior of the modeled systems, or with data, for the parameters that cannot be valued the first way or that the modeler chooses not to derive from knowledge. The first is knowledge-driven scoring and the second data-driven. The border runs between knowledge-derived values, whether by direct valuation or through scoring, and data-derived values. The relative proportion of the two has a major impact on generalizability.
Take the case by contradiction. Suppose every descriptor distribution, that is, their joint distribution, comes from calibration run on one set of in vivo data. First, the resulting VP inherits the CoU of the calibration set, the same population in the same context. Second, there is little room to use the model and the VP to extrapolate, for instance in search of a larger target population, even though the whole point of Nova's approach is to predict, which means to generalize and extrapolate.
In that case we are in the same paradigm as a statistical model fitted to a dataset. The only difference between the M&S approach and traditional statistical modeling is what the models are founded on: mathematical equations representing knowledge about each biological interaction of interest, against a global statistical model assumed to fit the data and the way it was collected. Leo Breiman distinguished two statistical cultures, "statistical modeling" and "algorithmic modeling", the latter accounting more for the phenomenon behind the data [4]. M&S is a third culture, a step beyond Breiman's algorithmic one. In the case illustrated here, the third culture moves considerably closer to the other two, because every parameter value results from a data fit. In that extreme case the model outputs certainly fit the CoU, but its generalizability would probably be no better than that of a statistical or algorithmic model.
Since generalizability is a major objective and a claimed benefit of Nova's approach, in a form and scope that vary with the R&D step, this absurd case shows something. To justify a prediction, always the result of generalizing or extrapolating, the proportion of data-driven calibrated parameters has to be as low as possible, the rest being calibrated from knowledge. The counterpart is that the CoU then depends on how the validation population is constituted, which will have to include the CoU patients and others too, which in turn helps enlarge the domain of validity.
Finding the right compromise between knowledge-derived and data-derived parameter valuation is therefore a key factor of generalizability.
Besides the role of the VP and the way it is designed (see Box 6), reliable scoring of the strength of evidence matters. The generalizability of a knowledge-based model can be assumed to increase with the average SoE of the assertions it rests on.
An M&S project results in more than one VP
Before the final VP, the one that corresponds to the CoU and answers the research questions, a series of VPs, that is, sets of inputs to the model, are used (see Box 6).
Box 6: various Virtual Populations
Legend: M = model; P = parameter; De = descriptor; VP = virtual population; JD = descriptors’ joint distribution; CM = computational model
Stage Model M Development Population Parameters Descriptors Objectives & Comments A uncalibrated M a single Virtual Patient all variables a single set of De values 1.find a reference Virtual Patient; 2.M proof of concept B uncalibrated M VP all variables (plausible) uniform Distributions (no correlation) JD find the parameter space (or set) for which the CM converge numerically (not necessarily biologically plausible) = finding the function domain C uncalibrated M VP some are set constant (plausible) uniform JD for the remaining De idem D uncalibrated M VP some are set constant JD derived from knowledge & data 1.calibration on virtual VP; 2.step #2 M proof of concept E uncalibrated M → type 1 calibrated M VP some are set constant JD derived from knowledge & data and adjusted through scoring 1.calibration on a virtual VP with constraints; 2.set of scores resulting in JD adjustment F not fully calibrated M → type 2 calibrated M real aggregated data some are set constant JD derived from knowledge & data and adjusted through scoring 1.calibration on a virtual VP with constraints; 2.set of scores resulting in JD adjustment G not fully calibrated M → type 3 calibrated M real individual data some are set constant JD derived from knowledge & data and adjusted through scoring 1.calibration on a virtual VP with constraints; 2.set of scores resulting in JD adjustment H calibrated M → type 1 validated M real aggregated data some are set constant JD derived from knowledge & data and adjusted through scoring 1.validation on summarized data; 2.results in validation metrics values I calibrated M → type 2 (ultimate) validated M real individual data some are set constant JD derived from knowledge & data and adjusted through scoring 1.validation on summarized data; 2.results in validation metrics values; 3. and in a final validated M J validated M real individual data some are set constant JD derived from knowledge & data and adjusted through scoring 1.simulation(s); 2.answering the questions
Calibration and validation: strictly non-overlapping knowledge and datasets
The sets of knowledge and data used for calibration, that is, valuing the VP descriptors' joint distribution, and for validation, the main component of assessing prediction credibility, should not overlap. Otherwise the derivation is tautological and the prediction less credible.
Validation
Why validation is crucial
A model has to be validated for its predictions to be credible. The validity of simulation outputs is also central to regulators' acceptance of in silico clinical trials (see below). The principles of science, the experimental method, and inferential statistics applied to therapeutic efficacy and safety justifiably shape how people assess evidence from in vivo or in vitro experiments and observations[5] [6] [7]. Beyond the need for unbiased comparison, weighing strength of evidence means removing random findings that come from the inherent variability of living systems, hence the statistical paradigm.
That paradigm hardly applies in silico, because you are not working with a sample drawn from a theoretical parent or real population but with the whole virtual population. Imagining a parent population for a virtual population is far-fetched, and its only benefit would be to apply the statistical paradigm of life and earth sciences without thinking. Run enough simulations and any tiny difference in system behavior will emerge, clinically relevant or not. A model can also be retro-engineered by building the expected results into it, which demonstrates nothing but itself.
The only acceptable way to assess simulation output validity is a standard certification process that stakeholders accept. A standard quality control procedure evaluates whether a pre-specified set of rules has been followed. Table 2 summarizes those principles.
Table 2: In silico approach validation principles
| Principle | Definition |
|---|---|
| Traceability | Knowledge incorporated into the disease model needs to be fully traceable back to each primary source. |
| Consistency | Knowledge incorporated into the disease model needs to be consistent with current science. |
| Bias | The mathematical and computational models need to be unbiased representations of the selected knowledge. |
| Internal validity | The disease model needs to be capable of “reproducing” knowledge and data that have not been used to design it in the first place. |
| Prospective validation | Simulated predictions of clinical trials or in vitro results must be qualitatively and quantitatively in line with future trial/experiment/observation outputs. |
Validation beneath the surface
Validating a model is an ongoing process that never ends. There is no real final validation step. When a model is built to answer a question, which is the right way round, its final validation has to be designed so that the predictions the answer rests on are credible enough. Those predictions come from the context of use that matches the research question, the transparency of the M&S process has made it possible to check their relevance, and the metrics measuring the gap between predicted and real data fall within acceptable confidence intervals.
A new validation becomes necessary when a new context of use is formulated to address a new research question. New data more relevant than what the final validation used can justify one too. That would be the case if the context of use was a phase 3 trial and data from a post-marketing study became available. In most cases the post-marketing study is integrated into the M&S process as a new final validation step, especially when Health Technology Assessment (HTA) decisions rest on the model's predictions.
Context of use and validity domain
Validation is the process of determining how accurately a model, its associated VP, and its simulations represent the real world, and particularly the part of it that matters to the R&D project at hand. It should be viewed from the perspective of the intended application of the model or simulations under similar conditions. The aim is to demonstrate that the predictions the computational approach generates for this project, on this question, are reliable. The question defines the context of use (CoU): the role of M&S in the project, and the patients, their conditions, and the care environment the question concerns. The output of validation defines the credibility and the domain of validity of the model, and so of the predictions simulations produce with it. Since simulations run on a virtual population, the Virtual Population has to be part of validation. Depending on the outputs, the validity domain can narrow the context of use.
Validation and data
As with calibration, the quality of validation depends heavily on the quality and relevance of the available experimental, clinical, and epidemiological data. The final validation step should compare outcome predictions from simulations to real patient outcomes in the CoU of interest.
The thorny question of metrics
Validation is a series of steps. Each provides either a qualitative judgement, such as "one can trace back from this simulation output to the pieces of knowledge that were inputs in the model design", or a measure of fit between reference data and model predictions. A metric expresses that distance. Since the probability that reference data and predictions are rigorously equal is effectively zero, the metric has to be compared against an acceptable value, which means setting a threshold for deciding whether the model counts as valid at that step.
Because there are several steps, model validity is a global judgement summarizing the credibility of the model.
Nova validates the model and the virtual population in accordance with the latest guidelines [VV40, 2018] and publications [Erdemir 2020, Viceconti 2019, Friedrich 2016], comparing them to relevant preclinical, observational, and experimental clinical data, depending on the question the model addresses and on what data is available.
Regulatory guidance on M&S validity
Nova validation plans follow the regulatory guidance on reporting modeling. The closest guidance to Nova's disease and treatment models is:
- the ASME V&V40 framework, summarized for application to mechanistic PBPK models by Kuemmel et al. and treatable as a regulatory reference for the FDA [VV40, 2018]
- the EMA's specific guidance on reporting PBPK models [8]
Validation is a progressive process
Validating a model and a VP at Nova is a continuous process rather than a one-time procedure. It follows the general principles above and applies during almost every step of M&S, starting with the KM.
Validation is necessary after:
- any modification, removal, or addition of at least one functional relation in the Computational Model or in the Virtual Population
- a new use of the model, meaning a change of context of use
Somewhat arbitrarily, validating a model at Nova can be described as a series of questions in five categories.
1. Qualitative validation. Qualitative criteria supporting credibility can be assembled separately from the quantitative comparison of predictions with reference data. The FDA acknowledged them during Nova MIDD meetings in 2019. An assessment of these qualitative credibility goals is added to the validation report.
- Model content: how was the Knowledge Model validated? Is the model granularity adapted to the question of interest? Can you reach the knowledge that justifies the model form and the parameter values? This is the transparency check.
- Model reuse: has the model been used in a different context of use?
- Model inputs:
- Uncertainty management: are the uncertainties associated with the model form and inputs, and their impact on predictions, understood and controlled for?
- Sensitivity analysis: is the model's sensitivity to input parameters, and its effect on predictions, understood and controlled for?
- Risk of tautology: could a bias in the model's structural form or inputs produce an answer that favors the desired outcome?
- Simulation design: is the in silico experiment design relevant to the questions of interest?
- Relevance: are the M&S results relevant to the clinical purpose, and to the context of use?
In short, this demonstrates that the Computational Model, the Virtual Population, and the simulation protocol were developed appropriately to support the simulations, and so to make the predictions and possible extrapolations of interest.
2. Semi-quantitative validation of the prediction. This assures that submodels reproduce behaviors of the modeled systems consistent with their known behaviors, or at least plausible where those behaviors are not precisely known. An untreated tumor receiving sufficient nutrients and no antagonist signals, for instance, should grow.
3. Quantitative validation of the prediction. This is the last and most relevant step. It compares the full model's output predictions in a real group of patients similar to the CoU with the outcomes observed. The model runs on those patients' baseline data, and the simulations continue to the end of the real-world observation period. Depending on data quality, the chosen validation metrics, the size of the real patient group, and the questions addressed, the compared outputs are either averaged values, such as the rate of a binary outcome or the mean of a continuous one, or individual outcomes, such as whether an event occurred for each real patient.
4. Verification. Code verification of the solving algorithm runs automatic tests on randomly generated ordinary differential equation systems, comparing the explicit closed form solution with the output of the resolution.
Calculation verification estimates and minimizes the numerical error introduced by approximating the mathematical model, by solving the system at various tolerance values, comparing outputs, and checking convergence.
Verification also covers peer review of the code and continuous testing to catch user errors. Project-specific evaluations can be added, such as verifying a spatial discretization implemented to mimic diffusion processes.
5. Model input sensitivity and uncertainty analysis. Sensitivity analysis examines how sensitive the Computational Model outputs are to the model inputs. Combined with uncertainty analysis, it identifies the impact of parameter uncertainty on the variables of interest and the final predictions. The parameters concerned are those known or likely to carry a measurement error, those resting on a low strength of evidence assertion, or those whose value remains uncertain after calibration. The aim is to confirm that the model is not sensitive to parameters or variables carrying high uncertainty, from knowledge gaps for instance.
Simulations
Everything above exists for the sake of simulation, which means running the model to answer the question. It is an experiment, carried out in silico. Apart from that, it follows the rules of the experimental method, the same rules Claude Bernard set out in the 19th century and applied to in vivo and in vitro experiments to ensure their results hold[4:1].
The obvious and formidable difference is the experimental units: they are virtual. That has two consequences, both of which strengthen the validity of the comparison an experiment always makes between a new scenario and a control. First, the compared scenarios can be applied to exactly the same experimental units, which removes every source of bias. Second, the only limit on the number of experimental units is computation time, which is in any case far shorter than the duration of the corresponding in vitro or in vivo experiment.
As with in vitro or in vivo work, a protocol is written before the model runs, one of Claude Bernard's prerequisites of the experimental method.
The Jinkō platform
Nova built Jinkō to run the six steps shown in Figure 3 and to meet the demands of transparency and continuous model updating that the virtuous circle implies (see Figure 1). Jinkō is Nova's unified M&S, or in silico clinical trial simulation, platform. With mathematical models and virtual populations at the heart of the company's approach, the platform takes its name from a fortunate homonymic collision: the Japanese words for "artificial" (人工, Jinkō) and "population" (人口, Jinkō). Jinkō is written in Haskell [9], a functional language. It has two modules: Jinkō Knowledge, which stores, handles, and uses bibliographic references and literature reviews, and Jinkō SimWork, the computational management module.
Jinkō Knowledge: the knowledge management module
Jinkō Knowledge is a community-driven knowledge engine for curating and organizing biomedical knowledge extracted from white and grey literature, with the goal of building and maintaining state-of-the-art Knowledge Models of pathophysiological processes. Built on human investigation for curation and semantic web formalism for exploitation, it lets researchers and biomodelers curate, formalize, and share biomedical knowledge from the scientific literature in an open science environment. Through it, biomodelers can:
- gather the necessary scientific publications in a dedicated project
- extract pieces of knowledge from those publications by creating assertions
- grade each assertion's Strength of Evidence (SoE), and annotate it, for instance by linking a protein in the assertion to descriptive and gene data banks
- organize the assertions into a discursive format, in plain language, describing the pathophysiology of interest
- represent the findings graphically
Representing the assertions' relationships as a network produces a Knowledge Model, along with a list of knowledge gaps.
The platform is a community environment: non-modelers such as clinical experts are invited to contribute to a project and to give direct feedback on the literature review the biomodelers conducted.
Jinkō SimWork: the computational management module
The Jinkō Model Solver, at the heart of the platform, runs atomic simulation tasks, meaning a single virtual patient with a single experimental setting on a single model, in a stateless and purely functional way. The same input set always leads to the same outputs, which makes caching safe and efficient.
SimWork also hosts calibration and analytics modules. Calibration uses genetic algorithms to calibrate model parameters from relevant knowledge extracted from the literature and all available preclinical and clinical data. The analytics modules help analyze simulation results and produce reports in line with traditional clinical trial reports, extended with what the in silico approach makes possible, such as identifying biomarkers.
1. (n.d.). Assessing Credibility of Computational Modeling through ... - ASME. Retrieved May 30, 2021, from https://www.asme.org/codes-standards/find-codes-standards/v-v-40-assessing-credibility-computational-modeling-verification-validation-application-medical-devices ↩︎
2. Kuemmel C, Yang Y, Zhang X, Florian J, Zhu H, Tegenge M, Huang S-M, Wang Y, Morrison T, Zineh I. Consideration of a Credibility Assessment Framework in Model-Informed Drug Development: Potential Application to Physiologically-Based Pharmacokinetic Modeling and Simulation.CPT Pharmacometrics Syst. Pharmacol 2020, 9, 21–28 ↩︎
3. Walter A. Kukull, WA, Ganguli, M. Generalizability. The trees, the forest, and the low-hanging fruit. Neurology, 78; 1886-91 (2012) ↩︎
4. Breiman L Statistical Modeling: The Two Cultures. Statistical Science; 16, 199–231 (2001) ↩︎ ↩︎
5. Bernard, C. An introduction to the study of experimental medicine (1865) ↩︎
6. Neyman, J. & Pearson, E. S. in Breakthroughs in Statistics 73–108 (Springer, 1992) ↩︎
7. Popper, K. The logic of scientific discovery. (Routledge, 2005) ↩︎
8. (2018, December 13). (PBPK) modeling and simulation - European Medicines Agency |. Retrieved June 29, 2021, from https://www.ema.europa.eu/en/documents/scientific-guideline/guideline-reporting-physiologically-based-pharmacokinetic-pbpk-modelling-simulation_en.pdf ↩︎
9. Marlow, S., & others. (2010). Haskell 2010 language report. Available Online https://www.haskell.org/onlinereport/haskell2010/ ↩︎
