WO2008053499A1 - Method for calculating the probability of developing tumoral or degenerative diseases. - Google Patents

Method for calculating the probability of developing tumoral or degenerative diseases. Download PDF

Info

Publication number
WO2008053499A1
WO2008053499A1 PCT/IT2006/000763 IT2006000763W WO2008053499A1 WO 2008053499 A1 WO2008053499 A1 WO 2008053499A1 IT 2006000763 W IT2006000763 W IT 2006000763W WO 2008053499 A1 WO2008053499 A1 WO 2008053499A1
Authority
WO
WIPO (PCT)
Prior art keywords
value
expression
population
tumoral
binary values
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IT2006/000763
Other languages
French (fr)
Other versions
WO2008053499A8 (en
Inventor
Daniel Levi
Paolo Gaetani
Riccardo Rodriguez Y Baena
Giovanni Broggi
Enrico Aimar
Lorenzo Panella
Michele Tedeschi
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Humanitas Mirasole SpA
Original Assignee
Humanitas Mirasole SpA
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Humanitas Mirasole SpA filed Critical Humanitas Mirasole SpA
Priority to PCT/IT2006/000763 priority Critical patent/WO2008053499A1/en
Publication of WO2008053499A1 publication Critical patent/WO2008053499A1/en
Publication of WO2008053499A8 publication Critical patent/WO2008053499A8/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B5/00ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • G16B20/20Allele or variant detection, e.g. single nucleotide polymorphism [SNP] detection
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B25/00ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression
    • G16B25/10Gene or protein expression profiling; Expression-ratio estimation or normalisation
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B5/00ICT specially adapted for modelling or simulations in systems biology, e.g. gene-regulatory networks, protein interaction networks or metabolic networks
    • G16B5/20Probabilistic models
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B20/00ICT specially adapted for functional genomics or proteomics, e.g. genotype-phenotype associations
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B25/00ICT specially adapted for hybridisation; ICT specially adapted for gene or protein expression

Definitions

  • the present invention relates to a method for calculating the probability of contracting tumoral or degenerative pathologies in populations of healthy individuals .
  • the present invention further relates to a method for determining the prognosis on the progression of the disease in patients suffering from tumoral or degenerative pathologies, as well as a prognosis on the effectiveness of a therapy adopted for said pathology.
  • the method being the object of the present invention is based on the evidence that a purely clinic approach is not sufficient in the oncology field. It can be thus stated that the prognostic determination of the disease along with the probability of contracting a tumoral pathology by healthy patients generally elude forms of control and prevision that are really reliable.
  • the behaviour of a chaotic system results from the combination of many ordered behaviours, none of which prevails under normal conditions. It has been demonstrated that, by suitably perturbating a chaotic system it may be guided to follow one of the regular behaviours thereof.
  • the chaotic systems are deterministic, i.e. when two almost identical systems are pushed or guided by the same signal, they will produce the same "output” , even if it is impossible to know which is the "output” .
  • the complex chaotic systems such as biologic systems are however subjected to a further phenomenon that may be called anti-chaos.
  • anti-chaos In fact, several very disordered systems spontaneously "crystallize" into a highly ordered state.
  • complex systems in biology from the thousands of genes that self-organize within a cell to the cell network and molecules mediating the immune response, and the like.
  • genome self-regulating system which is a good example of how the anti-chaos can govern gene development and regulation.
  • Cell types differ because varied forms of genetic activity are active in them, not because they have different genes.
  • the genes mutually regulate their activity either directly or by means of gene products.
  • the coordinated behaviour of this system is at the heart of cell differentiation and integration of genes, including those involved in oncogenesis.
  • the first step is to couple a binary variable, which can be either active or inactive, to the behaviour of each element in the system (a gene, in the instant case) .
  • the rules that govern these mutual interaction systems are stochastic Boolean networks.
  • the mathematical models applied in this study for building a system of simulation of gene networks involved in the oncogenesis are the so-called autonomous stochastic Boolean NK networks, where N is the number of genes involved, each of which has K inputs. These networks are autonomous because all the inputs comes from inside the system.
  • a stochastic Boolean network has a finite number of states . A system must therefore eventually reenter a state that it has previously encountered, and as the system is deterministic, it will repeat the same state sequence as it did before and will indefinitely cycle repeatedly through the same cycle of states until a perturbation occurs in the system. These cycles are called the dynamic attractors of the network.
  • a minimal perturbation occurs when a binary element temporarily flips to its opposite state (in the binary language, 0 sets to 1 or vice versa) . If this event does not push the network outside its original basin of attraction (homeostasis cycle) , it will eventually return to its starting cycle; but if the network is pushed into a different basin of attraction
  • a tissue that turns from an acute to a chronic inflammatory state is a case where a shift to a new homeostasis is achieved.
  • the attractors have various levels of stability in relation with minimal perturbations. Some can neutralize any single perturbation, others can withstand only a few, still others are destabilized by any perturbation.
  • a structural perturbation is a permanent mutation in the connections or in the Boolean functions of a network
  • DNA polymerase is a nanomachine, which is capable of copying a DNA template into a new strand.
  • Computer laws can be thus applied to the genome regulation and evolution system, which is the foundation of several phenomena such as, inter alia, the oncogenesis process.
  • An error in the genome produces signals in the cell which result in different biological responses, such as the stop of the cycle, reparation of DNA, cell death, or uncontrolled proliferation. Controls are usually carried out by an opposite action of two gene systems. The entire lifespan of a cell can be seen as a planned execution of proliferation, stop-quiescence, differentiation and death programs, each corresponding to the activation or inactivation of specific packages of genes .
  • the DNA polymerase can make a mistake; or the DNA polymerase can properly operate, but the DNA to be replicated may be altered. Independently of the origin of this error, the cell that will inherit this DNA will have an altered length of DNA. If the portion of altered DNA codes for a protein, the thus produced protein may be, in turn, either inactive or even fatal.
  • the onset of a tumour can be related to the failure of a (also unknown) process, for a predetermined number of times.
  • a (also unknown) process for a predetermined number of times.
  • cells reproduce by duplicating and by performing a careful numerical control system on their population. This control acts such that the number of cell individuals is maintained constant and does not grow according to second powers as mitotic duplication would require.
  • This mechanism is the result of complicated gene expression and its failure can be considered as an event governed by statistical laws; it is thus possible to assign a finite probability other than zero to the birth of a neoplastic cell .
  • the tissue is divided into independent generation channels.
  • neoplasia After a tumour has generated within a generation channel, the evolution of this channel is exclusively characterized by neoplasia; at a macroscopic level, a tumour appears in time t only when the number of neoplastic generation channels is higher than a predetermined number.
  • the model considers the continuous expression of, at tissue level, at least three genes and simulates the general state of the tissue by monitoring the number of errors that have occurred in the duplication of the genes involved. Synergies are taken into consideration, if present.
  • the simulation provides that N cell families exist, each being independent from the others and from the surrounding cell system. Each family evolves by generating only one daughter cell: this simplification takes into account the collective self-regulation system of the tissue, which maintains constant the population of the cells belonging to the tissue itself. As stated above, the evolution of a family is completely independent from the surrounding conditions, except for the time factor: the probability that a healthy cell can experience a mutation to a neoplastic cell increases over time. Therefore, the simulation is the calculation of the probability that within a certain time (i.e. after a certain number of mitotic divisions) at least one neoplastic cell will be formed.
  • the incidence of a particular type of tumour is the percentage fraction of cell families that have a neoplastic progenies after the preset period of time.
  • tumour suppressor and oncogenic genes respectively.
  • the results depends on the number of genes.
  • a tumour is generated when the gene controlling neoplastic proliferation stops operating properly. Consequently, the cell stops responding to growth inhibitor signals and acquires the capacity of proliferating in an uncontrolled manner.
  • the present invention is directed to provide a method for assessing the prognosis of a tumour pathology, which is very useful in order to select each time the most proper therapy.
  • the various forms taken by a tumour are clearly the result of different patterns of molecular abnormalities . It is thus essential to be capable of determining which genetic abnormalities have occurred in every single case.
  • the tool allowing such assessment is the microarray, which uses the so-called DNA chips. This known technology allows identifying the differences in the activities exhibited by the hundreds of genes expressed by the tumoral cells upon the diagnosis.
  • the DNA chip is used to analyze the level of gene expression in a tissue sample and consists of a layer of single-strand molecules of DNA (the so- called probes) that is immobilized on a support (glass, synthetic material or silicon) .
  • the DNA chip uses the DNA property of complementary pairing its bases.
  • the base A (adenine) can be paired only to base T
  • each type of probe (either an entire gene or a shorter DNA sequence) is immobilized at a well-precise point in the support grid.
  • the sample of tissue to be examined is treated such as to isolate the mRNA thereof and to obtain the corresponding cDNA in labelled form (fluorescent label or the like) .
  • the microarray is then incubated with the thus-obtained cDNA solution, then it is separated and suitably analyzed, such as by means of laser scanning.
  • the cDNAs of the genes expressed in the sample bind to the immobilized DNA of the corresponding gene, which was in a predetermined position, as stated above .
  • the raw data are computer-processed and translated into a colour-coded map. From this map, one can thus identify which genes are active (expressed) in the examined tissue and which are "switched off” (unexpressed) .
  • the DNA chip contain thousands of genes on an individual microarray, which allows carrying out a broad parallel gene assessment, by particularly identifying which genes are expressed in the various types of tumour tissue as compared with healthy tissue.
  • These devices further allow assessing the expression level of a determined gene using a sample of tissue containing a known number of cells and assessing the percentage of cells in that tissue which express the gene being investigated.
  • a gene is said to be expressed when it is transcribed to a molecule of mRNA and subsequently translated into a protein. This assessment may be carried out, for example, by means of fluorimetry. A mapping of the expressed genes will be thus obtained, the relative percentage of expression being associated with each of them.
  • DNA chips are commercially available.
  • DNA chip supplying Companies are for example: Affymetrix (Santa Clara, California) , Agilent Technologies (Palo Alto, California) , Perkin Elmer (Boston, MA) or Rosetta Inpharmatics (Kirkland, Washington) .
  • Data of microarrays relating to patients suffering from various tumours are also available in the literature. These data are either in the form of tables or lists relative both to diseased tissue and healthy tissue of the several patients and are averaged among all the patients in the examined population.
  • the method of the present invention allows providing, for a determined individual, a probability of developing a determined tumoral pathology.
  • the method of the invention allows assessing the probability of developing a determined tumoral pathology in a population of individuals, either healthy or diseased, that are selected according to various criteria. These selection criteria may be, for example, geographic or racial origin, or a particular life and/or work environment; alternatively, the selection may be carried out age brackets or gender, or diet or lifestyle, or pre-existing genetic, family predisposition or particular pathologies to which the population is subjected. This list is merely exemplary and accordingly it cannot be considered as limiting the field of application of the invention.
  • the probability of developing a given tumoral pathology is meant, according to a first aspect, the probability that the healthy individual contracts a determined tumoral pathology.
  • the "probability of developing” is the probability that an individual who already suffers from a given tumoral pathology has a particular pathologic progression, for example he/she is experiencing an evolution from a tumoral form to a different one. The calculation of the "probability of developing" a given tumoral pathology will thus imply, in this case, the determination of the prognosis for the determined patient .
  • the method according to the present invention for determining the probability of developing a given tumoral pathology in an individual can comprise the following phases :
  • the present invention relates to a method for determining the probability of developing a given tumoral pathology in a population of individuals comprising the following phases:
  • level of expression is meant the percentage of DNA probes that, for a determined site, are coupled to cDNA of the tissue of the individual being examined, as results from the microarray data. In a population of individuals, the "level of expression” is given by the summation of the levels of expression of the single individuals in the population, divided by the number of individuals.
  • the first phase (phase (A) ) of the method according to the invention is selecting, for the tumoral pathology being the object of the method, a set of oncogenic and tumour suppressor genes that are characteristic of said tumoral pathology.
  • oncogenic is meant a gene deputed to controlling cell proliferation.
  • tumor suppressor is meant a gene deputed to controlling the suppression of cell proliferation.
  • microarray methodology is the most recent acquisition and the results thereof are published in the literature or provided on specific databases or however made available to the public.
  • a further source for obtaining these data is the studies that have been carried out by the same entity that embodies the method of the invention, with the proviso that the results are statistically significant.
  • tumoral pathology is meant both a type of tumour that has been identified as a function of the target organ, and a form of manifestation or evolution of the primary tumour being the object of a cell differentiation, or the like.
  • the tumoral pathology being the object of the present method will be both any tumoral pathology of which the probability of developing the latter by a healthy individual is desired to be assessed, and a possible evolutionary form of a primary tumour from which the individual being examined is suffering.
  • epilevolutionary form is meant an undifferentiated tumour that has developed from a primary tumour.
  • the phase A may also be carried out very upstream of the actuation of the method on a given individual, such as to build a table or database that contains the information required for each known type of tumoral pathology.
  • the selection of the set of oncogenic and tumour suppressor genes will be carried out by taking into account the level of expression of the gene in the determined pathology being examined, such as to have a substantially balance between the levels of expression of the oncogenic genes and the levels of expression of the tumour suppressor genes being selected.
  • the phase A of the inventive method will comprise the following steps: i) providing statistic data of gene mapping and levels of expression thereof for the diseased tissue and healthy tissue of a population of patients of a determined tumoral pathology; ii) selecting from said statistic data obtained according to step i) a number n of oncongenes expressed in the diseased tissue of said population of patients of said tumoral pathology, wherein n is equal to or greater than l.
  • step i) selecting from said statistic data obtained according to step i) a number m of tumour suppressors expressed in the healthy tissue of said population of patients of said tumoral pathology, wherein m is equal to or greater than 1, such that the summation of the levels of expression of said tumour suppressors in said healthy tissue substantially corresponds to the summation of the levels of expressions of said oncogenes in said diseased tissue, thereby obtaining a set of oncogenes and tumour suppressors characteristic of said tumoral pathology.
  • step iv) assigning to each of said oncogenes and tumour suppressors of the set of genes that is selected according to the steps ii) and iii) the relative level of expression in the diseased tissue of said population of patients of said tumoral pathology; v) creating a set of binary values from said set of genes in step iv) , according to the following principle: - assigning the value of 1 to each oncogene having a level of expression in said diseased tissue that is greater than or equal to a preset value, preferably 50%, and the value of 0 to each oncogene having a level of expression in said diseased tissue that is lower than said preset value;
  • n + m will range between 2 and 5, more preferably it will be 3. While, in fact, selecting a set comprising a large number of genes is desired, such as to obtain a more accurate determination, the number of possible combination would become so high that a long processing time would be required with normally available computer equipment.
  • the tumour suppressors will be selected in the healthy tissue of the patient population such that the summation of the levels of expressions of the two tumour suppressors will have to be about 80%, for example 40% the first one and 40% the second one. The same will apply in the event that one tumour suppressor and two oncogenes are selected.
  • the set of binary values for the selected genes will be (1 1 0) .
  • the first tumour suppressor has a level of expression that is higher than half the relative level of expression in the healthy- tissue (>20%)
  • the second tumour suppressor the level of expression is lower than the respective half ( ⁇ 20%) .
  • the set of binary values obtained according to the step v) of phase A of the inventive method will be the so-called "crash set", i.e. the set of binary values corresponding to the occurrence, in the patient, of the conditions for developing the tumoral pathology under examination, as will be better detailed below.
  • a gene mapping may be similarly used which derives from a healthy population other than said population of patients, without however departing from the scopes of the present invention.
  • the second phase (phase B) of the method comprises the determination, in an individual being examined, of the level of expression of each of the genes in the set of oncogenic and tumour suppressing genes selected according to the phase A. Substantially, for each gene
  • the gene (oncogene and tumour suppressor) belonging to the selected set it is determined whether this gene is expressed or not expressed in the individual, and the level of expression thereof. If the gene is expressed with an expression level greater than a determined threshold level, then it is assigned the binary number 1, whereas if it is not expressed, it is assigned the binary number 0.
  • the phase B of the inventive method thus comprise the following steps: i) carrying out a gene mapping with the relative levels of expression for a tissue sample of an individual being examined; ii) assigning, for said individual, the level of expression relative to each gene in the set of oncogenic and tumour suppressor genes selected according to the phase A; iii) creating a set of binary values from the set of genes obtained according to step ii) , according to the following principle: - assigning the number 1 to each oncogene that has a level of expression in the tissue sample of said individual being equal to or greater than 1 A the relative level of expression in the diseased tissue of said patient population according to the phase A and the number 0 to each oncogene that has a level of expression in the tissue sample of said individual being lower than % the relative level of expression in the diseased tissue of said patient population according to phase A; i) - assigning the number 1 to each tumour suppressor that has a level of expression in the tissue of said individual being equal to or greater than 1 A the relative level of expression
  • the tissue sample of the individual being examined may be a sample of healthy tissue, or alternatively, a sample of diseased tissue in the case where the probability has to be determined that the tumoral pathology from which the individual is, in this case, already suffering will evolve to a different evolutionary form.
  • the tissue sample may be taken by any organ, particularly a target organ of a possible tumoral pathology.
  • the phase C of the inventive method provides calculating a simulation of gene evolution, in order to obtain a development probability value for the tumoral pathology being examined.
  • the simulational calculation according to the phase C is carried out starting from the set of binary values that is obtained according to the phase B, step iii) , as will be set forth below.
  • the probability of developing a tumoral pathology is calculated for a population of individuals (phases A' , B', and C).
  • the phases A', B' and C are carried out in an entirely similar manner to phases A, B, and C for the single individual, except that the level of expression of the genes will be an average value across all the population individuals, such as discussed above.
  • This embodiment will allow obtaining a global assessment of the probability that a population has of contracting a determined tumoral pathology.
  • This population may be selected either according to the criteria listed above or according to different criteria, without however departing from the scope of the present invention.
  • the calculation of the probability of developing a tumoral pathology is carried out using a random simulation algorithm, the initial input thereof consisting of the set of binary values obtained for the individual (phase B, step iii) or the population of individuals (phase B'), wherein the level of expression assigned to each gene represents the probability that the initial binary value for that gene will be maintained.
  • the simulation provides a number ⁇ of interactions, where ⁇ is a preset number depending on the type of tissue (and thus, of cell) being considered. Generally, 1 will be equal to the number of mitotic cycles that the type of cell being examined performs in the time unit multiplied by the desired number of time units to be considered. For example, the calculation may be carried out over the entire lifespan of an individual, or just on a limited period of time which may correspond to age brackets in which the risk of contracting a particular tumour is the highest, in yet another case, particularly in order to predict the progression of the disease in a patient who is already suffering from a tumoral pathology, a period of time of a few months may be considered.
  • This random simulation algorithm provides the random generation of numbers, using a number n + m of parent routines and one daughter routine.
  • the parent routines are completely independent and completely blind relative to the surrounding environment.
  • the behaviour of the daughter routine is, on the other hand, influenced by the behaviour of the parent routines on which it depends .
  • Each parent routine is assigned a gene from the set of genes that has been initially selected.
  • n + m corresponds to the number of oncogenic (n) and tumour suppressor genes (m) being initially selected.
  • n+m will be equal to three.
  • Each parent routine has a probability of giving an error result which corresponds to the random drawing of a particular number on a preset sample. This situation simulates the randomness by which a gene can express in an erroneous manner.
  • the error result occurs due to the generation of the binary set corresponding to the tumoral pathology being examined, which can be deduced by the initial data obtained on a patient population according to the phase A.
  • the method of random drawing of the binary value during the iterations is corrected by an algorithm, which takes into consideration the different probabilities for the various genes to maintain the value of 1 or 0 that was initially assigned thereto.
  • This different probability can be related with the level of expression of the individual genes in the starting set of genes of the individual or population of individuals being examined.
  • this random simulation algorithm will be a so-called "genetic algorithm", such as described in Sankara K.Pal et al., Neuro Fuzzy Pattern Recognition: Methods in soft computing, appendix a genetic algorithms: basic principles, features, Winley - Interscience publication, 1999.
  • the random simulation algorithm randomly generates, at each iteration, a parameter value (1 or 0) . This value may be thus either equal to the one of the preceding iteration or other than the latter.
  • the algorithm has been changed such as to take into consideration the probability that the initial value is maintained, and accordingly it comprises the following steps:
  • espr n and espr m are the levels of expression of the genes in the individual being examined that are determined according to phase B and each value of par n and par m will depend on the relative parent routine.
  • the value of par ⁇ or par m can be 0 or 1 and corresponds to the binary value assigned to the gene.
  • the "behaviour" function may accordingly do not range between i-th and i+l-th iteration if a set of values of par n and par m all being equal to 1 or 0 is generated (situation of perfect balance between oncogenes and tumour suppressors) , or it may increase or decrease by a small amount when a value results to differ from the others.
  • the "behaviour" function will vary between more or less than a basal value. Accordingly, an individual abnormal event is not catastrophic, because it alters the behaviour value only at the i-th step, the behaviour value will not thereafter change until a further abnormal event occurs. This further abnormal event may also correct or compensate the first abnormal event, thus resulting in the behaviour value varying between more or less than the initial value .
  • the selection of the maximum absolute value will be done by adopting the fixed delta principle of the "behaviour” function between basal value and i-th iteration.
  • the absolute value of delta above which an event will be considered as catastrophic will be 30 relative to the initial value of the "behaviour” function.
  • the "behaviour” function assumes a value equal to the initial value plus 30 or more, the system will be clearly unbalanced towards the uncontrolled proliferation, and thus tumorigenesis.
  • the value is equal to the initial value less 30 or more, there will be an unbalance towards apoptosis. This condition will occur in the case of degenerative diseases, such as Alzheimer's disease.
  • the phase C thus provides a step of calculating a gene evolution simulation with a random simulation algorithm, such as stated above.
  • the phase C of the inventive method further provides the following steps : i) obtaining a list of binary values for each of the parent routines relative to all iterations; ii) searching in said list the set of binary values corresponding to the maximum absolute value of the "behaviour” function and compare the same with the set of binary values typical of the tumoral pathology being examined according to phase A, step v) or according to phase A' ; iii) when said binary values being compared according to step ii) are equal, counting the number N of occurrences of the set of binary values corresponding to the maximum absolute value of the "behaviour” function; or iv) when said binary values compared according to step ii) are different, counting the number N of occurrences both of the set of binary values corresponding to the maximum absolute value of the "behaviour” function and of the set of binary values typical of the tumoral pathology being
  • the inventive method will also comprise the calculation of the probability that abnormal events may occur, which correspond to an intermediate situation between the equilibrium situation and the crash situation (development of the tumoral pathology) , but that are however indicative of a predisposition to developing said pathology. This is the typical pre-tumoral situation.
  • This calculation is carried out by means of a subroutine that provides the following steps : a) obtaining the set of binary values corresponding to the maximum absolute value of the "behaviour" function from the list in phase C 7 step i) ; b) changing one value at a time of said set of binary values, thus generating n + m subsets of binary values; c) calculating the value of the "behaviour” function for each of said n + m subsets of binary values according to the formula
  • CDehaV eguppressor and selecting the subsets of binary values for which the value of said "behaviour" function is intermediate between the basal value and said maximum value; d) counting the number N abnormal of occurrences of the subsets of binary values selected according to the step c) in the list obtained according to phase C, step i) ; e) calculating the probability Pabnormai of occurrence of abnormal events by applying the following algorithm
  • a personal computer may be used which comprises a collector connecting processing means, for example a central processing unit (CPU) , to memory means that include, for example, a RAM work memory, a read-only memory (ROM) - which includes a base program for starting the computer, a magnetic hard disk, optionally a drive (DRV) for reading optical disks (CD- ROMs) , optionally a floppy disk read/write drive for.
  • the computer can comprise a MODEM or other network means for controlling the communication with a telematic network, a keyboard controller, a mouse controller and a video controller. A keyboard, a mouse, and a monitor are connected to the respective controllers.
  • the acquisition means of the data obtained by the microarray scanning are connected to the collector by means of an interface port (ITF) .
  • ITF interface port
  • a program (PRG) that is loaded to the work memory during the running stage, and a respective database are stored within the hard disk.
  • the program (PRG) is distributed over one or more CR-ROMs for being installed onto hard disk.
  • the processing system has a different structure, for example if it consists of a central unit to which the various terminals are connected, or a telematic computer network (such as Internet, Intranet, VPN) , if it has other units (such as a printer) , etc.
  • a telematic computer network such as Internet, Intranet, VPN
  • the program is provided on floppy disk, pre-loaded onto hard disk, or stored on any other substrate that can be computer-read, sent to a user's computer by means of a telematic network, transmitted by a radio or more generally is provided in any form that can be directly loaded in the work memory of a user's computer.

Landscapes

  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Engineering & Computer Science (AREA)
  • Biotechnology (AREA)
  • Biophysics (AREA)
  • Molecular Biology (AREA)
  • Theoretical Computer Science (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Evolutionary Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Medical Informatics (AREA)
  • Genetics & Genomics (AREA)
  • Physiology (AREA)
  • Analytical Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Probability & Statistics with Applications (AREA)
  • Chemical & Material Sciences (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present invention relates to a method for calculating the probability of contracting tumoral or degenerative pathologies in populations of healthy- individuals. The present invention further relates to a method for determining the prognosis on the progression of the disease in patients suffering from tumoral or degenerative pathologies. Particularly, the method comprises the following phases : (A) selecting a set of oncogenic and tumour suppressor genes characteristic of a given tumoral pathology and assigning said genes the relative level of expression in the tumoral pathology; (B) determining in the individual being examined or said population of individuals per each one of the genes selected according to step A, if this gene is expressed or unexpressed and assign each one of said genes the relative level of expression; calculating a gene evolution simulation with a random simulation algorithm, thereby obtaining a probability value of developing said tumoral or degenerative pathology.

Description

DESCRIPTION
METHOD FOR CALCULATING THE PROBABILITY OF DEVELOPING TUMORAL OR DEGENERATIVE DISEASES.
The present invention relates to a method for calculating the probability of contracting tumoral or degenerative pathologies in populations of healthy individuals . The present invention further relates to a method for determining the prognosis on the progression of the disease in patients suffering from tumoral or degenerative pathologies, as well as a prognosis on the effectiveness of a therapy adopted for said pathology.
The method being the object of the present invention is based on the evidence that a purely clinic approach is not sufficient in the oncology field. It can be thus stated that the prognostic determination of the disease along with the probability of contracting a tumoral pathology by healthy patients generally elude forms of control and prevision that are really reliable.
At the heart of this evidence there is the consideration that the genetic network that determines each structural and functional aspect of a living body is a complex, and thus chaotic, dynamic system. A chaotic system, in the long term, does not exhibit observable regularity or recurrent configurations. The behaviour of a system, even a simple one, can be so sensitive to initial conditions as to make the final result uncertain; in other words, the distinctive characteristic of chaotic systems is that they exhibit a considerable sensibility to initial conditions .
The behaviour of a chaotic system results from the combination of many ordered behaviours, none of which prevails under normal conditions. It has been demonstrated that, by suitably perturbating a chaotic system it may be guided to follow one of the regular behaviours thereof. The chaotic systems are deterministic, i.e. when two almost identical systems are pushed or guided by the same signal, they will produce the same "output" , even if it is impossible to know which is the "output" .
The complex chaotic systems such as biologic systems are however subjected to a further phenomenon that may be called anti-chaos. In fact, several very disordered systems spontaneously "crystallize" into a highly ordered state. There are many complex systems in biology: from the thousands of genes that self-organize within a cell to the cell network and molecules mediating the immune response, and the like. Among these there is the genome self-regulating system, which is a good example of how the anti-chaos can govern gene development and regulation. Cell types differ because varied forms of genetic activity are active in them, not because they have different genes. In a genome, the genes mutually regulate their activity either directly or by means of gene products. The coordinated behaviour of this system is at the heart of cell differentiation and integration of genes, including those involved in oncogenesis.
Mathematical models help researchers understand the characteristics of these complex systems. The first step is to couple a binary variable, which can be either active or inactive, to the behaviour of each element in the system (a gene, in the instant case) .
The rules that govern these mutual interaction systems are stochastic Boolean networks. The mathematical models applied in this study for building a system of simulation of gene networks involved in the oncogenesis are the so-called autonomous stochastic Boolean NK networks, where N is the number of genes involved, each of which has K inputs. These networks are autonomous because all the inputs comes from inside the system. A stochastic Boolean network has a finite number of states . A system must therefore eventually reenter a state that it has previously encountered, and as the system is deterministic, it will repeat the same state sequence as it did before and will indefinitely cycle repeatedly through the same cycle of states until a perturbation occurs in the system. These cycles are called the dynamic attractors of the network.
There are two types of perturbations, minimal and structural perturbations. A minimal perturbation occurs when a binary element temporarily flips to its opposite state (in the binary language, 0 sets to 1 or vice versa) . If this event does not push the network outside its original basin of attraction (homeostasis cycle) , it will eventually return to its starting cycle; but if the network is pushed into a different basin of attraction
(for example, a new cell cycle) , its behaviour will be definitely changed: the network will flow into a new state cycle and will adopt a different recurrent pattern of behaviour (new homeostasis) . To make a clinical example, a tissue that turns from an acute to a chronic inflammatory state is a case where a shift to a new homeostasis is achieved. The attractors have various levels of stability in relation with minimal perturbations. Some can neutralize any single perturbation, others can withstand only a few, still others are destabilized by any perturbation. A structural perturbation is a permanent mutation in the connections or in the Boolean functions of a network
(a logical operator "or" switches to "and" and vice versa) . It is generated by exchanging the inputs of two elements and networks have a varying level of stability against them. Ordered systems can become chaotic and vice versa (normal to neoplastic or neoplastic to normal tissue) .
These minimal or structural alterations are caused by the inherent possibility of error existing in any type of system, from simplest machines to computers, and of course biological systems.
The idea at the heart of the present invention is that the behaviour of DNA polymerase has been found to be substantially similar to a computer, which processes the data and provides an output based on determined inputs. In fact, DNA polymerase is a nanomachine, which is capable of copying a DNA template into a new strand. Computer laws can be thus applied to the genome regulation and evolution system, which is the foundation of several phenomena such as, inter alia, the oncogenesis process. An error in the genome produces signals in the cell which result in different biological responses, such as the stop of the cycle, reparation of DNA, cell death, or uncontrolled proliferation. Controls are usually carried out by an opposite action of two gene systems. The entire lifespan of a cell can be seen as a planned execution of proliferation, stop-quiescence, differentiation and death programs, each corresponding to the activation or inactivation of specific packages of genes .
During DNA replication, the DNA polymerase can make a mistake; or the DNA polymerase can properly operate, but the DNA to be replicated may be altered. Independently of the origin of this error, the cell that will inherit this DNA will have an altered length of DNA. If the portion of altered DNA codes for a protein, the thus produced protein may be, in turn, either inactive or even fatal.
In view of the above, the onset of a tumour can be related to the failure of a (also unknown) process, for a predetermined number of times. In the organism, within a tissue, cells reproduce by duplicating and by performing a careful numerical control system on their population. This control acts such that the number of cell individuals is maintained constant and does not grow according to second powers as mitotic duplication would require. This mechanism is the result of complicated gene expression and its failure can be considered as an event governed by statistical laws; it is thus possible to assign a finite probability other than zero to the birth of a neoplastic cell . In a simplified model, the tissue is divided into independent generation channels. After a tumour has generated within a generation channel, the evolution of this channel is exclusively characterized by neoplasia; at a macroscopic level, a tumour appears in time t only when the number of neoplastic generation channels is higher than a predetermined number. The model considers the continuous expression of, at tissue level, at least three genes and simulates the general state of the tissue by monitoring the number of errors that have occurred in the duplication of the genes involved. Synergies are taken into consideration, if present.
Practically, cell simulation is merely the creation of several simultaneously evolving subroutines, their evolution may influence the evolution of the entire system or not. The error can be associated with a determined abnormal value appearing in a subroutine. This abnormal value may be either neglected by the system
(irrelevant error or point mutation) , or it may be read.
In the latter case, there are three possible consequences: 1) the error is corrected by a mechanism (gene check point) which triggers ad hoc (reversible error) ; 2) another subroutine provides an abnormal value that cancels the first abnormality (reversible error) ; 3) the abnormal value causes the block of all subroutines (fatal error) .
The simulation provides that N cell families exist, each being independent from the others and from the surrounding cell system. Each family evolves by generating only one daughter cell: this simplification takes into account the collective self-regulation system of the tissue, which maintains constant the population of the cells belonging to the tissue itself. As stated above, the evolution of a family is completely independent from the surrounding conditions, except for the time factor: the probability that a healthy cell can experience a mutation to a neoplastic cell increases over time. Therefore, the simulation is the calculation of the probability that within a certain time (i.e. after a certain number of mitotic divisions) at least one neoplastic cell will be formed. In addition to the above, the calculation of the probability that a number NT of tumoral cells can derive from N cell families at time T, i.e. the incidence of the tumour starting from time t=0. The incidence of a particular type of tumour is the percentage fraction of cell families that have a neoplastic progenies after the preset period of time.
Throughout the simulation, the maintenance of a normal state of a tissue derives from the synergic action by at least two families X and Y of tumour suppressor and oncogenic genes, respectively. The results depends on the number of genes. A tumour is generated when the gene controlling neoplastic proliferation stops operating properly. Consequently, the cell stops responding to growth inhibitor signals and acquires the capacity of proliferating in an uncontrolled manner.
It is a clinical evidence that, in the oncology field, and mainly in the case of tumours of the central nervous system, different patients apparently suffering from the same tumour have a completely different pathological progression, as in some cases the disease takes a particularly aggressive form. At present, there was no way to predict a tumour progression, and thus to decide whether the patient should be subjected to the massive dose therapy, which is more dangerous though indispensable in extreme cases, or milder though safer therapies. The present invention is directed to provide a method for assessing the prognosis of a tumour pathology, which is very useful in order to select each time the most proper therapy.
The various forms taken by a tumour are clearly the result of different patterns of molecular abnormalities . It is thus essential to be capable of determining which genetic abnormalities have occurred in every single case. The tool allowing such assessment is the microarray, which uses the so-called DNA chips. This known technology allows identifying the differences in the activities exhibited by the hundreds of genes expressed by the tumoral cells upon the diagnosis.
As stated above, the DNA chip is used to analyze the level of gene expression in a tissue sample and consists of a layer of single-strand molecules of DNA (the so- called probes) that is immobilized on a support (glass, synthetic material or silicon) . The DNA chip uses the DNA property of complementary pairing its bases. For example, the base A (adenine) can be paired only to base T
(thymine) of the opposite strand, whereas the base C (cytosine) can be paired only to base G (guanine) . In the DNA chip, each type of probe (either an entire gene or a shorter DNA sequence) is immobilized at a well-precise point in the support grid. The sample of tissue to be examined is treated such as to isolate the mRNA thereof and to obtain the corresponding cDNA in labelled form (fluorescent label or the like) . The microarray is then incubated with the thus-obtained cDNA solution, then it is separated and suitably analyzed, such as by means of laser scanning. The cDNAs of the genes expressed in the sample bind to the immobilized DNA of the corresponding gene, which was in a predetermined position, as stated above .
The raw data are computer-processed and translated into a colour-coded map. From this map, one can thus identify which genes are active (expressed) in the examined tissue and which are "switched off" (unexpressed) .
As stated above, the DNA chip contain thousands of genes on an individual microarray, which allows carrying out a broad parallel gene assessment, by particularly identifying which genes are expressed in the various types of tumour tissue as compared with healthy tissue. These devices further allow assessing the expression level of a determined gene using a sample of tissue containing a known number of cells and assessing the percentage of cells in that tissue which express the gene being investigated. A gene is said to be expressed when it is transcribed to a molecule of mRNA and subsequently translated into a protein. This assessment may be carried out, for example, by means of fluorimetry. A mapping of the expressed genes will be thus obtained, the relative percentage of expression being associated with each of them.
The microarray methodology is widely known and therefore it will not be described below in greater detail. DNA chips are commercially available. DNA chip supplying Companies are for example: Affymetrix (Santa Clara, California) , Agilent Technologies (Palo Alto, California) , Perkin Elmer (Boston, MA) or Rosetta Inpharmatics (Kirkland, Washington) . Data of microarrays relating to patients suffering from various tumours are also available in the literature. These data are either in the form of tables or lists relative both to diseased tissue and healthy tissue of the several patients and are averaged among all the patients in the examined population.
As stated above, the method of the present invention allows providing, for a determined individual, a probability of developing a determined tumoral pathology. In a further embodiment, the method of the invention allows assessing the probability of developing a determined tumoral pathology in a population of individuals, either healthy or diseased, that are selected according to various criteria. These selection criteria may be, for example, geographic or racial origin, or a particular life and/or work environment; alternatively, the selection may be carried out age brackets or gender, or diet or lifestyle, or pre-existing genetic, family predisposition or particular pathologies to which the population is subjected. This list is merely exemplary and accordingly it cannot be considered as limiting the field of application of the invention.
By the term "probability of developing" a given tumoral pathology is meant, according to a first aspect, the probability that the healthy individual contracts a determined tumoral pathology. In a second aspect, the "probability of developing" is the probability that an individual who already suffers from a given tumoral pathology has a particular pathologic progression, for example he/she is experiencing an evolution from a tumoral form to a different one. The calculation of the "probability of developing" a given tumoral pathology will thus imply, in this case, the determination of the prognosis for the determined patient .
The method according to the present invention for determining the probability of developing a given tumoral pathology in an individual can comprise the following phases :
(A) selecting a set of oncogenic and tumour suppressor genes characteristic of a given tumoral pathology and assigning the relative level of expression in the tumoral pathology to said genes;
(B) determining in the individual being examined per each one of the genes in the set of oncogenic and tumour suppressor genes that have been selected according to phase A, whether this gene is expressed or unexpressed and assign each one of said genes the relative level of expression; (C) calculating a simulation of gene evolution with a random simulation algorithm, thereby obtaining a probability value of developing said tumoral pathology.
Similarly, the present invention relates to a method for determining the probability of developing a given tumoral pathology in a population of individuals comprising the following phases:
(A' ) selecting a set of oncogenic and tumour suppressor genes characteristic of a given tumoral pathology and assigning the relative level of expression in the tumoral pathology to said genes;
(B') determining in a population of individuals, per each one of the genes in the set of oncogenic and tumour suppressor genes selected according to phase A' , if this gene is expressed or unexpressed and assign each one of said genes the relative level of expression;
(C) calculating a simulation of a gene evolution with a random simulation algorithm, thereby obtaining a probability value of developing said tumoral pathology.
By the term "level of expression" is meant the percentage of DNA probes that, for a determined site, are coupled to cDNA of the tissue of the individual being examined, as results from the microarray data. In a population of individuals, the "level of expression" is given by the summation of the levels of expression of the single individuals in the population, divided by the number of individuals.
The first phase (phase (A) ) of the method according to the invention is selecting, for the tumoral pathology being the object of the method, a set of oncogenic and tumour suppressor genes that are characteristic of said tumoral pathology. By the term "oncogenic" is meant a gene deputed to controlling cell proliferation. By the term "tumour suppressor" is meant a gene deputed to controlling the suppression of cell proliferation. The selection of said genes is operated based on public or proprietary data and deriving from genetic studies on patients suffering from the specific tumoral pathology- being the object of the method. These studies on the genetic- inheritance of a determined population of individuals are conveniently carried out using various methods, such as PCR, PCR Real Time, immunohistochemistry, Western Blot and microarray. The microarray methodology is the most recent acquisition and the results thereof are published in the literature or provided on specific databases or however made available to the public. A further source for obtaining these data is the studies that have been carried out by the same entity that embodies the method of the invention, with the proviso that the results are statistically significant.
By the term "tumoral pathology" is meant both a type of tumour that has been identified as a function of the target organ, and a form of manifestation or evolution of the primary tumour being the object of a cell differentiation, or the like. Particularly, the tumoral pathology being the object of the present method will be both any tumoral pathology of which the probability of developing the latter by a healthy individual is desired to be assessed, and a possible evolutionary form of a primary tumour from which the individual being examined is suffering. By the term "evolutionary form" is meant an undifferentiated tumour that has developed from a primary tumour.
The phase A may also be carried out very upstream of the actuation of the method on a given individual, such as to build a table or database that contains the information required for each known type of tumoral pathology.
The selection of the set of oncogenic and tumour suppressor genes will be carried out by taking into account the level of expression of the gene in the determined pathology being examined, such as to have a substantially balance between the levels of expression of the oncogenic genes and the levels of expression of the tumour suppressor genes being selected. The phase A of the inventive method will comprise the following steps: i) providing statistic data of gene mapping and levels of expression thereof for the diseased tissue and healthy tissue of a population of patients of a determined tumoral pathology; ii) selecting from said statistic data obtained according to step i) a number n of oncongenes expressed in the diseased tissue of said population of patients of said tumoral pathology, wherein n is equal to or greater than l. iii) selecting from said statistic data obtained according to step i) a number m of tumour suppressors expressed in the healthy tissue of said population of patients of said tumoral pathology, wherein m is equal to or greater than 1, such that the summation of the levels of expression of said tumour suppressors in said healthy tissue substantially corresponds to the summation of the levels of expressions of said oncogenes in said diseased tissue, thereby obtaining a set of oncogenes and tumour suppressors characteristic of said tumoral pathology. iv) assigning to each of said oncogenes and tumour suppressors of the set of genes that is selected according to the steps ii) and iii) the relative level of expression in the diseased tissue of said population of patients of said tumoral pathology; v) creating a set of binary values from said set of genes in step iv) , according to the following principle: - assigning the value of 1 to each oncogene having a level of expression in said diseased tissue that is greater than or equal to a preset value, preferably 50%, and the value of 0 to each oncogene having a level of expression in said diseased tissue that is lower than said preset value;
- assigning the value of 1 to each tumour suppressor having a level of expression in said diseased tissue that is greater than or equal to a half of the relative level of expression in said healthy tissue and assigning the value of 0 to each tumour suppressor having a level of expression in said diseased tissue that is lower than a half of the relative level of expression in said healthy tissue.
By the term "diseased tissue" is meant the tissue taken from the tumour to which said population is subjected. By the term "healthy tissue" is meant, on the other hand, the tissue taken from a part of the organ of the same patient which is not affected by the tumour. The sampling of the healthy or diseased tissue is carried out according to methods and criteria that are well known to the physician skilled in the art. Preferably, n + m will range between 2 and 5, more preferably it will be 3. While, in fact, selecting a set comprising a large number of genes is desired, such as to obtain a more accurate determination, the number of possible combination would become so high that a long processing time would be required with normally available computer equipment.
For example, if the set of genes selected consists of 3 genes, i.e. one oncogene and two tumour suppressors, if the- oncogene has 80% level of expression in the tumoral tissue of the patient population, the tumour suppressors will be selected in the healthy tissue of the patient population such that the summation of the levels of expressions of the two tumour suppressors will have to be about 80%, for example 40% the first one and 40% the second one. The same will apply in the event that one tumour suppressor and two oncogenes are selected. If then we assume that the levels of expression, in the diseased tissue of said patient population, are 80% (as stated above) for the oncogene, 22% for the first tumour suppressor and 13% for the second tumour suppressor, respectively, the set of binary values for the selected genes will be (1 1 0) . In fact, the first tumour suppressor has a level of expression that is higher than half the relative level of expression in the healthy- tissue (>20%) , whereas for the second tumour suppressor the level of expression is lower than the respective half (<20%) . It should be noted that the set of binary values obtained according to the step v) of phase A of the inventive method will be the so-called "crash set", i.e. the set of binary values corresponding to the occurrence, in the patient, of the conditions for developing the tumoral pathology under examination, as will be better detailed below.
It should be also said that, instead of using the gene mapping of the healthy tissue of the patient population described above, a gene mapping may be similarly used which derives from a healthy population other than said population of patients, without however departing from the scopes of the present invention.
The second phase (phase B) of the method comprises the determination, in an individual being examined, of the level of expression of each of the genes in the set of oncogenic and tumour suppressing genes selected according to the phase A. Substantially, for each gene
(oncogene and tumour suppressor) belonging to the selected set it is determined whether this gene is expressed or not expressed in the individual, and the level of expression thereof. If the gene is expressed with an expression level greater than a determined threshold level, then it is assigned the binary number 1, whereas if it is not expressed, it is assigned the binary number 0.
The phase B of the inventive method thus comprise the following steps: i) carrying out a gene mapping with the relative levels of expression for a tissue sample of an individual being examined; ii) assigning, for said individual, the level of expression relative to each gene in the set of oncogenic and tumour suppressor genes selected according to the phase A; iii) creating a set of binary values from the set of genes obtained according to step ii) , according to the following principle: - assigning the number 1 to each oncogene that has a level of expression in the tissue sample of said individual being equal to or greater than 1A the relative level of expression in the diseased tissue of said patient population according to the phase A and the number 0 to each oncogene that has a level of expression in the tissue sample of said individual being lower than % the relative level of expression in the diseased tissue of said patient population according to phase A; i) - assigning the number 1 to each tumour suppressor that has a level of expression in the tissue of said individual being equal to or greater than 1A the relative level of expression in the healthy tissue of said patient population according to the phase A and the number 0 to each tumour suppressor that has a level of expression in the tissue sample of said individual being lower than % the relative level of expression in the healthy tissue of said patient population according to phase A. Preferably, the gene mapping is obtained according to the microarray technique with a DNA probe.
As stated above, the tissue sample of the individual being examined may be a sample of healthy tissue, or alternatively, a sample of diseased tissue in the case where the probability has to be determined that the tumoral pathology from which the individual is, in this case, already suffering will evolve to a different evolutionary form. The tissue sample may be taken by any organ, particularly a target organ of a possible tumoral pathology.
The phase C of the inventive method provides calculating a simulation of gene evolution, in order to obtain a development probability value for the tumoral pathology being examined.
The simulational calculation according to the phase C is carried out starting from the set of binary values that is obtained according to the phase B, step iii) , as will be set forth below.
As stated above, in a second embodiment, the probability of developing a tumoral pathology is calculated for a population of individuals (phases A' , B', and C). The phases A', B' and C are carried out in an entirely similar manner to phases A, B, and C for the single individual, except that the level of expression of the genes will be an average value across all the population individuals, such as discussed above. This embodiment will allow obtaining a global assessment of the probability that a population has of contracting a determined tumoral pathology. This population may be selected either according to the criteria listed above or according to different criteria, without however departing from the scope of the present invention.
The calculation of the probability of developing a tumoral pathology is carried out using a random simulation algorithm, the initial input thereof consisting of the set of binary values obtained for the individual (phase B, step iii) or the population of individuals (phase B'), wherein the level of expression assigned to each gene represents the probability that the initial binary value for that gene will be maintained.
The simulation provides a number ± of interactions, where ± is a preset number depending on the type of tissue (and thus, of cell) being considered. Generally, 1 will be equal to the number of mitotic cycles that the type of cell being examined performs in the time unit multiplied by the desired number of time units to be considered. For example, the calculation may be carried out over the entire lifespan of an individual, or just on a limited period of time which may correspond to age brackets in which the risk of contracting a particular tumour is the highest, in yet another case, particularly in order to predict the progression of the disease in a patient who is already suffering from a tumoral pathology, a period of time of a few months may be considered. This random simulation algorithm provides the random generation of numbers, using a number n + m of parent routines and one daughter routine. The parent routines are completely independent and completely blind relative to the surrounding environment. The behaviour of the daughter routine is, on the other hand, influenced by the behaviour of the parent routines on which it depends . Each parent routine is assigned a gene from the set of genes that has been initially selected.
The number n + m corresponds to the number of oncogenic (n) and tumour suppressor genes (m) being initially selected. In the preferred case of three genes, n+m will be equal to three.
Each parent routine has a probability of giving an error result which corresponds to the random drawing of a particular number on a preset sample. This situation simulates the randomness by which a gene can express in an erroneous manner. In the instant case, the error result occurs due to the generation of the binary set corresponding to the tumoral pathology being examined, which can be deduced by the initial data obtained on a patient population according to the phase A.
The method of random drawing of the binary value during the iterations is corrected by an algorithm, which takes into consideration the different probabilities for the various genes to maintain the value of 1 or 0 that was initially assigned thereto. This different probability can be related with the level of expression of the individual genes in the starting set of genes of the individual or population of individuals being examined. Preferably, this random simulation algorithm will be a so-called "genetic algorithm", such as described in Sankara K.Pal et al., Neuro Fuzzy Pattern Recognition: Methods in soft computing, appendix a genetic algorithms: basic principles, features, Winley - Interscience publication, 1999. The random simulation algorithm randomly generates, at each iteration, a parameter value (1 or 0) . This value may be thus either equal to the one of the preceding iteration or other than the latter. The algorithm has been changed such as to take into consideration the probability that the initial value is maintained, and accordingly it comprises the following steps:
- randomly generating a binary parameter value; - comparing this value with the value generated during the preceding iteration: if this value is the same, maintaining this value in the set of binary values; if this value is different, then - maintaining, in the set of binary values, the value of the preceding iteration to the M-th time that said different value is drawn, wherein M is the level of expression of the gene of which said binary value is representative; - replace, in the set of binary values, the new value for the value of the preceding iteration after the M-th time in which said different value is drawn.
For example, if a gene has 80% level of expression in the tissue of the individual being examined, for 80 iterations out of 100 in which a value is randomly drawn that is other than the initial one, the initial value will be maintained in the set of binary values, whereas the replacement will be carried out after the eightieth iteration. The behaviour of the daughter routine at each subsequent iteration is given by the following expression: behavi+1 = behavi + (∑n exprn x eoncogenepar n - ∑m exprm x par ΠIΛ
-suppressor in which esprn and esprm are the levels of expression of the genes in the individual being examined that are determined according to phase B and each value of par n and par m will depend on the relative parent routine. The value of par π or par m can be 0 or 1 and corresponds to the binary value assigned to the gene. The "behaviour" function may accordingly do not range between i-th and i+l-th iteration if a set of values of par n and par m all being equal to 1 or 0 is generated (situation of perfect balance between oncogenes and tumour suppressors) , or it may increase or decrease by a small amount when a value results to differ from the others. Generally, the "behaviour" function will vary between more or less than a basal value. Accordingly, an individual abnormal event is not catastrophic, because it alters the behaviour value only at the i-th step, the behaviour value will not thereafter change until a further abnormal event occurs. This further abnormal event may also correct or compensate the first abnormal event, thus resulting in the behaviour value varying between more or less than the initial value .
Only when a maximum or minimum value is achieved for the "behaviour" function, a catastrophic event will occur. This will occur when more than one abnormal event is added, such as when an oncogene is activated (a gene shifting from unexpressed to expressed form) and a tumour suppressor switches off (a gene shifting from expressed to unexpressed form) .
The selection of the maximum absolute value will be done by adopting the fixed delta principle of the "behaviour" function between basal value and i-th iteration. Conveniently, the absolute value of delta above which an event will be considered as catastrophic will be 30 relative to the initial value of the "behaviour" function. This means that, when at the i-th iteration the "behaviour" function assumes a value equal to the initial value plus 30 or more, the system will be clearly unbalanced towards the uncontrolled proliferation, and thus tumorigenesis. On the contrary, when the value is equal to the initial value less 30 or more, there will be an unbalance towards apoptosis. This condition will occur in the case of degenerative diseases, such as Alzheimer's disease.
The phase C thus provides a step of calculating a gene evolution simulation with a random simulation algorithm, such as stated above. The phase C of the inventive method further provides the following steps : i) obtaining a list of binary values for each of the parent routines relative to all iterations; ii) searching in said list the set of binary values corresponding to the maximum absolute value of the "behaviour" function and compare the same with the set of binary values typical of the tumoral pathology being examined according to phase A, step v) or according to phase A' ; iii) when said binary values being compared according to step ii) are equal, counting the number N of occurrences of the set of binary values corresponding to the maximum absolute value of the "behaviour" function; or iv) when said binary values compared according to step ii) are different, counting the number N of occurrences both of the set of binary values corresponding to the maximum absolute value of the "behaviour" function and of the set of binary values typical of the tumoral pathology being examined; v) calculating the probability P of developing said tumoral pathology by applying said algorithm:
P = N/number of iterations x 100 Preferably, the inventive method will also comprise the calculation of the probability that abnormal events may occur, which correspond to an intermediate situation between the equilibrium situation and the crash situation (development of the tumoral pathology) , but that are however indicative of a predisposition to developing said pathology. This is the typical pre-tumoral situation.
This calculation is carried out by means of a subroutine that provides the following steps : a) obtaining the set of binary values corresponding to the maximum absolute value of the "behaviour" function from the list in phase C7 step i) ; b) changing one value at a time of said set of binary values, thus generating n + m subsets of binary values; c) calculating the value of the "behaviour" function for each of said n + m subsets of binary values according to the formula
CDehaV
Figure imgf000033_0001
eguppressor and selecting the subsets of binary values for which the value of said "behaviour" function is intermediate between the basal value and said maximum value; d) counting the number Nabnormal of occurrences of the subsets of binary values selected according to the step c) in the list obtained according to phase C, step i) ; e) calculating the probability Pabnormai of occurrence of abnormal events by applying the following algorithm
Pabnormai = Nabnormai/number of iterations x 100 The method of the invention will be conveniently implemented by means of a computer.
A personal computer (PC) , for example, may be used which comprises a collector connecting processing means, for example a central processing unit (CPU) , to memory means that include, for example, a RAM work memory, a read-only memory (ROM) - which includes a base program for starting the computer, a magnetic hard disk, optionally a drive (DRV) for reading optical disks (CD- ROMs) , optionally a floppy disk read/write drive for. Furthermore, the computer can comprise a MODEM or other network means for controlling the communication with a telematic network, a keyboard controller, a mouse controller and a video controller. A keyboard, a mouse, and a monitor are connected to the respective controllers. The acquisition means of the data obtained by the microarray scanning are connected to the collector by means of an interface port (ITF) . A program (PRG) that is loaded to the work memory during the running stage, and a respective database are stored within the hard disk. Typically, the program (PRG) is distributed over one or more CR-ROMs for being installed onto hard disk.
The same considerations apply when the processing system has a different structure, for example if it consists of a central unit to which the various terminals are connected, or a telematic computer network (such as Internet, Intranet, VPN) , if it has other units (such as a printer) , etc. Alternatively, the program is provided on floppy disk, pre-loaded onto hard disk, or stored on any other substrate that can be computer-read, sent to a user's computer by means of a telematic network, transmitted by a radio or more generally is provided in any form that can be directly loaded in the work memory of a user's computer.

Claims

1. A method for calculating the probability of developing a tumoral or degenerative pathology in an individual or a population of individuals, comprising the following phases:
(A) selecting a set of oncogenic and tumour suppressor genes characteristic of a given tumoral pathology and assigning said genes the relative level of expression in the tumoral pathology;
(B) determining in the individual being examined or said population of individuals per each one of the genes in the set of oncogenic and tumour suppressor genes selected according to phase A, if this gene is expressed or unexpressed and assign each one of said genes the relative level of expression;
(C) calculating a simulation of a gene evolution with a random simulation algorithm, thereby obtaining a probability value of developing said tumoral or degenerative pathology.
2. The method according to claim 1, wherein said phase A comprises the following steps : i) providing statistic data of gene mapping and the levels of expression thereof for the diseased tissue and healthy . tissue of a population of patients of a determined tumoral pathology; ii) selecting from said statistic data obtained according to step i) a number n of oncongenes expressed in the diseased tissue of said population of patients of said tumoral pathology, wherein n is equal to or greater than 1. iii) selecting from said statistic data obtained according to step i) a number m of tumour suppressors expressed in the healthy tissue of said population of patients of said tumoral pathology, wherein m is equal to or greater than 1, such that the summation of the levels of expression of said tumour suppressors in said healthy tissue substantially corresponds to the summation of the levels of expressions of said oncogenes in said diseased tissue, thereby obtaining a set of oncogenes and tumour suppressors characteristic of said tumoral pathology. iv) assigning each of said oncogenes and tumour suppressors of the set of genes that is selected according to the steps ii) and iii) the relative level of expression in the diseased tissue of said population of patients of said tumoral pathology; v) creating a set of binary values from said set of genes in step iv) , according to the following principle: - assigning the value of 1 to each oncogene having a level of expression in said diseased tissue that is greater than or equal to a preset value, preferably 50%, and the value of 0 to each oncogene having a level of expression in said diseased tissue that is lower than said preset value;
- assigning the value of 1 to each tumour suppressor having a level of expression in said diseased tissue that is greater than or equal to a half of the relative level of expression in said healthy tissue and assigning the value of 0 to each tumour suppressor having a level of expression in said diseased tissue that is lower than half the relative level of expression in said healthy tissue.
3. The method according to claim 2 , wherein n + m will range between 2 and 5.
4. The method according to claim 3 , wherein n + m is 3.
5. The method according to any claim 1 to 4 , wherein said phase B comprises the following steps: i) carrying out a gene mapping with the relative levels of expression for a tissue sample of an individual being examined or a population of individuals; ii) assigning, for said individual or said population of individuals, the level of expression relative to each gene in the set of oncogenic and tumour suppressor genes selected according to the phase A; iii) creating a set of binary values from the set of genes obtained according to step ii) , according to the following principle: - assigning the number 1 to each oncogene that has a level of expression in the tissue sample of said individual or population of individuals being equal to or greater than % the relative level of expression in the diseased tissue of said patient population according to the phase A and the number 0 to each oncogene that has a level of expression in the tissue sample of said individual or population of individuals being lower than %. the relative level of expression in the diseased tissue of said population of patients according to phase A; - assigning the number 1 to each tumour suppressor that has a level of expression in the tissue of said individual or population of individuals being equal to or greater than % the relative level of expression in the healthy tissue of said population of patients according to the phase A and the number 0 to each tumour suppressor that has a level of expression in the tissue sample of said individual or said population of individuals being lower than % the relative level of expression in the healthy tissue of said population of patients according to phase A.
6. The method according to claim 5, wherein said gene mapping is obtained according to the microarray technique with DNA probe.
7. The method according to any claim 1 to 6, wherein said phase C uses a random simulation algorithm, having an initial input consisting of the set of binary values that is obtained according to the phase B, step iii, and wherein said random simulation provides a number 1 of iterations, wherein ± is a preset number that depends on the type of tissue; wherein the probability P of developing said tumoral or degenerative pathology is related with the number of times that said random simulation algorithm generates the binary set corresponding to the tumoral or degenerative pathology being examined, which can be deduced by the initial data obtained on a population of patients according to phase A.
8. The method according to claim 7, wherein 1 is equal to the number of mitotic cycles that the type of cell being examined performs in the time unit multiplied by a preset number of time units.
9. The method according to claim 7 or 8 , wherein said random simulation algorithm is a "genetic algorithm" .
10. The method according to any claim 7 to 9, wherein said random simulation algorithm provides the random generation of numbers, using a number n + m of parent routines and one daughter routine, each parent routine being assigned with a gene from the set of genes that has been initially selected, wherein the number n + m corresponds to the number of oncogenic (n) and tumour suppressor (in) genes that has been initially selected; said daughter routine being determined by the following expression: behavi+i = behavi + (∑n exprn x eonCogenepar n - ∑m exprm x •^suppressor / wherein behavi+i and behav± are the value of the "behaviour" function at i+l-th and i-th iterations, exprn and exprm are the levels of expression of the genes in the individual being examined that are determined according to phase B and each value of par n and par m can be 0 or 1 and corresponds to the binary value assigned to the gene by the random simulation algorithm in the corresponding parent routine.
11. The method according to any claim 7 to 10, wherein said random simulation algorithm comprises the following steps:
- randomly generating a binary parameter value: comparing this value with the value generated during the preceding iteration; if this value is the same, maintaining this value in the set of binary values; if this value is different, then
- maintaining, in the set of binary values, the value of the preceding iteration to the M-th time that said different value is drawn, wherein M is the level of expression of the gene of which said binary value is representative;
- replace, in the set of binary values, the new value for the value of the preceding iteration after the M-th time in which said different value is drawn.
12. The method according to any claim 7 to 11, wherein said method comprises the selection of a maximum absolute value of the "behaviour" function based on the principle of fixed delta of the "behaviour" function between basal value and i-th iteration, wherein said maximum absolute value will correspond to an absolute value of delta of 30 relative to the initial value (0-th iteration) of the "behaviour" function.
13. The method according to claim 12 , wherein said phase C further comprises the following steps: i) obtaining a list of binary values for each of parent routines relative to all the iterations; ii) searching in said list the set of binary values corresponding to the maximum absolute value of the "behaviour" function and compare the same with the set of binary values typical of the tumoral pathology being examined according to phase A, step v) ; iii) when said binary values being compared according to step ii) are equal, counting the number N of occurrences of the set of binary values corresponding to the maximum absolute value of the "behaviour" function; or iv) when said binary values that are compared according to step ii) are different, counting the number N of occurrences both of the set of binary values corresponding to the maximum absolute value of the "behaviour" function and of the set of binary values typical of the tumoral pathology being examined; vi) calculating the probability P of developing said tumoral pathology by applying the following algorithm:
P = N/number of iterations x 100
14. The method according to any claim 7 to 13 , said method further comprising the following steps : a) obtaining the set of binary values corresponding to the maximum absolute value of the "behaviour" function from the list in phase C, step i) ; b) changing one value at a time of said set of binary values, thus generating n + m subsets of binary values; c) calculating the value of the "behaviour" function for each of said n + m subsets of binary values according to the formula
JDenaV =, ^n exprn X eoncOgene ~ ^m ®2CP^"m X ©suppressor and selecting the subsets of binary values for which the value of said "behaviour" function is intermediate between the basal value and said maximum value; d) counting the number Nabnormai of occurrences of the subsets of binary values selected according to step c) in the list obtained according to phase C, step i) ; e) calculating the probability Pabπormai of occurrence of abnormal events by applying the following algorithm:
Pabπormai = Nabnormai/number of iterations x 100 15. A program (PRG) for electronic processor for carrying out the method according to any claim 1 to 14.
16. An electronic processor-readable drive comprising a program (PRG) for carrying out the method according to any claim 1 to 14.
PCT/IT2006/000763 2006-10-30 2006-10-30 Method for calculating the probability of developing tumoral or degenerative diseases. Ceased WO2008053499A1 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
PCT/IT2006/000763 WO2008053499A1 (en) 2006-10-30 2006-10-30 Method for calculating the probability of developing tumoral or degenerative diseases.

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/IT2006/000763 WO2008053499A1 (en) 2006-10-30 2006-10-30 Method for calculating the probability of developing tumoral or degenerative diseases.

Publications (2)

Publication Number Publication Date
WO2008053499A1 true WO2008053499A1 (en) 2008-05-08
WO2008053499A8 WO2008053499A8 (en) 2008-08-14

Family

ID=38191147

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IT2006/000763 Ceased WO2008053499A1 (en) 2006-10-30 2006-10-30 Method for calculating the probability of developing tumoral or degenerative diseases.

Country Status (1)

Country Link
WO (1) WO2008053499A1 (en)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030225718A1 (en) * 2002-01-30 2003-12-04 Ilya Shmulevich Probabilistic boolean networks

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20030225718A1 (en) * 2002-01-30 2003-12-04 Ilya Shmulevich Probabilistic boolean networks

Non-Patent Citations (3)

* Cited by examiner, † Cited by third party
Title
SEUNGCHAN KIM: "CAN MARKOV CHAIN MODELS MIMIC BIOLOGICAL REGULATION?", JOURNAL OF BIOLOGICAL SYSTEMS, WORLD SCIENTIFIC, SINGAPORE, SG, vol. 10, no. 4, 2002, pages 337 - 357, XP007902599, ISSN: 0218-3390 *
SHMULEVICH I ET AL: "Gene perturbation and intervention in Probabilistic Boolean Networks", BIOINFORMATICS, OXFORD UNIVERSITY PRESS, OXFORD,, GB, vol. 18, no. 10, October 2002 (2002-10-01), pages 1319 - 1331, XP002316689, ISSN: 1367-4803 *
SHMULEVICH I ET AL: "Probabilistic boolean networks: a rule-based uncertainty model for gene regulatory networks", BIOINFORMATICS, OXFORD UNIVERSITY PRESS, OXFORD,, GB, vol. 18, no. 2, October 2001 (2001-10-01), pages 261 - 274, XP002960824, ISSN: 1367-4803 *

Also Published As

Publication number Publication date
WO2008053499A8 (en) 2008-08-14

Similar Documents

Publication Publication Date Title
Ching et al. Cox-nnet: an artificial neural network method for prognosis prediction of high-throughput omics data
Markowetz et al. Non-transcriptional pathway features reconstructed from secondary effects of RNA interference
Ruths et al. The signaling petri net-based simulator: a non-parametric strategy for characterizing the dynamics of cell-specific signaling networks
Rhrissorrakrai et al. Understanding the limits of animal models as predictors of human biology: lessons learned from the sbv IMPROVER Species Translation Challenge
JP2004524604A (en) Expert system for the classification and prediction of genetic diseases and for linking molecular genetic and clinical parameters
Luan et al. Group additive regression models for genomic data analysis
US20240274226A1 (en) Molecular evaluation methods
Park et al. Identifying functional gene regulatory network phenotypes underlying single cell transcriptional variability
Cho et al. Bayesian hierarchical error model for analysis of gene expression data
Peng et al. A component overlapping attribute clustering (COAC) algorithm for single-cell RNA sequencing data analysis and potential pathobiological implications
Magouliotis et al. In-depth computational analysis reveals the significant dysregulation of key gap junction proteins (GJPs) driving thoracic aortic aneurysm development
US11435357B2 (en) System and method for discovery of gene-environment interactions
Jo et al. Interpretation of SNP combination effects on schizophrenia etiology based on stepwise deep learning with multi-precision data
Mohsenizadeh et al. Optimal objective-based experimental design for uncertain dynamical gene networks with experimental error
Lin et al. PMINR: pointwise mutual information-based network regression–with application to studies of lung cancer and Alzheimer’s disease
Lin et al. Theoretical and computational studies of the glucose signaling pathways in yeast using global gene expression data
Huang et al. Identification of immune-related signatures and pathogenesis differences between thoracic aortic aneurysm patients with bicuspid versus tricuspid valves via weighted gene co-expression network analysis
Kernfeld et al. Model-X knockoffs reveal data-dependent limits on regulatory network identification
Marchetti et al. A methodology based on MP theory for gene expression analysis
Wu et al. Single-cell Ca2+ parameter inference reveals how transcriptional states inform dynamic cell responses
Gu et al. Bayesian variable selection for high dimensional predictors and self-reported outcomes
Frost et al. An independent filter for gene set testing based on spectral enrichment
Yalamanchili et al. Quantifying gene network connectivity in silico: scalability and accuracy of a modular approach
WO2002095654A1 (en) Methods for predicting the activities of cellular constituents
Herrero et al. An approach to inferring transcriptional regulation among genes from large‐scale expression data

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 06821753

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 06821753

Country of ref document: EP

Kind code of ref document: A1