[Skip Global Navigation]

Customer Success Stories

Success Stories Home

Children’s Memorial Research Center

Using predictive analytics for DNA micro-array analysis

The situation

Each year, nearly 3,000 children in the USA are diagnosed with brain tumours. Almost half will die within five years, making it the most fatal cancer among children. If a child does survive a brain tumour, the long-term effects can be significant and include neurological disabilities, retardation and psychological problems.

Beyond surgery, successful treatments for pediatric brain tumours are rare. In fact, clinicians must often conduct experimental research and protocols to develop the treatments. Dr Eric Bremer, a former director of brain tumour research at Children’s Memorial Research Center in Chicago, is regarded as one of the leading scientists searching for a better way to treat brain tumours. One of his priority projects involved building a gene expression knowledge‑base for pediatric brain tumours.

The challenge

One of the first requisites for this cancer treatment database is accurately classifying the tumour. While the vast majority of pediatric brain tumours can be categorized as gliomas, ependymomas and medulloblastomas, altogether there are 12 or so tumour types including subtypes. Classification is very subjective because pediatric brain tumours are often difficult to distinguish by appearance, and there are few objective markers such as those found for other childhood cancers such as leukemia.

In addition to tumour type, it's important to classify a tumour's stage or grade. Cancers are generally stratified into four stages, from the most benign (stage I) to malignant (stage IV). Neuropathologists often find it difficult to distinguish between intermediate stages. However, the treatment for different stages can be drastically different, and incorrect staging can have dramatic consequences for the patient. For example, if a child with a stage two tumour is misdiagnosed with a more advanced grade, he would unnecessarily receive more aggressive treatment. This not only results in needless pain, but also could lead to long-term damage.

The solution

A key to accurate classification of pediatric brain tumours lies at the molecular level. Just as a skin cell and a liver cell vary in their gene expression patterns, the same is true for different tumours or tumour grades. Dr Bremer was able to capture these differences with gene expression micro-array experiments ­– but analysing these differences to identify and classify the tumour is a more formidable obstacle. Said Dr Bremer: "You can easily get 7,000 to 30,000 data points for each sample. The problem is how to make sense of them".

Dr Bremer ran into the solution to his problem at the Micro-array Data Analysis Conference, where he met two SPSS Inc. employees. Later, they consulted and worked with Dr Bremer to analyse his micro-array data using SPSS Inc's data mining solution, IBM SPSS Modeler*, the leading data mining workbench.

Traditionally used in customer oriented business intelligence applications, IBM SPSS Modeler is now being used in the life sciences to study micro-array data.

There is no one right way to analyse the data; you want to try several ways. What's so great about IBM SPSS Modeler is that it's a workbench. It's easy to try different methodologies in order to find a method, or combination of methods that works best.

Dr Eric Bremer
Former Director of Brain Tumor Research
Children's Memorial Hospital, Chicago

The power of IBM SPSS Modeler

Dr Bremer was won over by IBM SPSS Modeler for one simple reason: "it worked." Dr Bremer combined his own data with that of the publicly available Pomeroy et. al. data set (Nature vol. 415, pp. 436-442, 2002), resulting in a total of 133 tumour samples from the six major pediatric brain tumour types.

IBM SPSS Modeler successfully classified these tumours with greater than 95 percent accuracy. While these samples were well described pathologically, they served as a test case that bodes well for future pediatric brain tumour classification, especially of difficult-to-classify tumours. In addition, sub-classification of gliomas and medulloblastomas appeared to be very reliable.

The key was the software’s flexibility. Before using IBM SPSS Modeler, Dr Bremer used up to five different software packages to analyse micro-array gene expression data. IBM SPSS Modeler replaced all of those packages.

"There is no one right way to analyse the data; you want to try several ways," said Dr Bremer. "What's so great about IBM SPSS Modeler is that it's a workbench. It's easy to try different methodologies in order to find a method, or combination of methods that works best."

Dr Bremer used two of IBM SPSS Modeler’s predictive techniques, artificial neural networks and decision trees, to analyse and classify his data. Information from one model complemented the other. While the neural network resulted in a more accurate classification, it didn't show how it actually accomplished the classification. The decision tree, however, showed precisely how the tumours were classified, and revealed potential gene markers that characterised certain cancers. Once these markers are validated, labs without micro-array technology could use this information to develop antibodies against the gene product (protein) as an alternate means of diagnosis.

Just as important as the model is the data that it is based on. Micro-array data analysis presents a number of challenges, given the small number of samples and large number of genes. Dr Bremer's data set had 133 samples, but nearly 7,000 variables. The Micro-array IBM SPSS Modeler Application Templates are based on real-world experiences to overcome these challenges.

It's also important that differences in gene expression values genuinely reflect biological variation and not artificial differences introduced during sample preparation. IBM SPSS Modeler can help assure this is the case by processing the data quality parameters along with the samples. It also simplifies the task of feeding data into the model. Before using IBM SPSS Modeler, Dr Bremer had to organise the expression values in a specific format depending on the requirements of the analysis package. Dr Bremer now skips this step because IBM SPSS Modeler streams can automatically prepare and accept gene expression data directly from his database – a huge time saving.

More is better

Classification using data obtained from gene expression micro-array experiments is just part of a bigger picture. Expression measures RNA levels and not the final gene products. Dr. Bremer's database will eventually incorporate clinical, pathological and biochemical information to provide as complete and accurate a picture as possible. In addition, patient outcome information will be added so that the database can reveal which treatments work best with brain tumours sharing genetic and pathological characteristics. This is the point at which bio-informatics transition to ‘biomedical informatics’.

Dr Bremer also realises the critical need to share data and knowledge. To this end, he plans to use IBM SPSS Modeler streams via a Web server so that any pediatric brain tumour researcher can submit files from their own micro-array experiments for analysis, and then receive predictive responses. IBM SPSS Modeler will enable Dr Bremer's data to be combined with data from other researchers. In turn, the added data can be used to update and add value to the original dataset: the more samples his database contains, the more accurate it will be – and the more lives that will be saved.

The bigger database will also help reveal the multitude of genes that are involved in tumour development. According to Nature, in 2001 alone more than 21,000 articles on characterising, diagnosing and treating malignancies were published. Dr Bremer will use SPSS Inc. software to sift through this literature and extract patterns that, for example, when combined with data from his tumour database, might help him evaluate prime drug targets that would form the basis for a cure for cancer. This, after all, is Dr. Bremer's ultimate goal.

Interested in data mining for medical applications? Download the Children’s Memorial Research Center PDF here

*IBM SPSS Modeler, formerly called Clementine®, is part of SPSS Inc.’s Predictive Analytics Software portfolio.

back to top