Search NASA⌕ Search

Engineering topics

Sales, Ana Paula

Publications and source records attributed to Sales, Ana Paula.

Improving five-year survival prediction via multitask learning across HPV-related cancers

Oncology is a highly siloed field of research in which sub-disciplinary specialization has limited the amount of information shared between researchers of distinct cancer types. This can be attributed to legitimate differences in the physiology and carcinogenesis of cancers affecting distinct anatomical sites. However, underlying processes that are shared across seemingly disparate cancers probably affect prognosis. The objective of the current study is to investigate whether multitask learning improves 5-year survival cancer patient survival prediction by leveraging information across anatomically distinct HPV related cancers. Furthermore, data were obtained from the Surveillance, Epidemiology, and End Results (SEER) program database. The study cohort consisted of 29,768 primary cancer cases diagnosed in the United States between 2004 and 2015. Ten different cancer diagnoses were selected, all with a known association with HPV risk. In the analysis, the cancer diagnoses were categorized into three distinct topography groups of varying specificity. The most specific topography grouping consisted of 10 original cancer diagnoses differentiated by the first two digits of the ICD-O-3 topography code. The second topography grouping consisted of cancer diagnoses categorized into six distinct organ groups. Finally, the third topography grouping consisted of just two groups, head-neck cancers and ano-genital cancers. The tasks were to predict 5-year survival for patients within the different topography groups using 14 predictive features which were selected among descriptive variables available in the SEER database. The information from the predictive features was shared between tasks in three different ways, resulting in three distinct predictive models: 1) Information was not shared between patients assigned to different tasks (single task learning); 2) Information was shared between all patients, regardless of task (pooled model); 3) Only relevant information was shared between patients grouped to different tasks (multitask learning). Prediction performance was evaluated with Brier scores. All three models were evaluated against one another on each of the three distinct topography-defined tasks. The results showed that multitask classifiers achieved relative improvement for the majority of the scenarios studied compared to single task learning and pooled baseline methods. In this study, we have demonstrated that sharing information among anatomically distinct cancer types can lead to improved predictive survival models.

59 BASIC BIOLOGICAL SCIENCES↗

Multitask Recommender Systems for Cancer Drug Response

The problem we are currently trying to address is that there are many types of cancer drugs and many types of cancers and there is not always experimental data for a specific cancer type and cancer drug interaction. While there is a large possible set of feasible drug and cancer combinations, testing each pair is not realistic due to the high monetary cost of cell-based assays. Thus, this leaves researchers with a difficult choice of what drugs they should test on specific cancer types. This issue is known as the cold-start problem. Our focus is on developing recommender systems capable of addressing the cold-start problem as it relates to interaction between cancer types and cancer drugs. One of the most effective ways to address the cold-start problem is through large data analysis, however due to the cost prohibitive nature of cancer research the largest available data set size is the Genomics of Drug Sensitivity in Cancer with 494,973 genomic associations. To achieve optimal model performance on the cold-start problem, it is advantageous to employ multitask algorithms that are capable of transferring information between cancer datasets. The aim of this report is to draw from adaptations and state of the art developments in both algorithms for recommender systems and multitask learning to model the interaction between cancer cell lines and cancer drugs. Cancer cell lines are defined by the US National Cancer Institute as "cancer cells that keep dividing and growing over time, under certain conditions in a laboratory". This paper will focus on evaluating the performance of Neural Collaborative Filtering and Gaussian Processes, as well as their multitask adaptations, on cancer datasets from CCLE, NCI60, GDSC and CTRP. These methods will be evaluated on model performance in regression prediction but also in interpretability.

60 APPLIED LIFE SCIENCES↗

Functional and transcriptional characterization of complex neuronal co-cultures

Brain-on-a-chip systems are designed to simulate brain activity using traditional in vitro cell culture on an engineered platform. It is a noninvasive tool to screen new drugs, evaluate toxicants, and elucidate disease mechanisms. However, successful recapitulation of brain function on these systems is dependent on the complexity of the cell culture. In this study, we increased cellular complexity of traditional (simple) neuronal cultures by co-culturing with astrocytes and oligodendrocyte precursor cells (complex culture). We evaluated and compared neuronal activity (e.g., network formation and maturation), cellular composition in long-term culture, and the transcriptome of the two cultures. Compared to simple cultures, neurons from complex co-cultures exhibited earlier synapse and network development and maturation, which was supported by localized synaptophysin expression, up-regulation of genes involved in mature neuronal processes, and synchronized neural network activity. Also, mature oligodendrocytes and reactive astrocytes were only detected in complex cultures upon transcriptomic analysis of age-matched cultures. Functionally, the GABA antagonist bicuculline had a greater influence on bursting activity in complex versus simple cultures. Collectively, the cellular complexity of brain-on-a-chip systems intrinsically develops cell type-specific phenotypes relevant to the brain while accelerating the maturation of neuronal networks, important features underdeveloped in traditional cultures.

59 BASIC BIOLOGICAL SCIENCES↗

Generation and evaluation of synthetic patient data

Background: Machine learning (ML) has made a significant impact in medicine and cancer research; however, its impact in these areas has been undeniably slower and more limited than in other application domains. A major reason for this has been the lack of availability of patient data to the broader ML research community, in large part due to patient privacy protection concerns. High-quality, realistic, synthetic datasets can be leveraged to accelerate methodological developments in medicine. By and large, medical data is high dimensional and often categorical. These characteristics pose multiple modeling challenges. Methods: In this paper, we evaluate three classes of synthetic data generation approaches; probabilistic models, classification-based imputation models, and generative adversarial neural networks. Metrics for evaluating the quality of the generated synthetic datasets are presented and discussed. Results: While the results and discussions are broadly applicable to medical data, for demonstration purposes we generate synthetic datasets for cancer based on the publicly available cancer registry data from the Surveillance Epidemiology and End Results (SEER) program. Specifically, our cohort consists of breast, respiratory, and non-solid cancer cases diagnosed between 2010 and 2015, which includes over 360,000 individual cases. Conclusions: We discuss the trade-offs of the different methods and metrics, providing guidance on considerations for the generation and usage of medical synthetic data.

59 BASIC BIOLOGICAL SCIENCES↗