Search NASA⌕ Search

SEARCH · Search NASA

Results for “medical data architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

A Phenome-Wide Association Study of genes associated with COVID-19 severity reveals shared genetics with complex diseases in the Million Veteran Program

The study aims to determine the shared genetic architecture between COVID-19 severity with existing medical conditions using electronic health record (EHR) data. We conducted a Phenome-Wide Association Study (PheWAS) of genetic variants associated with critical illness (n = 35) or hospitalization (n = 42) due to severe COVID-19 using genome-wide association summary data from the Host Genetics Initiative. PheWAS analysis was performed using genotype-phenotype data from the Veterans Affairs Million Veteran Program (MVP). Phenotypes were defined by International Classification of Diseases (ICD) codes mapped to clinically relevant groups using published PheWAS methods. Among 658,582 Veterans, variants associated with severe COVID-19 were tested for association across 1,559 phenotypes. Variants at the ABO locus (rs495828, rs505922) associated with the largest number of phenotypes (n rs495828 = 53 and n rs505922 = 59); strongest association with venous embolism, odds ratio (OR rs495828 1.33 (p = 1.32 x 10 –199 ), and thrombosis OR rs505922 1.33, p = 2.2 x10 -265 . Among 67 respiratory conditions tested, 11 had significant associations including MUC5B locus (rs35705950) with increased risk of idiopathic fibrosing alveolitis OR 2.83, p = 4.12 × 10 –191 ; CRHR1 (rs61667602) associated with reduced risk of pulmonary fibrosis, OR 0.84, p = 2.26× 10 –12 . The TYK2 locus (rs11085727) associated with reduced risk for autoimmune conditions, e.g., psoriasis OR 0.88, p = 6.48 x10 -23 , lupus OR 0.84, p = 3.97 x 10 –06 . PheWAS stratified by ancestry demonstrated differences in genotype-phenotype associations. LMNA (rs581342) associated with neutropenia OR 1.29 p = 4.1 x 10 –13 among Veterans of African and Hispanic ancestry but not European. Overall, we observed a shared genetic architecture between COVID-19 severity and conditions related to underlying risk factors for severe and poor COVID-19 outcomes. Differing associations between genotype-phenotype across ancestries may inform heterogenous outcomes observed with COVID-19. Divergent associations between risk for severe COVID-19 with autoimmune inflammatory conditions both respiratory and non-respiratory highlights the shared pathways and fine balance of immune host response and autoimmunity and caution required when considering treatment targets.

59 BASIC BIOLOGICAL SCIENCES↗

Know Your Space: Inlier and Outlier Construction for Calibrating Medical OOD Detectors

This software offers methods and functions for training calibrated out-of-distribution detectors for medical image classification tasks. It includes functionalities for training, synthesizing data augmentations, calibration, and out-of-distribution detection. Developed using PyTorch, this software is compatible with standard neural network architectures used for imaging data. Additionally, it provides capabilities to compute evaluation metrics for assessing the performance and quality of the detectors.

Narayanaswamy, VivekSivaraman↗

Genetics of varicose veins reveals polygenic architecture and genetic overlap with arterial and venous disease

Varicose veins represent a common cause of cardiovascular morbidity, with limited available medical therapies. Although varicose veins are heritable and epidemiologic studies have identified several candidate varicose vein risk factors, the molecular and genetic basis remains uncertain. Here we analyzed the contribution of common genetic variants to varicose veins using data from the Veterans Affairs Million Veteran Program and four other large biobanks. Among 49,765 individuals with varicose veins and 1,334,301 disease-free controls, we identified 139 risk loci. We identified genetic overlap between varicose veins, other vascular diseases and dozens of anthropometric factors. Using Mendelian randomization, we prioritized therapeutic targets via integration of proteomic and transcriptomic data. Finally, topological enrichment analyses confirmed the biologic roles of endothelial shear flow disruption, inflammation, vascular remodeling and angiogenesis. Further, these findings may facilitate future efforts to develop nonsurgical therapies for varicose veins.

59 BASIC BIOLOGICAL SCIENCES↗

Incubating advances in integrated photonics with emerging sensing and computational capabilities

As photonic technologies grow in multidimensional aspects, integrated photonics holds a unique position and continuously presents enormous possibilities for research communities. Applications include data centers, environmental monitoring, medical diagnosis, and highly compact communication components, with further possibilities continuously growing. Herein, we review state-of-the-art integrated photonic on-chip sensors that operate in the visible to mid-infrared wavelength region on various material platforms. Among the different materials, architectures, and technologies leading the way for on-chip sensors, we discuss the optical sensing principles that are commonly applied to biochemical and gas sensing. Our focus is on passive optical waveguides, including dispersion-engineered metamaterial-based structures, which are essential for enhancing the interaction between light and analytes in chip-scale sensors. We harness a diverse array of cutting-edge sensing technologies, heralding a revolutionary on-chip sensing paradigm. Our arsenal includes refractive-index-based sensing, plasmonics, and spectroscopy, which forge an unparalleled foundation for innovation and precision. Furthermore, we include a brief discussion of recent trends and computational concepts, incorporating Artificial Intelligence & Machine Learning (AI/ML) and deep learning approaches over the past few years to improve the qualitative and quantitative analysis of sensor measurements.

Jain, Sourabh (ORCID:0000000279923275)↗

Development of Analysis Methods that Integrate Numeric and Textual Equipment Reliability Data

Within the Light Water Reactor Sustainability (LWRS) program, the Risk-Informed Systems Analysis (RISA) Pathway is performing collaborative research on the development and deployment of technologies designed to assist operating nuclear power plants (NPPs) to reduce operating costs improve plant reliability and availability. One of the RISA research areas is focusing on the development of methods and tools designed to optimize plant operations (e.g., maintenance/replacement schedules, optimal maintenance postures for plant structures, systems, and components [SSCs]) in a manner that is more cost effective than current approaches and makes better use of available SSC health data. The Risk-Informed Asset Management (RIAM) project targets this research area by creating a direct bridge between component equipment reliability (ER) data and system engineer decision making regarding maintenance activity scheduling and component aging management. In this respect, one challenge that NPP system engineers are facing is that the amount of ER data being continuously generated is not only extremely large in size, but it comes in different forms: textual (e.g., condition or maintenance reports) and numeric (e.g., generated by monitoring systems). All these data elements provide them with valuable insights and information regarding: 1) the discovery of anomalous behaviors or degradation trends, 2) the identification of the possible causes behind such behaviors/trends, and 3) the prediction of their direct consequences. However, several challenges have proved to be roadblocks to this process. While some of these challenges are technical in nature (i.e., data are often distributed over several physical servers/databases), others are conceptual in nature: data elements come in different formats (e.g., numeric or textual), and measured values have different scales (e.g., vibration spectra and oil temperature). The activities performed by the RIAM project during FY23 directly tackles the need to simultaneously integrate the analysis of ER data in all its forms, numeric and textual. Note that such task has never been performed before due to the complexity of the systems under consideration but, most importantly, because of the technical challenges behind the harmonization of ER data formats and the lack of adequate computational methods to analyze them. Our approach borrows ideas and concepts from the medical field where integration of several data sources is vital to assist medical practitioners to perform correct diagnosis and indicate optimal treatments. In our view a NPP asset is equivalent to a patient in a medical context. The main difference is the complexity of a human body is a magnitude more complex when compared to typical assets commonly present in NPPs (e.g., centrifugal pumps, or motor operated valves). This simplifies our first requirement when analyzing heterogenous ER data formats: to put data into “context”. Context is here intended as the additional piece of information that is needed by ER data analysis tools to understand what these data elements are referring to, i.e., which king of knowledge they are generating. In our context, this knowledge can be translated into models that capture the form and functional architecture of assets/systems, their dependencies, and how they interact. These models actually emulate the knowledge that that NPP system engineers possess about assets and systems; this is their key of success when analyzing ER data, their challenge is ability to handle large amount of data. Here, we employ model-based system engineering (MBSE) models of systems and assets to represent and capture their architecture and functional, i.e. cause-effect, relations. Then, ER data elements are processed by identifying first of all which elements of the developed MBSE elements they are referring to. For numeric ER data this task is fairly easy since it is possible to precisely pinpoint what MBSE elements the corresponding sensor are observing (e.g., bearing temperature of a centrifugal pump). Task is much harder for textual data since the information contained in issue or maintenance reports needs to “be understood” by a computational tool. Here we called this process as “knowledge extraction”. Once again, we borrow the experience in the medical field where methods to extract knowledge from textual data have been developed in the past decade. The missing element for us is the availability of a complete dictionary of NPP related concepts (in addition to the MBSE models presented earlier) that can put “text into context”. In FY23, such dictionary has been developed along with all the computational elements required for knowledge extraction. Lastly, once numeric and textual ER data elements have been processed and “understood”, then the last step is the discovery of possible cause-effect relations among them. This is performed by observing if a logical connection through the MBSE models exists, and if the

97 MATHEMATICS AND COMPUTING↗

Intracardiac Electrical Imaging using the 12-lead ECG: A Machine Learning Approach using Synthetic Data

Current state-of-the-art techniques for non-invasive imaging of cardiac electrical phenomena require voltage recordings from dozens of different torso locations and anatomical models built from expensive medical diagnostic imaging procedures. Here this study aimed to assess if recent machine learning advances could alternatively reconstruct electroanatomical maps at clinically relevant resolutions using only the standard 12-lead electrocardiogram (ECG) as input. To that end, a computational study was conducted to generate a dataset of over 16000 detailed cardiac simulations, which was then used to train neural network (NN) architectures designed to exploit both spatial and temporal correlations in the ECG signal. Analysis over a validation set showed average errors in activation map reconstruction below 1.7 msec over 75 intracardiac locations. Furthermore, phenotypical patterns of activation and the morphology of the activation potential were correctly reconstructed. The approach offers opportunities to stratify patients non-invasively, both retrospectively and prospectively, using metrics otherwise only available through invasive clinical procedures.

59 BASIC BIOLOGICAL SCIENCES↗

Evaluation of Deep Learning Model Architectures for Point-of-Care Ultrasound Diagnostics

Point-of-care ultrasound imaging is a critical tool for patient triage during trauma for diagnosing injuries and prioritizing limited medical evacuation resources. Specifically, an eFAST exam evaluates if there are free fluids in the chest or abdomen but this is only possible if ultrasound scans can be accurately interpreted, a challenge in the pre-hospital setting. In this effort, we evaluated the use of artificial intelligent eFAST image interpretation models. Widely used deep learning model architectures were evaluated as well as Bayesian models optimized for six different diagnostic models: pneumothorax (i) B- or (ii) M-mode, hemothorax (iii) B- or (iv) M-mode, (v) pelvic or bladder abdominal hemorrhage and (vi) right upper quadrant abdominal hemorrhage. Models were trained using images captured in 27 swine. Using a leave-one-subject-out training approach, the MobileNetV2 and DarkNet53 models surpassed 85% accuracy for each M-mode scan site. The different B-mode models performed worse with accuracies between 68% and 74% except for the pelvic hemorrhage model, which only reached 62% accuracy for all model architectures. These results highlight which eFAST scan sites can be easily automated with image interpretation models, while other scan sites, such as the bladder hemorrhage model, will require more robust model development or data augmentation to improve performance. With these additional improvements, the skill threshold for ultrasound-based triage can be reduced, thus expanding its utility in the pre-hospital setting.

47 OTHER INSTRUMENTATION↗

From 2D to 4D: a containerized workflow and browser to explore dynamic chromatin architecture

Background Characterizing the physical organization of the genome is essential for understanding long-range gene regulation, chromatin compartmentalization, and epigenetic accessibility. Hi-C experiments generate two-dimensional (2D) genome-wide contact maps of chromatin interactions by capturing the spatial proximity between genomic loci, which reveal interaction frequencies but lack the spatial resolution needed to interpret the three-dimensional (3D) genome structure(s). Emerging evidence suggests that epigenetic regulation is closely linked to 3D genome architecture, and that structural changes over time (4D) drive key biological processes in development, disease, and environmental response. Thus, integrating 3D structure with functional data is critical for a more complete understanding of genome regulation. Previous work, most notably the 4DHiC chromosome modeling framework, has shown that physical multi-dimensional modeling approaches rooted in polymer physics and molecular dynamics can resolve these structures at biologically meaningful resolutions by integrating temporal Hi-C data with physical constraints to uncover dynamic chromosome reorganization. Thus, molecular dynamics simulations, constrained by Hi-C contact matrices, can resolve fine-scale structural changes and reveal functionally significant transitions in chromatin conformation. Results Herein, we present the 4D Genome Browser Workflow (4DGBWorkflow) and the 4D Genome Browser (4DGB). The algorithm is based on the 4DHiC method, and the containerized tool is an end-to-end workflow that can transform, filter, and view 4D epigenomics and chromatin datasets, allowing non-specialists to apply three-dimensional modeling principles to diverse datasets and experimental conditions. The software executes on a laptop running macOS, Linux or Windows. From input Hi-C files (.hic), the 4DGBWorkflow produces 3D reconstructions of chromosomes, integrates the reconstruction with track data (e.g., epigenetic marks, transcriptome profiles), and provides comparative visualization of the results in a single workflow. Conclusions The 4DGBWorkflow and 4D Genome Browser are open-source tools for comparative analysis and visualization of 4D chromosome datasets, including chromatin architecture and epigenomic signals. Automatic integration of Hi-C data with molecular dynamics democratizes the construction of time resolved 3D genome structures, simplifying complex simulations and data integration schemes.

3D Genome Browser↗

Dynamic molecular architecture of the synaptonemal complex

During meiosis, pairing between homologous chromosomes is stabilized by the assembly of the synaptonemal complex (SC). The SC ensures the formation of crossovers between homologous chromosomes and regulates their distribution. However, how the SC regulates crossover formation remains elusive. We isolated an unusual mutation in Caenorhabditis elegans that disrupts crossover interference but not SC assembly. This mutation alters the unique C terminal domain of an essential SC protein, SYP-4, a likely ortholog of the vertebrate SC protein SIX6OS1. We use three-dimensional stochastic optical reconstruction microscopy (3D-STORM) to interrogate the molecular architecture of the SC from wild-type and mutant C. elegans animals. Using a probabilistic mapping approach to analyze super-resolution image data, we detect changes in the organization of the synaptonemal complex in wild-type animals that coincide with crossover designation. We also found that our syp-4 mutant perturbs SC architecture. Our findings add to growing evidence that the SC is an active material whose molecular organization contributes to chromosome-wide crossover regulation.

59 BASIC BIOLOGICAL SCIENCES↗

EHR-BERT: A BERT-based model for effective anomaly detection in electronic health records

Objective: Physicians and clinicians rely on data contained in electronic health records (EHRs), as recorded by health information technology (HIT), to make informed decisions about their patients. The reliability of HIT systems in this regard is critical to patient safety. Consequently, better tools are needed to monitor the performance of HIT systems for potential hazards that could compromise the collected EHRs, which in turn could affect patient safety. In this paper, we propose a new framework for detecting anomalies in EHRs using sequence of clinical events. This new framework, EHR-Bidirectional Encoder Representations from Transformers (BERT), is motivated by the gaps in the existing deep-learning related methods, including high false negatives, sub-optimal accuracy, higher computational cost, and the risk of information loss. EHR-BERT is an innovative framework rooted in the BERT architecture, meticulously tailored to navigate the hurdles in the contemporary BERT method; thus, enhancing anomaly detection in EHRs for healthcare applications.Methods: The EHR-BERT framework was designed using the Sequential Masked Token Prediction (SMTP) method. This approach treats EHRs as natural language sentences and iteratively masks input tokens during both training and prediction stages. This method facilitates the learning of EHR sequence patterns in both directions for each event and identifies anomalies based on deviations from the normal execution models trained on EHR sequences.Results: Extensive experiments on large EHR datasets across various medical domains demonstrate that EHR-BERT markedly improves upon existing models. It significantly reduces the number of false positives and enhances the detection rate, thus bolstering the reliability of anomaly detection in electronic health records. This improvement is attributed to the model’s ability to minimize information loss and maximize data utilization effectively.Conclusion: EHR-BERT showcases immense potential in decreasing medical errors related to anomalous clinical events, positioning itself as an indispensable asset for enhancing patient safety and the overall standard of healthcare services. The framework effectively overcomes the drawbacks of earlier models, making it a promising solution for healthcare professionals to ensure the reliability and quality of health data.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Experimental Evaluation of Interference in 2.4 GHz Wireless Network

To attain automation across different applications, nuclear power plants are beginning to leverage advancements in wireless communication technologies. A “one-size-fits-all” solution cannot be applied since wireless technologies are selected according to application needs, quality of service requirements, and economic restrictions. To balance the trade-off between technical and economic requirements, a multi-band heterogeneous wireless network architecture is needed. Numerous wireless technologies including Wi-Fi, Zigbee, and Bluetooth share the 2.4 GHz industrial, scientific, and medical band. However, due to different channel access mechanisms and transmit power levels, and very importantly, uncoordinated use, coexistence of these devices in the same vicinity can cause interference and degradation in performance. This report provides the technical basis for understanding the coexistence of these wireless technologies through an experimental evaluation of their performance. This report investigates interactions encompassing variables such as transmission power level, distance between the devices, data rates, and the utilization of co-channel or adjacent channels. The results show that the operation of both Zigbee and Bluetooth is severely compromised when coexisting with Wi-Fi within the same frequency spectrum. On the other hand, the performance of Bluetooth is not impaired by Zigbee and vice versa unless there exists any external interference from Wi-Fi.

2.4 GHz↗

Sequencing the Genomes of the First Terrestrial Fungal Lineages: What Have We Learned?

The first genome sequenced of a eukaryotic organism was for Saccharomyces cerevisiae, as reported in 1996, but it was more than 10 years before any of the zygomycete fungi, which are the early-diverging terrestrial fungi currently placed in the phyla Mucoromycota and Zoopagomycota, were sequenced. The genome for Rhizopus delemar was completed in 2008; currently, more than 1000 zygomycete genomes have been sequenced. Genomic data from these early-diverging terrestrial fungi revealed deep phylogenetic separation of the two major clades—primarily plant—associated saprotrophic and mycorrhizal Mucoromycota versus the primarily mycoparasitic or animal-associated parasites and commensals in the Zoopagomycota. Genomic studies provide many valuable insights into how these fungi evolved in response to the challenges of living on land, including adaptations to sensing light and gravity, development of hyphal growth, and co-existence with the first terrestrial plants. Genome sequence data have facilitated studies of genome architecture, including a history of genome duplications and horizontal gene transfer events, distribution and organization of mating type loci, rDNA genes and transposable elements, methylation processes, and genes useful for various industrial applications. Pathogenicity genes and specialized secondary metabolites have also been detected in soil saprobes and pathogenic fungi. Novel endosymbiotic bacteria and viruses have been discovered during several zygomycete genome projects. Overall, genomic information has helped to resolve a plethora of research questions, from the placement of zygomycetes on the evolutionary tree of life and in natural ecosystems, to the applied biotechnological and medical questions.

59 BASIC BIOLOGICAL SCIENCES↗

FAIR Data and Interpretable AI Framework for Architectured Metamaterials (Final Report)

This research program established a transformative framework for the discovery and design of mechanical metamaterials, which are architected structures engineered to control physical phenomena like sound and vibration in ways natural materials cannot. To overcome the traditional reliance on trial-and-error, the project developed an interpretable Artificial Intelligence (AI) framework that moves beyond "black box" models to reveal the specific geometric patterns—such as "unit-cell templates"—that govern a material’s performance. A major breakthrough was the development of a hierarchical design method, which allows a single material to block vibrations across multiple frequency ranges simultaneously by layering patterns at different scales without them interfering with one another. This was further expanded to include irregular, graph-based designs that use spanning tree algorithms to ensure structural connectivity while allowing for customized, direction-dependent properties like stiffness and acoustic impedance. Beyond design, the project addressed the practicalities of real-world production by developing uncertainty quantification techniques that account for manufacturing defects and material variability, reducing the need for expensive physical testing by orders of magnitude. To speed up the discovery process, the team implemented Gaussian Process Regression and other surrogate models that provide accurate performance predictions at a fraction of the traditional computational cost. The AI-generated designs were successfully validated through fabrication of physical samples and wave propagation experiments, confirming their ability to accurately guide or reflect waves as predicted. By contributing these tools and high-quality FAIR benchmark datasets to the wider scientific community, this work provides a scalable foundation for advancing technologies in aerospace vibration control, medical imaging, and noise reduction.

36 MATERIALS SCIENCE↗

Identifying Heterogeneous Micromechanical Properties of Biological Tissues via Physics–Informed Neural Networks

The heterogeneous micromechanical properties of biological tissues have profound implications across diverse medical and engineering domains. However, identifying full-field heterogeneous elastic properties of soft materials using traditional engineering approaches is fundamentally challenging due to difficulties in estimating local stress fields. Recently, there has been a growing interest in data-driven models for learning full-field mechanical responses, such as displacement and strain, from experimental or synthetic data. However, research studies on inferring full-field elastic properties of materials, a more challenging problem, are scarce, particularly for large deformation, hyperelastic materials. Here, a physics-informed machine learning approach is proposed to identify the elasticity map in nonlinear, large deformation hyperelastic materials. This study reports the prediction accuracies and computational efficiency of physics-informed neural networks (PINNs) in inferring the heterogeneous elasticity maps across materials with structural complexity that closely resemble real tissue microstructure, such as brain, tricuspid valve, and breast cancer tissues. Further, the improved architecture is applied to three hyperelastic constitutive models: Neo-Hookean, Mooney Rivlin, and Gent. Furthermore, the improved network architecture consistently produces accurate estimations of heterogeneous elasticity maps, even when there is up to 10% noise present in the training data.

59 BASIC BIOLOGICAL SCIENCES↗

Introduction: Neuromorphic Materials

The explosive growth in data collection and the need to process it efficiently, as well as the desire to automate increasingly complex tasks in transportation, medical care, manufacturing, security and many other fields have motivated a growing interest in neuromorphic computing. Unlike the binary, transistorbased ON/OFF logic gates and separate logic and memory functionalities employed in digital computing, neuromorphic computing is inspired by animal brains that use interconnected synapses and neurons to perform processing, storage and transmission of information at the same location, while only consuming ~20 W or less of power. Motivated by the brain’s efficiency, adaptability, self-learning and resiliency qualities, neuromorphic computing can be broadly defined as an approach to processing and storing information using hardware and algorithms inspired by models of biological neural systems. Present research in neuromorphic computing encompasses approaches that vary significantly in their degree of neuro-inspiration, from systems that only incorporate features such as asynchronous, event-driven operation or use crossbar arrays of non-volatile memory (NVM) elements to accelerate deep neural networks (DNNs), to designs that embrace the extreme parallelism, sparsity, reconfigurability, adaptability, complexity and stochasticity observed in nervous systems. The term ‘neuromorphic’ computing is often credited to Carver Mead, who in the 1980s investigated Si-based analog electronics to replicate functions of the animal retina. Earlier important advances in this field include the work of Frank Rosenblatt, who proposed the concept of the perceptron, Bernard Widrow, who used this concept to build one of the first analog neural networks, the Adaline and many other researchers (see ref. 6 for an historical perspective on neuromorphic computing). With the recent increase in the use of artificial intelligence and large language models, and rising concerns over the associated energy costs, interest in neuromorphic hardware has expanded rapidly. According to some estimates, driven largely by the drastic growth in the training use of artificial intelligence (AI) models using the current computing architectures, the energy cost of computing is projected to reach the energy supply worldwide by 2045. Furthermore, while this is not a realistic outcome, it means that, if more efficient computing technologies are not developed -- soon -- the world will soon become one where demand for energy and market constraints limit the continued increase of societal access to AI and cloud services from data centers. Data centers used for training and use of these models consume hundreds of terawatt hours of electricity, already past 4% of the US electricity demand.

Circuits↗

SLAB: simultaneous labeling and binding affinity prediction for protein–ligand structures

Machine learning models are often used as scoring functions to predict the binding affinity of a protein–ligand complex. These models are trained with limited amounts of data with experimentally measured binding affinity values. A large number of compounds are labeled inactive through single-concentration screens without measuring binding affinities. These inactive compounds, along with the active ones, can be used to train binary classification models, while regression models are trained using compounds with binding affinities only. However, the classification and regression tasks are often handled separately, without sharing the learned feature representations. In this paper, we propose a novel model architecture that jointly performs regression and classification objectives, aiming to maximize data utilization and improve predictive performance by leveraging two complementary tasks. In our setup, the regression yields the binding affinity, whereas the classification task yields the label as active or inactive. We demonstrate our method using PDBbind, the standard 3D structure database, as well as a dataset of flavivirus protease compounds with binding affinity data. Our experiments show that the new joint training strategy improves the accuracy of the model, increasing applicability in various practical drug screening scenarios.

Biological and medical sciences↗

Synchrotron-based diffraction-enhanced imaging and diffraction-enhanced imaging combined with CT X-ray imaging systems to image seeds at 30 keV

Utilized the upgraded Synchrotron-based non-destructive Diffraction-enhanced imaging and Diffraction-enhanced imaging coupled with CT X-ray imaging systems to image the chickpea seeds, to enhance the contrast in plant root architecture, visibility of fine structures of root architecture growth and some aspects of physiology at acceptable level. DEI-CT images were acquired with 30 keV synchrotron X-rays. A series of DEI-CT slices were assembled together, to form a 3D data set. DEI-CT images explored more structural information and morphology. Noticed detailed anatomical, physiological observations, and contrast mechanisms. Furthermore, with these systems, some of the complex plant traits, root morphology, growth of laterals and subsequent laterals can be visualized directly.

36 MATERIALS SCIENCE↗

Strictly Enforcing Invertibility and Conservation in CNN-Based Super Resolution for Scientific Datasets

Abstract Recently, deep convolutional neural networks (CNNs) have revolutionized image “super resolution” (SR), dramatically outperforming past methods for enhancing image resolution. They could be a boon for the many scientific fields that involve imaging or any regularly gridded datasets: satellite remote sensing, radar meteorology, medical imaging, numerical modeling, and so on. Unfortunately, while SR-CNNs produce visually compelling results, they do not necessarily conserve physical quantities between their low-resolution inputs and high-resolution outputs when applied to scientific datasets. Here, a method for “downsampling enforcement” in SR-CNNs is proposed. A differentiable operator is derived that, when applied as the final transfer function of a CNN, ensures the high-resolution outputs exactly reproduce the low-resolution inputs under 2D-average downsampling while improving performance of the SR schemes. The method is demonstrated across seven modern CNN-based SR schemes on several benchmark image datasets, and applications to weather radar, satellite imager, and climate model data are shown. The approach improves training time and performance while ensuring physical consistency between the super-resolved and low-resolution data. Significance Statement Recent advancements in using deep learning to increase the resolution of images have substantial potential across the many scientific fields that use images and image-like data. Most image super-resolution research has focused on the visual quality of outputs, however, and is not necessarily well suited for use with scientific data where known physics constraints may need to be enforced. Here, we introduce a method to modify existing deep neural network architectures so that they strictly conserve physical quantities in the input field when “super resolving” scientific data and find that the method can improve performance across a wide range of datasets and neural networks. Integration of known physics and adherence to established physical constraints into deep neural networks will be a critical step before their potential can be fully realized in the physical sciences.

54 ENVIRONMENTAL SCIENCES↗