SEARCH · Search NASA
Results for “computer science”
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Sequential Inference of Hospitalization Electronic Health Records Using Probabilistic Models
In the dynamic hospital setting, decision support can be a valuable tool for improving patient outcomes. Data-driven inference of future outcomes is challenging in this dynamic setting, where long sequences such as laboratory tests and medications are updated frequently. This is due in part to heterogeneity of data types and mixed-sequence types contained in variable length sequences. In this work we design a probabilistic unsupervised model for multiple arbitrary-length sequences contained in hospitalization Electronic Health Record (EHR) data. The model uses a latent variable structure and captures complex relationships between medications, diagnoses, laboratory tests, neurological assessments, and medications. It can be trained on original data, without requiring any lossy transformations or time binning. Inference algorithms are derived that use partial data to infer properties of the complete sequences, including their length and presence of specific values. We train this model on data from subjects receiving medical care in the Kaiser Permanente Northern California integrated healthcare delivery system. The results are evaluated against held-out data for predicting the length of sequences and presence of Intensive Care Unit (ICU) in hospitalization bed sequences. Our method outperforms a baseline approach, showing that in these experiments the trained model captures information in the sequences that is informative of their future values.
Efficient Preparation of Dicke States
Here, we present an algorithm utilizing midcircuit measurement and feedback that prepares Dicke states with polylogarithmically many ancillae and polylogarithmic depth. Our algorithm uses only global midcircuit projective measurements and adaptively chosen global rotations. This improves over prior work that was only efficient for Dicke states of low weight or was not efficient in both depth and width. Our algorithm can also naturally be implemented in a cavity QED context using logarithmic time, zero ancillae, and atom-photon coupling scaling with the square root of the system size.
Predictive Model for Starlink Maritime Performance Using Multi-Horizon RandomForest
Low Earth orbit (LEO) satellite systems have become a crucial enabler of broadband access for maritime industries, where traditional networks are unavailable. However, the high mobility of LEO constellations and constantly changing weather conditions result in unpredictable link fluctuations, limiting the ability of maritime platforms to plan bandwidth usage proactively. To the best of our knowledge, no prior work has developed a short-term predictive model for maritime LEO connectivity using real experimental field measurements. This paper proposes a data-driven forecasting model that predicts future downlink throughput using multi-horizon RandomForest regression. The model is trained using real experimental coastal measurement data incorporating recent throughput history, network-layer indicators, and environmental variables. The proposed approach reduces mean absolute error by approximately 31% compared to a persistence baseline for 15-minute horizons. It maintains a measurable improvement at 30 minutes, despite increased stochasticity. These findings confirm that proactive bandwidth awareness is feasible on maritime platforms and can effectively support operational decisions such as adaptive streaming, routing, and resource scheduling. The performance gap between forecasting horizons also highlights the need for expanded offshore datasets to improve prediction robustness under harsher maritime environments.
23.2% efficient low band gap perovskite solar cells with cyanogen management
The role of thiocyanates in minimising organohalide diffusion into PEDOT:PSS whilst accelerating device degradation is identified and a route towards improving both device efficiency and stability is demonstrated.
What to Support When You’re Compressing
Over the last nearly 20 years, lossy compression has become an essential aspect of HPC applications’ data pipelines, allowing them to overcome limitations in storage capacity and bandwidth and, in some cases, increase computational throughput and capacity. However, with the adoption of lossy compression comes the requirement to assess and control the impact lossy compression has on scientific outcomes. In this work, we take a major step forward in describing the state of practice and by characterizing workloads. We examine applications’ needs and compressors’ capabilities across 9 different supercomputing application domains. We present 24 takeaways that provide best practices for applications, operational impacts for facilities achieving compressed data, and gaps in application needs not addressed by production compressors that point towards opportunities for future compression research.
Epitaxial Growth of 2D Core‐Crown SnS 2 /SnSe 2 Heterostructure Through Interfacial Modification with Polyvinylpyrrolidone
Abstract Developing generalized strategies for controlled synthesis of 2D heterostructures remains a significant challenge because the existing approaches often suffer from poor reproducibility and scalability. In this study, a solution synthesis approach for epitaxial core‐crown heterostructures with controlled band alignment, that overcomes these challenges is reported. Polyvinylpyrrolidone (PVP) is used as a structure‐directing agent to reduce lattice mismatch between SnS 2 and SnSe 2 (10‐10) surfaces and direct epitaxial growth of SnSe 2 crown on SnS 2 seed. Additionally, PVP adsorption to the basal plane prevents van der Waals stacking and stabilizes 2D heterostructures during synthesis. Driven by interfacial thermodynamics, the formation of the core‐crown heterostructure is highly reproducible and the size of the 2D heterostructure and relative areas of the core and the crown can be precisely controlled in a two‐step process by varying synthesis times for the seed and the crown. The identified growth pathway for 2D heterostructures can be generalized to other combinations of van der Waals materials to provide a platform for synthesizing micron‐size epitaxial heterostructures with a desired electronic structure for catalysis and microelectronics.
Atomate2: modular workflows for materials science
High-throughput density functional theory (DFT) calculations have become a vital element of computational materials science, enabling materials screening, property database generation, and training of “universal” machine learning models. While several software frameworks have emerged to support these computational efforts, new developments such as machine learned force fields have increased demands for more flexible and programmable workflow solutions. This manuscript introduces atomate2, a comprehensive evolution of our original atomate framework, designed to address existing limitations in computational materials research infrastructure. Key features include the support for multiple electronic structure packages and interoperability between them, along with generalizable workflows that can be written in an abstract form irrespective of the DFT package or machine learning force field used within them. Our hope is that atomate2's improved usability and extensibility can reduce technical barriers for high-throughput research workflows and facilitate the rapid adoption of emerging methods in computational material science.
Persistent Classification: Understanding Adversarial Attacks by Studying Decision Boundary Dynamics
ABSTRACT There are a number of hypotheses underlying the existence of adversarial examples for classification problems. These include the high‐dimensionality of the data, the high codimension in the ambient space of the data manifolds of interest, and that the structure of machine learning models may encourage classifiers to develop decision boundaries close to data points. This article proposes a new framework for studying adversarial examples that does not depend directly on the distance to the decision boundary. Similarly to the smoothed classifier literature, we define a (natural or adversarial) data point to be ( γ , σ)‐stable if the probability of the same classification is at least for points sampled in a Gaussian neighborhood of the point with a given standard deviation . We focus on studying the differences between persistence metrics along interpolants of natural and adversarial points. We show that adversarial examples have significantly lower persistence than natural examples for large neural networks in the context of the MNIST and ImageNet datasets. We connect this lack of persistence with decision boundary geometry by measuring angles of interpolants with respect to decision boundaries. Finally, we connect this approach with robustness by developing a manifold alignment gradient metric and demonstrating the increase in robustness that can be achieved when training with the addition of this metric.
Performance analysis and data reduction for exascale scientific workflows
Chimbuko is the first in situ, scalable, workflow-level performance analysis tool for trace-level analysis and visualization of application performance. This tool was developed by the Co-design Center for Online Data Analysis and Reduction and funded by the U.S. Department of Energy’s Exascale Computing Project. We provide a detailed description of Chimbuko’s architecture and illustrate our online and offline visualization with multiple use cases. We also present results for the deployment and scalability of the tool as applied to a high-energy physics workflow running at large scale on the Frontier supercomputer.
Integrative mapping reveals molecular features underlying the mechanism of nucleocytoplasmic transport
Nuclear pore complexes (NPCs) enable rapid, selective, and robust nucleocytoplasmic transport. To explain how transport emerges from the system components and their interactions, we used experimental data and theoretical information to construct an integrative Brownian dynamics model of transport through an NPC, coupled to a kinetic model of transport in the cell. The model recapitulates key aspects of transport for a wide range of molecular cargoes, including preribosomes and viral capsids. Our model quantifies how flexible phenylalanine-glycine (FG) repeat proteins create an entropic barrier to passive diffusion and how this barrier is selectively lowered in facilitated diffusion by the many transient interactions of nuclear transport receptors with the FG repeats. Selective transport is enhanced by “fuzzy” multivalent interactions, redundant FG repeat mass, coupling to the energy-dependent RanGTP concentration gradient, and exponential dependence of transport kinetics on the transport barrier. Our model will facilitate rational modulation of the NPC and its artificial mimics.
Data-Conforming Data-Driven Control: Avoiding Premature Generalizations Beyond Data
Data-driven and adaptive control approaches face the problem of introducing sudden distributional shifts beyond the distribution of data encountered during learning. Therefore, they are prone to invalidating the very assumptions used in their own construction. This is due to the linearity of the underlying system, inherently assumed and formulated in most data-driven control approaches, which may falsely generalize the behavior of the system beyond the behavior experienced in the data. This article seeks to mitigate these problems by enforcing consistency of the newly designed closed-loop systems with data and slowing down any distributional shifts in the joint state-input space. This is achieved through incorporating affine regularization terms and linear matrix inequality constraints to data-driven approaches, resulting in convex semi-definite programs that can be efficiently solved by standard software packages. We discuss the optimality conditions of these programs and then conclude this article with a numerical example that further highlights the problem of premature generalization beyond data and shows the effectiveness of our proposed approaches in enhancing the safety of data-driven control methods.
Efficient Client Selection in Federated Learning
Federated Learning (FL) enables decentralized machine learning while preserving data privacy. This paper proposes a novel client selection framework that integrates differential privacy and fault tolerance. The adaptive client selection adjusts the number of clients based on performance and system constraints, with noise added to protect privacy. Evaluated on the UNSW-NB15 and ROAD datasets for network anomaly detection, the method improves accuracy by 7% and reduces training time by 25 % compared to baselines. Fault tolerance enhances robustness with minimal performance trade-offs.
Adaptive Client Selection in Federated Learning: A Network Anomaly Detection Use Case
Federated Learning (FL) has become a ubiquitous approach for training machine learning models on decentralized data, addressing the myriad privacy concerns inherent in traditional centralized methods. However, the efficiency of FL depends on effective client selection and robust privacy preservation mechanisms. Inadequate client selection may lead to suboptimal model performance, while insufficient privacy measures risk exposing sensitive data. This paper proposes a client selection framework for FL that integrates differential privacy and fault tolerance. Our adaptive approach dynamically adjusts the number of selected clients based on model performance and system constraints, ensuring privacy through calibrated noise addition. We evaluate our method on a network anomaly detection use case using the UNSW-NB15 and ROAD datasets. Results show up to a 7% increase in accuracy and a 25% reduction in training time compared to FedL2P. Moreover, we highlight the trade-offs between privacy budgets and model performance, with higher privacy budgets reducing noise and improving accuracy. Our fault tolerance mechanism, while causing a slight performance drop, enhances robustness to client failures. Statistical validation using Mann-Whitney U tests confirms the significance of these improvements (p < 0.05).
Federated Learning for Efficient Condition Monitoring and Anomaly Detection in Industrial Cyber-Physical Systems
Detecting and localizing anomalies in cyber-physical systems (CPS) has become increasingly challenging as systems grow in complexity, particularly due to varying sensor reliability and node failures in distributed environments. While federated learning (FL) offers a foundation for distributed model training, existing approaches lack mechanisms to handle these CPS-specific challenges. This paper presents an enhanced FL framework that introduces three key innovations: adaptive model aggregation based on sensor reliability, dynamic node selection for resource optimization, and Weibull-based checkpointing for fault tolerance. Our framework enables reliable condition monitoring while addressing the computational and reliability challenges of industrial CPS deployments. Experiments on NASA Bearing and Hydraulic System Datasets demonstrate superior performance over state-of-the-art FL methods, achieving 99.5% AUC-ROC in anomaly detection and maintaining accuracy under node failures. Statistical validation using Mann-Whitney (U) test confirms significant improvements (p < 0.05) in both detection accuracy and computational efficiency across diverse operational scenarios.1
A Route to Design Novel Functional Peptides by Applying a Denoising Diffusional Model to mRNA Display Libraries
In vitro directed evolution techniques, such as mRNA display, enable peptide ligand discovery and optimization. However, physical libraries that rely on a genetic code can only search a small fraction of sequence space due to inherent biases in the genetic code and experimental limitations. To address this challenge, denoising diffusion implicit models (DDIMs) are applied to generate novel peptide ligands against B‐cell lymphoma extra‐large (Bcl‐x L ), a key cancer target. Starting with high‐throughput sequencing data from previous selections, a DDIM is trained to produce novel sequences with high affinity binding. Experimental validation confirms that most generated sequences are functionally equivalent to the original library members for Bcl‐x L binding and demonstrated comparable binding kinetics and affinity relative to the wildtype and nearest original neighbors. Importantly, this approach generated rare sequences not easily accessible via mutation and directed evolution. These results indicate that DDIMs can complement and expand directed evolution data, efficiently exploring underrepresented regions of sequence space. This approach provides a broadly applicable framework for accelerating ligand discovery and optimizing molecular properties across diverse targets.
High-Resolution Mapping of Photocatalytic Activity by Diffusion-Based and Tunneling Modes of Photo-Scanning Electrochemical Microscopy
Not Available
US Department of Energy, Office of Science, High-Performance Computing Facility: 2023 Operational Assessment Oak Ridge Leadership Computing Facility
The Oak Ridge Leadership Computing Facility (OLCF) was established to accelerate scientific discovery by providing world-leading computational performance and advanced data infrastructure. As a US Department of Energy (DOE) Office of Science user facility, the OLCF has managed the successful deployment and operation of a succession of leadership-class resources dedicated to open science. In addition to these resources, the OLCF staff continually strive to develop innovative processes and technologies, improve security, and empower users through effective allocation management and comprehensive user support and training. These efforts support the advancement of science by the OLCF users and benefit high-performance computing (HPC) facilities around the world. In calendar year (CY) 2023, the OLCF supported 1,676 users and 598 projects and exceeded all targets for user satisfaction. The facility received an average satisfaction score of 4.52 out of 5 on the annual user survey, and 94% of respondents reported a high satisfaction rate with the OLCF overall. Of the 3,619 user tickets submitted in CY 2023, OLCF staff resolved 97% within 3 business days. The facility opened Frontier to full scientific operations this year. Two projects conducted on Frontier received the Association for Computing Machinery (ACM) Gordon Bell Prize and the Gordon Bell Special Prize for Climate Modeling, and a third earned a nomination as a Gordon Bell Prize finalist. The facility’s previous flagship machine, Summit, gained new life and was extended through 2024 in part to help provide resources to the Integrated Research Infrastructure (IRI) projects and the National Artificial Intelligence Research Resource (NAIRR) pilot program. The facility instantiated an Advanced Computing Ecosystem testbed in part to support IRI workflows. OLCF made interactivity easier and more accessible to users than ever through tools like Jupyter notebooks and workflows.