CRISPR-CARB/nocap
Network Optimization and Causal Analysis of Perturb-seq (NOCAP) is a software package for causal inference of gene regulation networks using data from perturb-seq.
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Network Optimization and Causal Analysis of Perturb-seq (NOCAP) is a software package for causal inference of gene regulation networks using data from perturb-seq.
SAND2026-18878O The NN-OpInf tool is a PyTorch-based approach to operator inference that uses composable, structure-preserving neural networks to represent nonlinear operators. Operator inference is a machine learning method for inferring low-dimensional systems from data and polynomial models for system dynamics. However, many systems do not conform to polynomial structures, which NN-OpInf addresses by parameterizing operators with neural networks. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525.
The rapid expansion of artificial intelligence (AI) and machine learning is driving unprecedented electricity demand from data centers. It is predicted that by 2030, 90% of AI workloads will be inference-based, requiring interconnection of multiple low-latency edge data centers (<20 MW) sited closer to end users - often on already constrained distribution feeders. Although individually small, these loads can aggregate to large loads per feeder, straining infrastructure, creating multi-year interconnection delays, and driving up customer costs. This paper proposes a data center-focused grid-integration framework that combines feeder hosting capacity analysis with building energy efficiency, building load flexibility, and waste heat reuse to expand effective feeder and substation headroom. Such approaches can reduce interconnection delays, lower costs for ratepayers, and accelerate AI-ready infrastructure deployment.
Atmospheric retrievals are essential tools for interpreting exoplanet transmission and eclipse spectra, enabling quantitative constraints on the chemical composition, aerosol properties, and thermal structure of planetary atmospheres. The James Webb Space Telescope (JWST) offers unprecedented spectral precision, resolution, and wavelength coverage, unlocking transformative insights into the formation, evolution, climate, and potential habitability of planetary systems. However, this opportunity is accompanied by challenges: modeling assumptions and unaccounted-for noise or signal sources can bias retrieval outcomes and their interpretation. To address these limitations, we introduce a Gaussian process (GP)-aided atmospheric retrieval framework that flexibly accounts for unmodeled features and correlated noise in exoplanet spectra. We validate this method on synthetic JWST observations, and show that GP-aided retrievals reduce bias in inferred abundances and better capture model–data mismatches than traditional approaches. We also introduce the concept of mean squared error to quantify the trade-off between bias and variance, arguing that this metric more accurately reflects retrieval performance than bias alone. We then reanalyze the NIRISS/SOSS JWST transmission spectrum of WASP-96 b, finding that GP-aided retrievals yield broader constraints on CO 2 and H 2 O, possibly alleviating tension between previous retrieval results and equilibrium predictions. Our GP framework provides precise and accurate constraints while highlighting regions where models fail to explain the data. As JWST matures and future facilities come online, a deeper understanding of the limitations of both data and models will be essential, and GP-enabled retrievals like the one presented here offer a principled path forward.
Four post-irradiation heating tests of fuel compacts from the U.S. Advanced Gas Reactor (AGR)-3/4 irradiation experiment were completed. In addition to tristructural isotropic (TRISO)-coated driver fuel, each compact contained designed-to-fail (DTF) particles with fuel kernels coated only in pyrocarbon so as to simulate exposed kernels. Tests at 1600/1700°C, 1400°C, and 1200°C were performed to measure fission product release as a function of time and temperature. Silver release was highest in the 1200°C test, supporting the observation that silver release rates are highest in the 1100–1300°C range. Compared to tests of AGR-1 compacts with no exposed kernels, the Cs-134 and Kr-85 releases were noticeably higher in AGR-3/4. The exposed kernels’ contributions to Eu and Sr release are inconclusive, due to the difficulty in distinguishing among the combined effects of higher irradiation temperatures in these particular AGR-3/4 compacts, the presence of the DTF particles, and the Fuel Accident Condition Simulator (FACS) test temperatures. These data can be used to make inferences about fission product retention in exposed kernels as a function of time and temperature.
To enable an accurate determination of oscillation parameters, accelerator-based neutrino experiments require detailed simulations of nuclear interaction physics in the GeV regime. While substantial effort from both theory and experiment is currently being invested to improve the fidelity of these simulations, their present deficiencies typically oblige experimental collaborations to resort to empirical tuning of simulation model parameters. As the precision requirements of the field continue to become more stringent, machine learning techniques may provide a powerful means of handling corresponding growth in the complexity of future neutrino interaction model tuning exercises. To study the suitability of simulation-based inference (SBI) for this physics application, in this paper we revisit a tuned configuration of the GENIE neutrino event generator that was originally developed by the MicroBooNE collaboration. Despite closely reproducing the adopted values of four physics parameters when confronted with the tuned cross-section predictions as input, we find that our trained SBI algorithm prefers modestly different values (within MicroBooNE's assigned uncertainties) and achieves slightly better goodness-of-fit when inference is run on the experimental data set originally used by MicroBooNE. We also find that our trained algorithm can create a fair approximation of an alternative neutrino scattering simulation, NuWro, that shares only a subset of its physics model parameters with GENIE.
Modern scientific instruments operate under increasingly extreme constraints on bandwidth, latency, and power. Inference at the sensor edge determines experimental data collection efficiency by deciding which information to save for further analysis. Particle tracking detectors at the Large Hadron Collider exemplify this challenge: pixelated silicon sensors generate rich spatiotemporal ionization patterns, yet most of this information is discarded due to data-rate limitations. Concurrently, advancements in co-design tools provide rapid turn-around for incorporating machine learning into application-specific integrated circuits, motivating designs for particle detectors with new integrated technologies. We demonstrate that neural networks embedded in the front-end electronics can infer charged-particle kinematic parameters from a single silicon layer. We regress hit positions and incident angles with calibrated uncertainties, while satisfying stringent constraints on numerical precision, latency, and silicon area. Our results establish a path toward probabilistic inference directly at the edge, opening new opportunities for intelligent sensing in high-rate scientific instruments.
Network biology is an interdisciplinary field bridging computational and biological sciences that has proved pivotal in advancing the understanding of cellular functions and diseases across biological systems and scales. Although the field has been around for two decades, it remains nascent. It has witnessed rapid evolution, accompanied by emerging challenges. These stem from various factors, notably the growing complexity and volume of data together with the increased diversity of data types describing different tiers of biological organization. We discuss prevailing research directions in network biology, focusing on molecular/cellular networks but also on other biological network types such as biomedical knowledge graphs, patient similarity networks, brain networks, and social/contact networks relevant to disease spread. In more detail, we highlight areas of inference and comparison of biological networks, multimodal data integration and heterogeneous networks, higher-order network analysis, machine learning on networks, and network-based personalized medicine. Following the overview of recent breakthroughs across these five areas, we offer a perspective on future directions of network biology. Additionally, we discuss scientific communities, educational initiatives, and the importance of fostering diversity within the field. This article establishes a roadmap for an immediate and long-term vision for network biology.
Site characterization for underground injection and storage of gigatonne-scale CO₂ requires reliable and cost-effective methods to detect and characterize faults and fractures and to assess their stress state and fault activation potential. This is critical, as wastewater injection and disposal have been shown to activate faults and induce earthquakes, and CO₂ leakage remains a key concern for long-term storage. In this project, we developed seismic methods to detect and characterize large-scale sedimentary and crystalline basement faults and associated small-scale fractures below conventional seismic imaging resolution using multicomponent (9C) surface seismic data. Machine learning was used to automatically interpret large-scale faults, providing key information for estimating the maximum magnitude of potential induced earthquakes. High-fidelity imaging was achieved by exploiting redundancy across multiple elastic wave modes, where independent images from different modes and frequencies cross-validate each other. We also used our nonlinear signal comparison (NLSC) method for ground roll removal, improving data quality in complex near-surface conditions. The methods were validated using field data acquired in central Montana. Results show that basement faults extend into the sedimentary section and that small-scale fractures are widespread above the basement. The inferred stress orientation is consistent with regional stress data, and the estimated maximum induced earthquake magnitude is small (Mw ~2.3). The developed workflow provides a practical approach for fault and fracture characterization and for assessing induced seismicity and leakage risk. It is directly applicable to CO₂ storage site selection and to other subsurface systems.
Noise is a consistent problem for x-ray transmission images of High-Energy-Density (HED) experiments because it can significantly affect the accuracy of inferring quantitative physical properties from these images. We consider experiments that use x-ray area backlighting to image a thin layer of opaque material within a physics package to observe its hydrodynamic evolution. The spatial variance of the x-ray transmission across the system due to changing opacity serves as an analog for measuring density in this evolving layer. The noise in these images adds nonphysical variations in measured intensity, which can significantly reduce the accuracy of our inferred densities, particularly at small spatial scales. Denoising these images is thus necessary to improve our quantitative analysis, but any denoising method also affects the underlying information in the image. In this paper, we present a method for denoising HED x-ray images via a deep convolutional neural network model with a modified DenseNet architecture. In our denoising framework, we estimate the noise present in the real (data) images of interest and apply the inferred noise distribution to a set of natural images. These synthetic noisy images are then used to train a neural network model to recognize and remove noise of that character. We show that our trained denoiser network significantly reduces the noise in our experimental images while retaining important physical features.
During land model development, simulated carbon dynamics are often benchmarked against observational data sets to evaluate model performance. Functional relationship benchmarks are the relationship between a driving variable (e.g., temperature) and a response variable (e.g., ecosystem respiration) and are a promising tool for assessing model performance by evaluating modeled sensitivities to changing environmental conditions. However, observed functional relationships can be influenced by choices made during data collection and throughout the benchmarking process, impacting the inferred skill of land models. To avoid misrepresenting a model's true performance, it is necessary to systematically evaluate best practices when constructing functional relationship benchmarks. We developed a set of guidelines for constructing functional relationship benchmarks, considering the choice of data set, number of daily observations, temporal extent, and temporal resolution across Alaska and Canada over a 20-year period from 2001 to 2020. The temperature sensitivity of ecosystem respiration from observations, evaluated through an apparent Q 10 , is highly variable both spatially and as a result of the data processing approach applied in the benchmark formation. When benchmarking 13 models from the Warming Permafrost Model Intercomparison Project (WrPMIP), the range in inferred model skill is substantially impacted by the choices applied in constructing functional relationship benchmarks. The inferred performance of a given model is most sensitive to the number of daily observations and temporal extent, followed by choice of benchmark data set and temporal averaging. Results from this analysis can guide the development of consistent and robust functional relationships for future model evaluation studies.
The application of deep machine learning methods in astronomy has exploded in the last decade, with new models showing remarkably improved performance on benchmark tasks. Not nearly enough attention is given to understanding the models' robustness, especially when the test data are systematically different from the training data, or "out of domain." Domain shift poses a significant challenge for simulation-based inference, where models are trained on simulated data but applied to real observational data. In this paper, we explore domain shift and test domain adaptation methods for a specific scientific case: simulation-based inference for estimating galaxy cluster masses from X-ray profiles. We build datasets to mimic simulation-based inference: a training set from the Magneticum simulation, a scatter-augmented training set to capture uncertainties in scaling relations, and a test set derived from the IllustrisTNG simulation. We demonstrate that the Test Set is out of domain in subtle ways that would be difficult to detect without careful analysis. We apply three deep learning methods: a standard neural network (NN), a neural network trained on the scatter-augmented input catalogs, and a Deep Reconstruction-Regression Network (DRRN), a semi-supervised deep model engineered to address domain shift. Although the NN improves results by 17% in the Training Data, it performs 40% worse on the out-of-domain Test Set. Surprisingly, the Scatter-Augmented Neural Network (SANN) performs similarly. While the DRRN is successful in mapping the training and Test Data onto the same latent space, it consistently underperforms compared to a straightforward Yx scaling relation. These results serve as a warning that simulation-based inference must be handled with extreme care, as subtle differences between training simulations and observational data can lead to unforeseen biases creeping into the results.
Measurements of galaxy distributions at large cosmic distances capture clustering from the past. In this study, we use a cosmological model to translate these observations into the present-day galaxy distribution. Specifically, we reconstruct the 3D linear matter power spectrum at redshift z = 0 using Dark Energy Spectroscopic Instrument (DESI) Year 1 (DR1) galaxy clustering data and Cosmic Microwave Background (CMB) observations, assuming the ΛCDM model, and compare it to the result assuming the w 0 w a CDM model. Building on previous state-of-the-art methods, we apply Effective Field Theory (EFT) modelling of the galaxy power spectrum to account for small-scale effects in the 2-point statistics of galaxy data. Implementation of the EFT approach improves the modelling of the galaxy power spectrum, providing a more robust consistency test of the assumed cosmological model. By casting both CMB and galaxy clustering observations, spanning distinct redshift regimes, into k-space, we can identify discrepancies between the datasets of different redshifts, which would indicate potential inaccuracies in the assumed expansion history. While previous studies have shown consistency with ΛCDM, this work extends the analysis with higher-quality data to further test the expansion histories of both ΛCDM and w 0 w a CDM. Our findings show that both ΛCDM and w 0 w a CDM provide consistent fits to the linear matter power spectrum recovered from DESI DR1 data.
With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.
Smart grid technologies have rapidly become one of the largest and most comprehensive sources of data for the modern utility. For the most part, data streams are seen as an essential tool that enable utilities to carry their day-to-day business operations, but they also create the need for efficient and secure data management strategies. In the context of the smart grid, ensuring data privacy is becoming an increasing concern due to a combination of factors that range from shifts in operational paradigms and rapid technology evolution to changes in legislation. Furthermore, researchers have highlighted the risks associated with improperly protected energy records. For example, energy consumption data from homes could be used to infer the behaviors and habits of home occupants through activity recognition or user profiling (Fan, 2017), which may lead to unfair service pricing, targeted advertising, or other personal security violations. Similarly, Electric Vehicles’ (EVs) charging metadata could be used to reveal private information about the owner such as their payment methods, preferred charging stations, and other locational and timing information that could be used to reconstruct the vehicle owner’s behaviors. The privacy of user data, even when used for statistical analysis or machine learning training processes, also needs to be carefully considered, as an individual’s private traits may still be vulnerable if their inclusion/exclusion greatly impacts the result or could be linked to a public dataset through cross-reference. The breach of user privacy also has severe impacts for organizations that store, transmit, or work on the data in the form of diminishing the public’s trust in them while potentially incurring legal consequences (e.g., fines and suspensions under the European Union General Data Protection Regulation, Health Insurance Portability and Accountability Act, etc.). Because of these risks, several privacy-preserving mechanisms are available to help organizations comply with privacy legislations and prevent the unauthorized and malicious use of user data. In light of these concerns, this report focuses on performing a computational review of privacy-preserving mechanisms that have received a significant amount of interest in literature. It specifically focuses on 1) homomorphic encryption, 2) zero-knowledge proofs, 3) differential privacy, and 4) federated learning. It is worth noting that although many of the methods presented in this document rely on cryptographic primitives, their intent is not to provide perfect secrecy, but rather to enable users to maintain privacy, and thus they shall not be compared or equated to other constructs that are aimed to address cybersecurity constructs.
While typical validation and verification approaches focus on identifying the associations between data elements using statistical and machine learning methods, the novel methods in this paper focus instead on identifying causal relationships between data elements. Statistical and machine-learning-based approaches are strictly data-driven, meaning that they provide quantitative comparison measures between data sets without explicitly considering the hypotheses behind them. This can lead to the erroneous conclusion that, if two data sets are close enough, the models that generated them are similar. In addition, when experimental and simulated data differ to an extent that fails to meet the acceptance criteria, calibration techniques are used to tweak simulation model parameters to reduce the gap between the two types of data. This produces the false expectation that a simulation model will match reality. The methods presented in this paper move away from these strictly data-driven methods for validation and calibration toward more robust, model-driven methods based on causal inference. Causal inference aims to identify the possible mechanisms that might have generated data. Thus, this analysis targets the prediction of the effects when one (or more) of the identified mechanisms are altered. There are many approaches to identify, quantify, and illustrate causal relationships. For the scope of this paper, directed graphs are employed as causal models. If the directed graph lacks cycles, it is known as a directed acyclic graph. A node in such a graph represents an observed data element while a directed edge connecting two nodes represents a causal relationship between two variables. The developed causal methods are designed to extract causal models from simulation models and experimental data. Causal models capture the causal relationships between data elements (e.g., simulated and experimental data). In this context, validation and verification are performed by comparing causal models. The proposed approach does not only inform system analysts on how a simulation model matches real-world data, but also identifies elements of the simulation model that should be revised when discrepancies between simulation and experimental data are observed. Through these causal methods, analysts can identify the portion of the model equation(s) that are behind an edge connecting two variables. Hence, once the structural differences between causal models have been determined, model calibration can occur by changing only those model parameters that impact the identified causal relationships.
Photonic Doppler velocimetry (PDV) is an established technique for measuring the velocities of fast-moving surfaces in high-energy-density experiments. In the standard approach to PDV analysis, the short-time Fourier transform (STFT) is used to generate a spectrogram from which the velocity history of the target is inferred. The user chooses the form, duration, and separation of the window function. Here, in this study, we present a Bayesian approach to infer the velocity directly from the PDV oscilloscope trace, without using the spectrogram for analysis. This is clearly a difficult inference problem due to the highly periodic nature of the data, but we find that with carefully chosen prior distributions for the model parameters, we can accurately recover the injected velocity from synthetic data. We validate this method using PDV data collected at the STAR two-stage light gas gun at Sandia National Laboratories, recovering shock-front velocities in quartz that are consistent with those inferred using the STFT-based approach and are interpolated across regions of low signal-to-noise data. Although this method does not rely on the same user choices as the STFT, we caution that it can be prone to misspecification if the chosen model is not sufficient to capture the velocity behavior. Analysis using posterior predictive checks can be used to establish whether a better model is required, although more complex models come with additional computational cost, often taking more than several hours to converge when sampling the Bayesian posterior. We, therefore, recommend it be viewed as a complementary method to that of the STFT-based approach.
MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.