Cosmology from HSC Y1 weak lensing data with combined higher-order statistics and simulation-based inference
Not Available
SEARCH · Search NASA
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Spatiotemporal data has evolved in scale due to augmented use in cross-domain applications. Simultaneously, there is substantial growth in the availability of Geographic Information Systems (GIS) data provided by the United States Geological Survey (USGS) along with other federal, state, county, or local agencies through open-data portals and public access APIs. However, data availability does not equate with accessibility. Large-scale analyses and applications require robust, performant data management with co-location of data storage and computing. The insufficiency of data management infrastructure compels researchers to adopt ad hoc project- specific GIS data storage solutions (e.g., copying data to High-Performance computer file systems). As an ad hoc storage strategy does not scale, it hampers cross-domain analyses causing difficulty in data reuse and utilizing existing code bases. Furthermore, GIS data is complex and requires expertise to analyze and manipulate due to its intricate data structures and data-specific projection transformations. Despite the challenges, we recognize that derived GIS data products, e.g., satellite or LIDAR-based images, can be used in downstream applications such as AI by domain, but non-GIS experts. To address the data needs and overcome the challenges, we are working towards a GIS Data Platform focused on efficient data storage, data discovery and access, and an API to enable common workflows. We propose a knowledge-graph (KG) approach for data discovery, whereby datasets are semantically linked to higher- level constructs such as projects and research areas. The semantic data links enable researchers to explore datasets in a top-down approach by specifying relevant and meaningful terms (assists in finding hidden data). An advantage is that the nodes and edges in a knowledge graph create built-in semantic documentation. Deeper spatiotemporal connections between data sources can be encoded via Graph Neural Networks (GNN) (Zhang et al., 2021). The KG approach can be extended to integrate the data itself in a Virtual KG (VKG). Our work will derive inspiration from large-scale VKG efforts that have been undertaken or are currently underway as part of the OpenStreetMap project (Ding et al., 2021). For DOE Data Days, we share the proposed geospatial data platform hybrid (cloud/on-prem) architecture, our work-to-date on storing, retrieving, and transforming LiDAR and raster data relevant to two important NREL use-cases, including the Renewable Energy Potential (reV) Model, and present our proposal for a KG based data discovery engine.
The development, implementation, and evaluation of a new weakly coupled ocean data assimilation (WCODA) system for the fully coupled Energy Exascale Earth System Model version 2 (E3SMv2) utilizing the four-dimensional ensemble variational (4DEnVar) method are presented in this study. The 4DEnVar method, based on the dimension-reduced projection four-dimensional variational (DRP-4DVar) approach, replaces the adjoint model with the ensemble technique, thereby reducing computational demands. Monthly mean ocean temperature and salinity data from the EN4.2.1 reanalysis are integrated into the ocean component of E3SMv2 from 1950 to 2021 with the goal of providing realistic initial conditions for decadal predictions and predictability studies. The performance of the WCODA system is assessed using various metrics, including the reduction rate of the cost function, root mean square error (RMSE) differences, correlation differences, and model biases. Results indicate that the WCODA system effectively assimilates the reanalysis data into the climate model, consistently achieving negative reduction rates of the cost function and notable improvements in RMSE and correlation across various ocean layers and regions. Significant enhancements are observed in the upper ocean layers across the majority of global ocean regions, particularly in the north Atlantic, north Pacific, and Indian Ocean. Model biases in sea surface temperature and salinity are also substantially reduced. For sea surface temperature, cold biases in the north Pacific and north Atlantic are diminished by about 1–2 °C, and warm biases in the Southern Ocean are corrected by approximately 1.5–2.5 °C. In terms of salinity, improvements are observed with bias reductions of about 0.5–1 psu in the north Atlantic and north Pacific and up to 1.5 psu in parts of the Southern Ocean. The ultimate goal of the WCODA system is to advance the predictive capabilities of E3SM for subseasonal to decadal climate predictions, thereby supporting research on strategic energy-sector policies and planning.
ABSTRACT It has long been known that galaxy shapes align coherently with the large-scale density field. Characterizing this effect is essential to interpreting measurements of weak gravitational lensing, the deflection of light from distant galaxies by matter overdensities along the line of sight, as it also produces coherent galaxy alignments that we wish to interpret in terms of a cosmological model. Existing direct measurements of intrinsic alignments using galaxy samples with high-quality shape and redshift measurements typically use well-understood but sub-optimal projected estimators, which do not make good use of the information in the data when comparing those estimators to theoretical models. We demonstrate a more optimal estimator, based on a multipole expansion of the correlation functions or power spectra, for direct measurements of galaxy intrinsic alignments. We show that even using the lowest order multipole alone increases the significance of inferred model parameters using simulated and real data, without any additional modelling complexity. We apply this estimator to measurements of parameters of the non-linear alignment model using data from the Sloan Digital Sky survey, demonstrating consistent results with a factor of ∼2 greater precision in parameter fits to intrinsic alignments models. This result is functionally equivalent to quadrupling the survey area, but without the attendant costs – thereby demonstrating the value in using this new estimator in current and future intrinsic alignments measurements using spectroscopic galaxy samples.
Cloud base height (CBH) and cloud base vertical velocity (CBVV) are important variables that impact the overall climate in a region as they influence the formulation, longevity, and evolution of clouds. Retrieval of both parameters have long used ground instrumentation (e.g., Doppler lidar (DL), ground base radar); however, retrieving CBH from satellites is particularly challenging given that space-based instruments only observe cloud tops. In this manuscript, CBH is retrieved using a multi-linear regression equation, while CBVV used a random forests model. Both retrievals combine satellite and numerical weather prediction data. The satellite data used are the Visible Infrared Imaging Radiometer Suite imagery, while measurements of CBH and CBVV include DL and radiosonde data at the Southern Great Plains (SGP) Atmospheric Radiation Measurement observatory. Data from 83 summer days (May-August) in 2018–2021 featuring cumulus clouds forced by solar heating were examined and used to train the models, with years 2022–2023 used for validation. Various spatial domains were defined with one large (2.4° longitude by 2.0° latitude) SGP domain being split into smaller sections (smallest being 0.99° and 0.61° longitude and latitude respectably). CBH and CBVV values obtained from the DL as compared to the models show root mean square errors between 150 and 200 m, with CBVV values between 0.45 and 1 ms -1 . Finally, it was found that the CBH formulation performs well over all domains, while the CBVV retrievals become less accurate due to more turbulence being introduced into the observations as the number of DL stations decreases in the smaller domains.
Inverter-based resource (IBR) models are necessary to analyze modern power system stability and create effective control strategies. Modeling IBRs in converter-rich power systems is crucial, yet challenging due to the lack of commercial information on converter topologies and control parameters. This paper proposes novel convolutional neural network (CNN)–based data-driven techniques for modeling IBRs, addressing adaptability and proprietary concerns without requiring internal system physics knowledge. The proposed method is tested using real grid-tied commercial IBR transient data and demonstrates effectiveness and accuracy. Furthermore, the developed modeling approach is integrated and implemented in the open-source power distribution simulation and analysis tool, GridLAB-D, to illustrate the potentiality of dynamic analysis of large-scale power systems with high IBRs.
Science is and always has been based on data, but the terms ‘data-centric’ and the ‘4th paradigm’ of materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of artificial intelligence and its subset machine learning, has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.
Real-time detection and classification of distant objects is necessary for many national security applications. However, when objects are far from the sensor, they occupy only a small number of pixels in the captured video, limiting the amount of visual detail available for recognition. State-of-the-art classification methods typically rely on high-resolution (HR) video streams to capture characteristic object features, but obtaining such detail is challenging for distant objects that occupy only a few pixels. This motivates the development of video super-resolution (VSR) methods that enhance object classification by recovering fine details from low-pixel representations. Current VSR methods rely either on model-based optimization, which is interpretable but computationally expensive, or on learning-based approaches, which are efficient and high-performing but often lack flexibility and interpretability. In this report, we propose an end-to-end trainable unrolled VSR network, UVSRNet, which super-resolves each frame in a video by exploiting sub-pixel motion between neighboring low-resolution (LR) frames as well as incorporating high-frequency detail from previously super-resolved frames. In particular, by unrolling a plug-and-play (PnP) half-quadratic splitting (HQS) algorithm, we leverage a model-based data-fitting module alongside a learning-based autoregressive prior module. This combination yields a method that maintains the flexibility and interpretability of model-based methods while achieving the performance advantages of learning-based methods.
Distinguishing the most important features (e.g. proteins, metabolites, etc.) per group (e.g. control and treatment) is a critical challenge in feature-rich multi-omics experiments, especially in soil data. Traditional feature identification and ranking approaches, such as differential expression, are based on single omics and thus not directly translatable to multi-omics experiments. Here, 5 multi-omics integration models (DIABLO, JACA, MOFA, MultiMLP, and SLIDE) that were not explicitly built for soil data applications were tested using a soil-based multi-omics experiment. The data were obtained from an experimental setup of an autoclaved soil system inoculated with 8 bacteria and using chitin as the carbon source and including samples collected at 0- (control), 4-, 8-, and 12-weeks post-inoculation. The omics data included metaproteomics, 16S rRNA sequencing, and LC-MS/MS metabolomics (in positive and negative mode). Each multi-omics integration model was implemented, and top features were compared to differential univariate statistics per omic type, demonstrating that integration approaches cut the potential number of top features from 2957 identified by differential statistics to 13-224 (a 99.6% to 92.4% reduction). Interestingly, most top features across integration models were not shared; though, scaling and averaging ranks across models shared similar patterns. This work highlights the usefulness of multi-omics integration models in soil-based microbial studies and the power of using multiple integration models together to interpret results.
Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.
Travel behaviour and time-use data are two vital data sources for travel demand modelling. Travel behaviour is traditionally collected through household travel surveys, enhanced by using GPS-supported smartphone apps for passive location data collection. However, recruiting individuals willing to install these apps with sustained motivation to continue participation has been a critical challenge. This paper shares insights from a travel and time-use data collection procedure in Chicago and Sydney using the Fourstep app. Social media platforms were utilised as a solution to recruit participants in Chicago, where an international market research company failed to accomplish the task. This paper also discusses the challenges we faced and suggests ways to overcome them, offering valuable guidance to researchers in recruiting participants for smartphone application-based data collection. It also offers an analysis of travel, time-use, and travel-based multitasking behaviours based on the data collected from the Chicago and Sydney samples.
Advanced reactor designs including microreactors and small modular reactors will contribute to the clean production of cheap energy, and autonomous control for advanced reactors is an appealing option for reducing cost. However, there is a lack of industry experience applying autonomous control for advanced nuclear reactors. To accelerate the development and industry acceptance of autonomous control software for nuclear reactors, we aim to demonstrate autonomous control of the Purdue University research reactor (PUR-1) using INL-developed model predictive control (MPC) methods. To prepare for this demonstration, data-driven predictive models based on process data collected from PUR-1 have been developed and integrated with MPC and used to control a physics-based model of PUR-1. A data-driven dynamics model and a gated recurrent unit (GRU) network were both trained on process data from PUR-1. The dynamics model was shown to effectively control the reactor model with MPC when provided reactivity as a control variable but failed to control the model through the control rod positions. The GRU network produced more accurate predictions than the dynamics model when evaluated on operational data, and future work will include the evaluation of the GRU network in the controller.
Not Available
In many practical applications, the accuracy of neutron transport calculations depends on an estimate of the energy distribution provided in the source definition. While the energy distribution of neutrons emitted by spontaneous fission sources is well-defined, the spectra of neutrons produced in (α,n) reactions in compound mixtures, such as americium-beryllium (AmBe), depend on the microstructural properties of the source material, which may vary. Although the energy of neutrons produced in nuclear reactions can be calculated from first principles, such computations require detailed knowledge of the key characteristics of the source material, which are often unknown. Here, this study focuses on deriving the at-birth neutron energy distribution for an AmBe source by applying reverse transport calculations to the externally measured high-precision spectrum recommended in the latest revision of the international standard. The derived at-birth spectrum was validated through indirect energy-sensitive neutron emission rate measurements of AmBe sources calibrated at national primary metrology institutions. The resulting at-birth neutron energy distribution serves as essential input data for radiation transport modeling, providing a validated spectrum for broader experimental and computational applications.
Isotopic evidence of groundwater and stream water is frequently used to investigate water exchanges with groundwater. Monthly sampling of rain, stream water, and groundwater was conducted at Tims Branch watershed in South Carolina for the oxygen and hydrogen stable isotope (δ 2 H and δ 18 O) measurement, as well as pH and oxidation–reduction potential (ORP). Together with a mass balance perspective, it was determined that it takes a few weeks to one month for groundwater in the hyporheic zone to fully exchange with stream water. From hydrodynamic modelling, we show that substantial (up to 70 %) groundwater exchange occurs at gaining and losing sites. Groundwater exfiltration, i.e. inflow into stream water, contributes up to 4 % to stream water, with the remainder from upstream exfiltration. A 2–4 % per day renewal rate of adjacent groundwater would indirectly indicate a groundwater residence time in the order of half a month to a full month (assuming either a well-mixed case or large dispersion rate in pulse flow case), in agreement with a greatly reduced variability of δ 2 H and δ 18 O of groundwater compared to stream water and rain. This reduced variability of stable isotope signal from groundwater confirms our hypothesis that riparian groundwater mixing at Tims Branch is more of a mixed type rather than a pulse flow type. As a result, a monthly time scale is sufficient for groundwater to become anoxic at exit points into stream water resulting in the episodic production of natural organic matter- and iron-rich flocs upon oxidation.
Gigatonne-scale atmospheric carbon dioxide removal (CDR), alongside deep emission cuts, is critical to stabilizing the climate. However, some of the most scalable CDR technologies are also the most land intensive. Here, we examine whether adequate land resources exist in the contiguous United States to meet CDR targets when prioritizing grid emissions reduction, food production, and the protection of sensitive ecosystems. We focus on biomass carbon removal and storage (BiCRS) and direct air capture and storage (DACS) and show that suitable lands exceed the expected needs: 37.6 million hectares of land are available for BiCRS, resulting in 0.26 GtCO2 of CDR/year, and 34 million hectares are suitable for wind- and solar-powered DACS, resulting in 4.8 GtCO2 of CDR/year if facilities are co-located with geologic CO2 storage. We identify biomass and energy supply hotspots to meet CDR targets while ensuring land protection and minimizing land competition.
The Muon $g-2$ experiment at Fermilab aims to measure the muon magnetic moment anomaly, $a_{\mu}=(g-2)/2$, with a final accuracy of 0.14 parts per million (ppm). The experiment’s first result was published in 2021, based on data collected in 2018, and in 2023 a new result was published based on two more years of data taking, 2019 and 2020. The combination of the two results from Fermilab and the previous one from Brookhaven National Laboratory brought the uncertainty on the experimental measurement of $a_{\mu}$ to the unprecedented value of 0.19 ppm. This paper will present details about the improvements of statistical and systematic uncertainties on $a_{\mu}$ since the 2021 result.