Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A novel method may reveal bulk metallic glass compressive ductility trends in high data rate nanoindentation

Recent methods allow novel amorphous alloy compositions to be rapidly manufactured at small scale; however, obtaining materials properties such as compressive ductility from these smaller specimens has remained a challenge. Here, we suggest a potential high-throughput nanoindentation method that may be able to rapidly characterize the relative compressive ductility between these alloys based on their serration characteristics. The properties of emergent serrations, when interpreted in a simple micromechanical stress relaxation model, may order these materials by their compressive plastic strain to failure. These results are consistent with the ordering obtained from compressed specimens as well as with model simulations, suggesting that this model may be broadly useful for interpreting compressive ductility from nanoindentation serrations. After it is validated on more materials, this new method will match the rapid pace of amorphous alloy development, thus allowing metallic glass properties to be fine-tuned for each application prior to scale prototyping.

36 MATERIALS SCIENCE↗

Advanced flip-coil system for magnetic field integral measurements of insertion devices

A novel flip-coil measurement system has been developed for the National Synchrotron Light Source II (NSLS-II) at Brookhaven National Laboratory. This paper describes the design, implementation, and commissioning of the new measurement bench, highlighting its key features, including improved mechanical stability, advanced data acquisition, and enhanced reproducibility. The system enables precise characterization of field integrals and multipole components, ensuring the optimal performance of Insertion Devices (IDs) before installation in the NSLS-II storage ring. The flip-coil system incorporates an innovative approach to minimize mechanical and electrical errors, which significantly improves the reproducibility of measurements. In addition, the system features a state-of-the-art data acquisition system that enables real-time monitoring and analysis, further enhancing the efficiency and accuracy of the measurement process. Furthermore, preliminary tests have demonstrated that the new system meets the stringent requirements for magnetic field characterization of advanced insertion devices, making it an essential tool for future ID commissioning and quality assurance at NSLS-II.

36 MATERIALS SCIENCE↗

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

ARPA-E Grid Optimization (GO) Competition Challenge 1

The ARPA-E Grid Optimization (GO) Competition Challenge 1, from 2018 to 2019, focused on the basic Security Constrained AC Optimal Power Flow problem (SCOPF) for a single time period. The Challenge utilized sets of unique datasets generated by the ARPA-E GRID DATA program. Each dataset consisted of a collection of power system network models of different sizes with associated operating scenarios (snapshots in time defining instantaneous power demand, renewable generation, generator and line availability, etc.). The datasets were of two types: Real-Time, which included starting-point information, and Online, which did not. Week-Ahead data is also provided for some cases but was not used in the Competition. Although most datasets were synthetic and generated by GRIDDATA, a few came from industry and were only used in the Final Event. All synthetic Input Data and Team Results for the GO Competition Challenge 1 for the Sandbox, Trial Events 1 to 3, and the Final Event along with problem, format, scoring and rules descriptions are available here. Data for industry scenarios will not be made public. Challenge 1, a minimization problem, required two computational steps. Solver 1 or Code 1 solved the base SCOPF problem under a strict wall clock time limit, as would be the case in industry, and reported the base case operating point as output, which was used to compute the Objective Function value that was used as the scenario score. The feasibility of the solution was provided by the Solver 2 or Code 2, which solves the power flow problem for all contingencies based on the results from Solver 1. This is not normally done in industry, so the time limits were relaxed. In fact, there were no time limits for Trial Event 1. This proved to be a mistake, with some codes running for more than 90 hours, and a time limit of 2 seconds per contingency was imposed for all other events. Entrants were free to use their own Solver 2 or use an open-source version provided by the Competition. Containers, such as Docker, were considered to improve the portability of codes, but none that could reliably support a multi-node parallel computing environment, e.g., MPI, could be found. For more information on the competition and challenge see the "GO Competition Challenge 1 Information" and "GO Competition Challenge 1 Additional Information" resources below.

ACOPF↗

High-speed Boulders and the Debris Field in DART Ejecta

On 2022 September 26 the Double Asteroid Redirection Test (DART) spacecraft collided with Dimorphos, the moon of the near-Earth asteroid 65803 Didymos, in a full-scale demonstration of a kinetic impactor concept. The companion Light Italian Cubesat for Imaging of Asteroids (LICIACube) spacecraft documented the aftermath, capturing images of the expansion and evolution of the ejecta from 29 to 243 s after the impact. We present results from our analyses of these observations, including an improved reduction of the data and new absolute calibration, an updated LICIACube trajectory, and a detailed description of the events and phenomena that were recorded throughout the flyby. One notable aspect of the ejecta was the existence of clusters of boulders, up to 3.6 m in radius, that were ejected at speeds of up to 52 m s −1 . Our analysis of the spatial distribution of 104 of these boulders suggests that they are likely the remnants of larger boulders shattered by the DART spacecraft in the first stages of the impact. The amount of momentum contained in these boulders is more than 3 times that of the DART spacecraft, and it is directed primarily to the south, almost perpendicular to the DART trajectory. Recoil of Dimorphos from the ejection of these boulders has the potential to change its orbital plane by up to a degree and to impart a non-principal-axis component to its rotation state. Damping timescales for these phenomena are such that the Hera spacecraft, arriving at the system in 2026, should be able to measure these effects.

asteroid dynamics↗

Full-polarization millimeter wavelength variability of Sagittarius A * during the 2018 EHT campaign

Context. Sagittarius A* (Sgr A*), the supermassive black hole at the center of the Milky Way, provides a unique laboratory to study accretion dynamics and plasma processes near the event horizon. Aims. We investigated the variability and polarization properties of Sgr A* using ALMA observations during the 2018 Event Horizon Telescope campaign. Methods. We analyzed high-cadence full-polarization light curves from ALMA at millimeter wavelengths, performed time-series analysis, and investigated the temporal behavior during an X-ray flare observed by Chandra on 2018 April 24. The variability characteristics are compared with expectations from standard accretion flow models. Results. We find low variability in total intensity (σ/μ < 10%), but significantly higher variability in linear and circular polarization (∼30% and ∼50%, respectively). A time-series analysis reveals red-noise variability, with power spectral densities between −2 and −3 across all Stokes parameters. Polarized intensity shows stable intra-day timescales, while total intensity exhibits more variable timescales, suggesting distinct emission regions, with polarization likely arising from a coherent structure. On April 24, a statistically significant inter-band delay in polarized intensity coincides with a near-simultaneous X-ray and millimeter peak that deviates from the typical delayed flare scenario. This event also features enhanced millimeter variability and coherent polarization loop evolution. The observed simultaneity challenges standard models of transient synchrotron emission with cooling delays, favoring instead a scenario of continuous energy injection in an optically thin region. Conclusions. Our results offer new constraints on the physical mechanisms driving variability in Sgr A*, and provide key observational input for refining theoretical models of accretion and plasma behavior in the vicinity of supermassive black holes.

Galaxy: center↗

Spatially resolved polarization swings in the supermassive binary black hole candidate OJ 287 with first Event Horizon Telescope observations

We present the first Event Horizon Telescope 1.3 mm observations of the supermassive binary black hole candidate OJ 287. The observations achieved an unprecedented angular resolution of 18 μas and reveal significant structural and polarization variability over just five days, marking the shortest timescale on which such changes have been directly imaged in this source. The inner jet exhibits a twisted ridgeline structure, with features displaying apparent superluminal motions up to about 22 c. The linear polarization maps reveal three main polarized features whose electric-vector position angles (EVPAs) change substantially over the time span of our observations, including a component with a radial polarization consistent with being produced by a recollimation shock. Most notably, we directly resolved two innermost jet components whose EVPAs rotate in opposite directions. The faster component, moving at 2.4 ± 0.9 μas/day (17.4 ± 6.5 c), exhibits counterclockwise EVPA swings of roughly 3.7° per day, while the slower component, with a proper motion of 1.4 ± 0.3 μas/day (10.2 ± 2.2 c), rotates clockwise at approximately 2.5° per day. Previous studies inferred helical magnetic fields in AGN jets from time-resolved or integrated polarization variability but lacked the angular resolution to directly image this effect. Our results provide spatially resolved evidence that a helical magnetic field threads the jet’s collimation and acceleration zone, ruling out models based on the superposition of unresolved components. Our analysis suggests that propagating shocks interact with a Kelvin–Helmholtz plasma instability, illuminating different phases of the helical magnetic field and producing the observed polarization spatial and temporal variability. Moreover, our model naturally accounts for the more rapid polarization rotation observed in the faster moving component. Our model predicts even more rapid swings in polarization, which could be tested with future observations featuring a more densely sampled time coverage.

OJ 287↗

In Silico Screening of CO 2 –Dipeptide Interactions for Bioinspired Carbon Capture

Carbon capture, sequestration and utilization offers a viable solution for reducing the total amount of atmospheric CO 2 concentrations. On an industrial scale, amine-based solvents are extensively employed for CO 2 capture through chemisorption. Nevertheless, this method is marked by the high cost associated with solvent regeneration, high vapor pressure, and the corrosive and toxic attributes of by-products, such as nitrosamines. An alternative approach is the biomimicry of sustainable materials that have strong affinity and selectivity for CO 2 . Bioinspired approaches, such as those based on naturally occurring amino acids, have been proposed for direct air capture methodologies. In this study, we present a database consisting of 960 dipeptide molecular structures, composed of the 20 naturally occurring amino acids. Furthermore, those structures were analyzed with a novel computational workflow presented in this work that considers certain interaction sites that determine CO 2 affinity. Density functional theory (DFT) and symmetry-adapted perturbation theory (SAPT) computations were performed for the calculation of CO 2 interaction energies, which allowed to limit our search space to 400 unique dipeptide structures. Using this computational workflow, we provide statistical insights into dipeptides and their affinity for CO 2 binding, as well as design principles that can further enhance CO 2 capture through cooperative binding.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

RCSB protein data Bank: Next‐generation advanced search for exploration of experimental structures and computed structure models

Abstract The Protein Data Bank (PDB), established in 1971, is the primary global, open‐access archive for experimentally determined 3D macromolecular structures (proteins, RNA, DNA). The research‐focused RCSB.org web‐portal provides access to these data alongside more than one million machine‐learning‐predicted structure models, greatly expanding the available structural landscape. Rapid growth of both experimental and computational structures has increased the need for powerful yet accessible search tools that serve a broad and diverse scientific community. Herein, we describe a redesigned RCSB Protein Data Bank RCSB.org Advanced Search capability that supports intuitive discovery of 3D structures through a unified interface. This interface integrates annotation‐, sequence‐, and 3D structure‐based searches, embeds an interactive 3D viewer, and incorporates curated biological knowledge, such as catalytic site definitions from Mechanism and Catalytic Site Atlas and ligand‐guided structural motifs, for constructing geometry‐driven queries. A new Chemical Search tool allows definition of chemical queries via an integrated drawing tool or standard identifiers, seamlessly combining them with annotation filters. By allowing query definition directly within spatial and chemical contexts, these search interfaces reduce the need for detailed knowledge of residue numbering, chain identifiers, or external cheminformatics software. This capability enables efficient exploration of structures, chemical diversity, and structure–function relationships across all life domains. The redesigned interfaces can be accessed directly at rcsb.org/search/advanced for Advanced Search and rcsb.org/search/chemical for Chemical Search.

Rose, Yana [Research Collaboratory for Structural ↗

A novel methodology for gamma-ray spectra dataset procurement over varying standoff distances and source activities

The adoption of machine learning approaches for gamma-ray spectroscopy has received considerable attention in the literature. Many studies have investigated the deployment of various algorithm architectures to a specific task. However, little attention has been afforded to the development of the datasets leveraged to train the models. Such training datasets typically span a set of environmental or detector parameters to encompass a problem space of interest to a user. Variations in these measurement parameters will also induce fluctuations in the detector response, including expected pile-up and ground scatter effects. Fundamental to this work is the understanding that 1) the underlying spectral shape varies as the measurement parameters change and 2) the statistical uncertainties associated with two spectra impact their level of similarity. While previous studies attribute some arbitrary discretization to the measurement parameters for the generation of their synthetic training data, this work introduces a principled methodology for efficient spectral-based discretization of a problem space. A signal-to-noise ratio (SNR) respective spectral comparison measure and a Gaussian Process Regression (GPR) model are used to predict the spectral similarity across a range of measurement parameters. This innovative approach effectively showcased its capability by dividing a problem space, ranging from 5 cm to 100 cm standoff distances and 5 μCi–100 μCi of 137 Cs, into three unique combinations of measurement parameters. The findings from this work will aid in creating more robust datasets, which incorporate many possible measurement scenarios, reduce the number of required experimental test set measurements, and possibly enable experimental training data collection for gamma-ray spectroscopy.

data science↗

CRISPRi-ART enables functional genomics of diverse bacteriophages using RNA-binding dCas13d

Bacteriophages constitute one of the largest reservoirs of genes of unknown function in the biosphere. Even in well-characterized phages, the functions of most genes remain unknown. Experimental approaches to study phage gene fitness and function at genome scale are lacking, partly because phages subvert many modern functional genomics tools. Here we leverage RNA-targeting dCas13d to selectively interfere with protein translation and to measure phage gene fitness at a transcriptome-wide scale. We find CRISPR Interference through Antisense RNA-Targeting (CRISPRi-ART) to be effective across phage phylogeny, from model ssRNA, ssDNA and dsDNA phages to nucleus-forming jumbo phages. Using CRISPRi-ART, we determine a conserved role of diverse rII homologues in subverting phage Lambda RexAB-mediated immunity to superinfection and identify genes critical for phage fitness. CRISPRi-ART establishes a broad-spectrum phage functional genomics platform, revealing more than 90 previously unknown genes important for phage fitness.

59 BASIC BIOLOGICAL SCIENCES↗

Analysis of differential scanning calorimetry data for aged plutonium

Differential scanning calorimetry data for samples of a 52 year old plutonium alloy with 3.3 at. % Ga that were heated beyond the melting point is analyzed using transition state theory to find activation energies for the δ to ε and ε to liquid phase transitions. A Bayesian statistical method involving a Gaussian process model is used to find mean values and confidence intervals for the activation energies. The activation energy for the δ to ε phase transition increases by 3.3 ± 3.8% per decade, relative to the case when all age related plutonium lattice point defects have been removed through annealing. The corresponding increase in activation energy for the ε to liquid transition is shown to be 7.1 ± 1.8% per decade. It is postulated that the change in activation energy with age for both phase transitions is caused, in part, by the accumulation of the same type of lattice point defects associated with the observed increase in elastic bulk modulus over time.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

pop-cosmos : redshifts and physical properties of KiDS-1000 galaxies

ABSTRACT Principled Bayesian inference of galaxy properties has not previously been performed for wide-area weak-lensing surveys with millions of sources. We address this gap by applying the pop-cosmos generative model to perform spectral energy distribution (SED) fitting for 4 million KiDS (Kilo-Degree Survey)-1000 galaxies. Calibrated on deep COSMOS2020 photometric data, pop-cosmos specifies a physically motivated prior over the galaxy population up to $z \simeq 6$ in stellar population synthesis (SPS) parameter space. Using the Speculator SPS emulator with GPU (graphics processing unit)-accelerated Markov Chain Monte Carlo sampling, we perform full posterior inference at 8.2 GPU seconds per galaxy, obtaining joint constraints on galaxy redshifts and physical properties. We validate photometric redshifts against $\sim \!185\,\!000$ KiDS galaxies cross-matched to Dark Energy Spectroscopic Instrument Data Release 1 spectroscopic samples, achieving low bias ($2\times 10^{-3}$), scatter ($\sigma _{\mathrm{MAD}}=0.03$), and outlier fraction (3.2 per cent) for the Bright Galaxy Survey, with comparable performance (bias $3\times 10^{-2}$, $\sigma _{\mathrm{MAD}}=0.05$, 1.0 per cent outliers) for luminous red galaxies (LRGs). Within the LRG sample, we identify massive, dusty, star-forming contaminants at $z \simeq 0.4$ satisfying standard colour selections for quenched populations. We infer trends in stellar mass, star formation, metallicity, and dust across five tomographic redshift bins consistent with established scaling relations. Using specific star formation rate constraints, we identify $\sim$7 per cent of KiDS-1000 galaxies as quenched, versus 37 per cent implied by conservative colour cuts. This enables the construction of weak-lensing samples defined by physical properties while mitigating intrinsic alignment systematics and preserving statistical power. Our analysis validates pop-cosmos out of sample, establishing it as a scalable approach for galaxy evolution and cosmological analyses with photometric surveys.

Halder, Anik [Institute of Astronomy and Kavli Ins↗

Reimagining Disassembly Interfaces With Visualization: Combining Instruction Tracing and Control Flow With DisViz

In applications where efficiency is critical, developers may examine their compiled binaries, seeking to understand how the compiler transformed their source code and what performance implications that transformation may have. This analysis is challenging due to the vast number of disassembled binary instructions and the many-to-many mappings between them and the source code. These problems are exacerbated as source code size increases, giving the compiler more freedom to map and disperse binary instructions across the disassembly space. Interfaces for disassembly typically display instructions as an unstructured listing or sacrifice the order of execution. Here, we design a new visual interface for disassembly code that combines execution order with control flow structure, enabling analysts to both trace through code and identify familiar aspects of the computation. Central to our approach is a novel layout of instructions grouped into basic blocks that displays a looping structure in an intuitive way. We add to this disassembly representation a unique block-based mini-map that leverages our layout and shows context across thousands of disassembly instructions. Finally, we embed our disassembly visualization in a web-based tool, DisViz, which adds dynamic linking with source code across the entire application. DizViz was developed in collaboration with program analysis experts following design study methodology and was validated through evaluation sessions with ten participants from four institutions. Participants successfully completed the evaluation tasks, hypothesized about compiler optimizations, and noted the utility of our new disassembly view. Our evaluation suggests that our new integrated view helps application developers in understanding and navigating disassembly code.

Computer science↗

Establishing reference ranges for circulating biomarkers of drug‐induced liver injury in healthy human volunteers 1

Aims The potential of mechanistic biomarkers to improve prediction of drug‐induced liver injury (DILI) and hepatic regeneration is widely acknowledged. We sought to determine reference intervals for new biomarkers of DILI and regeneration, as well as to characterize their natural variability and impact of diurnal variation. Methods Serum samples from 227 healthy volunteers were recruited as part of a cross‐sectional study; of these, 25 subjects had weekly serial sampling over 3 weeks, while 23 had intensive blood sampling over a 24h period. Alanine aminotransferase (ALT), MicroRNA‐122 (miR‐122), High Mobility Group Box‐1 (HMGB1), total Keratin‐18 (K18), caspase‐cleaved Keratin‐18 (ccK18), Glutamate Dehydrogenase (GLDH) and Macrophage Colony‐Stimulating Factor‐1 (CSF‐1) were assayed. Results Reference intervals were established for each biomarker based on the 97.5% quantile (90% CI) following the assessment of fixed effects in univariate and multivariable models. Intra‐individual variability was found to be non‐significant, and there was no significant impact of diurnal variation. Conclusion Reference intervals for novel DILI biomarkers have been described. An upper limit of a reference range might represent the most appropriate mechanism to utilize these data. These data can now be used to interpret data from exploratory clinical DILI studies and to assist their further qualification as required by regulatory authorities.

Jorgensen, Andrea L. [Department of Health Data Sc↗

DMRG++

A free and open source implementation of the DMRG Algorithm

Alvarez, Gonzalo [Oak Ridge National Laboratory (O↗

Data for FUN-PROSE: A Deep Learning Approach to Predict Condition-Specific Gene Expression in Fungi

mRNA levels of all genes in a genome is a critical piece of information defining the overall state of the cell in a given environmental condition. Being able to reconstruct such condition-specific expression in fungal genomes is particularly important to metabolically engineer these organisms to produce desired chemicals in industrially scalable conditions. Most previous deep learning approaches focused on predicting the average expression levels of a gene based on its promoter sequence, ignoring its variation across different conditions. Here we present FUN-PROSE—a deep learning model trained to predict differential expression of individual genes across various conditions using their promoter sequences and expression levels of all transcription factors. We train and test our model on three fungal species and get the correlation between predicted and observed condition-specific gene expression as high as 0.85. We then interpret our model to extract promoter sequence motifs responsible for variable expression of individual genes. We also carried out input feature importance analysis to connect individual transcription factors to their gene targets. A sizeable fraction of both sequence motifs and TF-gene interactions learned by our model agree with previously known biological information, while the rest corresponds to either novel biological facts or indirect correlations.

Genomics↗

Carbon Storage Technical Viability Approach (CS TVA) Matrix

The Carbon Storage Technical Viability Approach (CS TVA) Matrix is a knowledge framework developed to outline the information needed for geologic carbon storage. The CS TVA Matrix contains 5 categories, 14 sub-categories, and 47 components. This framework can be leveraged to assess the availability of data and information needed for a carbon storage project. The information categories of the matrix are tied to a list of required data using weighted mapping, published herein.

carbon storage↗