Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Coassembly and binning of a twenty-year metagenomic time-series from Lake Mendota

Abstract The North Temperate Lakes Long-Term Ecological Research (NTL-LTER) program has been extensively used to improve understanding of how aquatic ecosystems respond to environmental stressors, climate fluctuations, and human activities. Here, we report on the metagenomes of samples collected between 2000 and 2019 from Lake Mendota, a freshwater eutrophic lake within the NTL-LTER site. We utilized the distributed metagenome assembler MetaHipMer to coassemble over 10 terabases (Tbp) of data from 471 individual Illumina-sequenced metagenomes. A total of 95,523,664 contigs were assembled and binned to generate 1,894 non-redundant metagenome-assembled genomes (MAGs) with ≥50% completeness and ≤10% contamination. Phylogenomic analysis revealed that the MAGs were nearly exclusively bacterial, dominated by Pseudomonadota (Proteobacteria, N = 623) and Bacteroidota (N = 321). Nine eukaryotic MAGs were identified by eukCC with six assigned to the phylum Chlorophyta. Additionally, 6,350 high-quality viral sequences were identified by geNomad with the majority classified in the phylum Uroviricota. This expansive coassembled metagenomic dataset provides an unprecedented foundation to advance understanding of microbial communities in freshwater ecosystems and explore temporal ecosystem dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Multi-Attribute Subset Selection enables prediction of representative phenotypes across microbial populations

The interpretation of complex biological datasets requires the identification of representative variables that describe the data without critical information loss. This is particularly important in the analysis of large phenotypic datasets (phenomics). Here we introduce Multi-Attribute Subset Selection (MASS), an algorithm which separates a matrix of phenotypes (e.g., yield across microbial species and environmental conditions) into predictor and response sets of conditions. Using mixed integer linear programming, MASS expresses the response conditions as a linear combination of the predictor conditions, while simultaneously searching for the optimally descriptive set of predictors. We apply the algorithm to three microbial datasets and identify environmental conditions that predict phenotypes under other conditions, providing biologically interpretable axes for strain discrimination. MASS could be used to reduce the number of experiments needed to identify species or to map their metabolic capabilities. The generality of the algorithm allows addressing subset selection problems in areas beyond biology.

59 BASIC BIOLOGICAL SCIENCES↗

Thoroughly testing and integrating hundreds of Pull Requests per month: ROOT’s new Cost-efficient and Feature Rich GitHub-based CI

ROOT is an open source framework, freely available on GitHub, at the heart of data acquisition, processing and analysis of HE(N)P experiments, and beyond. It is developed collaboratively: contributions are not authored only by ROOT team members, but also by the user community at large: developers and scientists from universities, labs as well as the private sector. More than 1500 GitHub Pull Requests are merged on average per year. It is in this context that code integration acquires a primary role. The review of code contributions isn’t enough: not only they need to be thoroughly reviewed, they also need to be thoroughly tested through a powerful CI infrastructure on several different platforms to comply with the high code quality standards of the project. Since the end of 2023, ROOT moved its continuous integration system from Jenkins to GitHub Actions. In this contribution, we characterise the transition to the GitHub CI, focussing on our strategy, its implementation and the lessons learned, as well as the advantages the new system offers with respect to the previous one. Particular emphasis will be given to the evaluation of the cost-benefit ratio for Jenkins and GitHub Actions for the ROOT project. We also describe how we manage to run in less than one hour thousands of unit, integration, functional and end-to-end tests on different flavours of Windows, four versions of macOS, as well as about ten of the most used Linux distributions, taking advantage of the CERN computing infrastructure.

Piparo, Danilo [CERN]↗

cclib 2.0: An updated architecture for interoperable computational chemistry

Interoperability in computational chemistry is elusive, impeded by the independent development of software packages and idiosyncratic nature of their output files. The cclib library was introduced in 2006 as an attempt to improve this situation by providing a consistent interface to the results of various quantum chemistry programs. The shared API across programs enabled by cclib has allowed users to focus on results as opposed to output and to combine data from multiple programs or develop generic downstream tools. Initial development, however, did not anticipate the rapid progress of computational capabilities, novel methods, and new programs; nor did it foresee the growing need for customizability. Here, we recount this history and present cclib 2, focused on extensibility and modularity. We also introduce recent design pivots—the formalization of cclib’s intermediate data representation as a tree-based structure, a new combinator-based parser organization, and parsed chemical properties as extensible objects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Novel data interpretation method for DIII-D divertor retarding field energy analyzer with 3-D particle-in-cell simulations

A novel data interpretation process that utilizes comprehensive particle-in-cell (PIC) simulations is developed for the new retarding field energy analyzer (RFEA) currently being constructed at DIII-D for the lower divertor using the Divertor Material Evaluation System. Furthermore, this probe is expected to survive a heat load of up to 100 MW/m 2 for up to 5 s and reliably measure the main ion temperature (T i ) on the divertor target ranging from 10 to 200 eV. These extreme conditions posed significant engineering limitations on the probe geometry, thus extensive validation work has been performed. The conventional fitting method for the RFEA I–V characteristics is based on a simplified 1-D model without considering the ion space charge inside the probe cavity and may not be sufficient for probes designed for the DIII-D divertor environment. In this article, a more realistic description of the particle propagation process within the RFEA cavity is achieved by including both 3-D geometric effects and ion space charge in the PIC simulations, and the capability to reconstruct the ion energy distribution functions is demonstrated with reasonable consistency.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Phase-based velocity extraction method for photonic Doppler velocimetry with potential higher time resolution

We present an extension of the [Takeda et al., J. Opt. Soc. Am. 72, 156 (1982)] phase extraction method to heterodyne photonic Doppler velocimetry applications. The method yields results equivalent to those obtained by the short-time Fourier transform (STFT), while offering potential improvements in time resolution. Unlike STFT, which relies on window functions, such as the Hamming window, that emphasize central data points and diminish the influence of edges, the extended Takeda method utilizes all data uniformly. This uniform treatment allows for the derivation of empirical equations that directly relate velocity error to the actual time resolution rather than to the local analysis duration. The established equation provides a useful metric for both optimizing hardware configuration and guiding data analysis. Simulation and experimental results confirm that, for a given dataset, specifying a target time resolution yields consistent velocity errors for both methods. These findings underscore the Takeda method’s advantages, particularly its potential higher time resolution and reduced computational burden, making it a valuable tool for high-throughput applications such as laser dynamic compression experiments.

Computer simulation↗

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks↗

Optimizing spin dressing sensitivity for the nEDMSF experiment

nEDMSF aims to measure the neutron electric dipole moment (d n ) with unprecedented precision. In this paper we explore the experiment's sensitivity when operating with an implementation of the critical dressing method in which the angle between the neutron and Helium-3 spins (ϕ 3n ) is subjected to a square modulation by an amount ϕ d (the “dressing angle”). Several parameters can be tuned to optimize sensitivity. We find roughly 10% improvement over a previous estimate, resulting primarily from the addition of a waiting period between the π/2 pulse that initiates d n -driven ϕ 3n growth and the start of ϕ3n modulation. We find negligible further improvement by allowing ϕ d to vary continuously over the course of a run, and no degradation resulting from the addition of an in situ background measurement into each ϕ3n modulation sequence. A complete simulation confirms a 300 live-day sensitivity ofσ = 1.45×10 -28 e ·cm. At this level of sensitivity, σ ϕ3n0 = 1 mrad precision on the initial n/ 3 He angle difference is not negligible.

47 OTHER INSTRUMENTATION↗

Tools for unbinned unfolding

Machine learning has enabled differential cross section measurements that are not discretized. Going beyond the traditional histogram-based paradigm, these unbinned unfolding methods are rapidly being integrated into experimental workflows. Here, in order to enable widespread adaptation and standardization, we develop methods, benchmarks, and software for unbinned unfolding. For methodology, we demonstrate the utility of boosted decision trees for unfolding with a relatively small number of high-level features. This complements state-of-the-art deep learning models capable of unfolding the full phase space. To benchmark unbinned unfolding methods, we develop an extension of existing dataset to include acceptance effects, a necessary challenge for real measurements. Additionally, we directly compare binned and unbinned methods using discretized inputs for the latter in order to control for the binning itself. Lastly, we have assembled two software packages for the OmniFold unbinned unfolding method that should serve as the starting point for any future analyses using this technique. One package is based on the widely-used RooUnfold framework and the other is a standalone package available through the Python Package Index (PyPI).

47 OTHER INSTRUMENTATION↗

SALSA: a new versatile readout chip for MPGD detectors

The SALSA chip is a future readout ASIC foreseen for the MPGD detectors, developed in the framework of the EIC collider project, to equip the MPGD trackers of the EPIC experiment. It is designed to be versatile, to be adapted to other usages of MPGD detectors like TPC or photon detectors. It integrates a frontend block and an ADC for each of the 64 channels, associated to a configurable DSP processor meant to correct data and reduce the raw data flux to limit the output bandwidth. It will be compatible with the continuous readout foreseen for the EPIC DAQ, but will also work in a triggered environment. Several prototypes are already produced in order to qualify the different blocks of the chip, in particular the frontend, the ADC and the clock generation. The next 32-channel prototype is currently under development and is planed to be produced in 2025. In conclusion, the final prototype will be produced and tested from 2026 for a production of the SALSA chip at the horizon of 2027.

Data processing↗

Neural posterior unfolding

Differential cross section measurements are the currency of scientific exchange in particle and nuclear physics. A key challenge for these analyses is the correction for detector distortions, known as deconvolution or unfolding. Binned unfolding of cross section measurements traditionally rely on the regularized inversion of the response matrix that represents the detector response, mapping pre-detector (`particle level') observables to post-detector (`detector level') observables. In this paper we introduce Neural Posterior Unfolding, a modern, Bayesian approach that leverages normalizing flows for unfolding. By using normalizing flows for neural posterior estimation, NPU offers several key advantages including implicit regularization through the neural network architecture, fast amortized inference that eliminates the need for repeated retraining, and direct access to the full uncertainty in the unfolded result. In addition to introducing NPU, we implement a classical Bayesian unfolding method called Fully Bayesian Unfolding (FBU) in modern Python so it can also be studied. These tools are validated on simple Gaussian examples and then tested on simulated jet substructure examples from the Large Hadron Collider (LHC). We find that the Bayesian methods are effective and worth additional development to be analysis ready for cross section measurements at the LHC and beyond.

Analysis and statistical methods↗

The Natural Products Magnetic Resonance Database (NP-MRD) for 2025

The Natural Products Magnetic Resonance Database or NP-MRD (https://np-mrd.org) is a comprehensive, freely accessible, web-based resource for the deposition, distribution, extraction and retrieval of nuclear magnetic resonance (NMR) data on natural products. The NP-MRD was initially established to support compound de-replication and data dissemination for the natural products community. However, that community has now grown to include many users from the metabolomics, microbiomics, foodomics and nutrition science fields. Indeed, since its launch in 2021, the NP-MRD has expanded enormously in size, scope and popularity. The current version of NP-MRD now contains nearly 7X more compounds (281,859 vs. 40,908) and 7X more NMR spectra (5.1 million vs. 817,000) than the first release. More specifically, an additional 4.6 million predicted spectra and another 11,000 spectra simulated from experimental chemical shifts were deposited into the database. Likewise, the number of NMR raw spectral data depositions has grown from a 165 spectra per year to more than 10,000 per year. As a result of this expansion, the number of monthly webpage views has grown from 55 to 20,000 and the number of monthly visitors has increased from 7 to 2500. To address this growth and to better support the expanding needs of its diverse community of users, many additional improvements to the NP-MRD have been made. These include significant enhancements to the data submission process, important improvements to the visualization and display of NMR spectra, notable updates to the database’s spectral search utilities and useful additions to support better NMR spectral analysis/prediction. Significant efforts have also been undertaken to remediate and update many of NP-MRD’s database entries. This manuscript describes these database improvements and expansion efforts, along with how they have been implemented and what future upgrades to the NP-MRD are planned.

Artifical Intelligence↗

Measurement of the A dependence of the ν μ charged-current quasielasticlike cross section as a function of muon and proton kinematics at ⟨E ν ⟩ ∼ 6 GeV

The first simultaneous measurements of the 𝜈 𝜇 quasielasticlike cross section on C, CH, H 2 ⁡O, Fe, and Pb targets as a function of kinematic imbalance variables in the plane transverse to the incoming neutrino direction are presented. These variables combine the muon and proton information to provide a new way to disentangle the effects of the nucleus in quasielasticlike processes. The data were obtained using a wideband 𝜈 𝜇 beam with ⟨E 𝜈 ⟩ ∼ 6 GeV. Cross-section ratios of the different target materials to CH are also shown. These measurements are used to explore the nature of the cross-section 𝐴 scaling, as well as initial and final state interaction effects. Comparisons are made to predictions from a number of commonly used neutrino Monte Carlo event generators. The range of predictions of the different models tends to cover the data but the degree and consistency of the agreement suffers in regions, and on higher 𝐴 targets, where the final state interactions are expected to be more pronounced.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Randomized low-rank decompositions of nuclear three-body interactions

First-principles simulations of many-fermion systems are commonly limited by the computational requirements of processing large data objects. As a remedy, we propose the use of low-rank approximations of three-body interactions, which are the dominant such limitation in nuclear physics. We introduce a randomized decomposition technique to handle the excessively large matrix dimensions and study the sensitivity of low-rank properties to interaction details. The developed low-rank three-nucleon interactions are benchmarked in ab initio simulations of few- and many-body systems. Exploiting low-rank properties provides a promising route to extend the microscopic description of atomic nuclei to large systems where storage requirements exceed the computational capacities of the most advanced high-performance computing facilities.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Traffic Signal Control for Large-Scale Urban Traffic Networks: Real-World Experiments using Vision-Based Sensors

Effective control of traffic signals plays a critical role in ensuring smooth vehicle flow in urban areas. Expertly engineered traffic signal controllers can considerably minimize travel delays and enhance sustainability. In this paper, the team proposes the Model Predictive Control (MPC) traffic signal control strategy using real-time traffic flow data from a vision-based camera as feedback information. Also, a realistic signal timing plan that considers National Electrical Manufacturers Association (NEMA) constraints has been developed to be applied to real-world scenarios. The primary aim is to reduce the number of vehicles across all links in the controlled area, thereby optimizing traffic flow and reducing energy consumption. To validate the proposed method, several real-life experiments were conducted at 24 intersections in Chattanooga, Tennessee, by collaborating with traffic field engineers. These experiments demonstrated significant performance improvements in comparison to the existing method.

data processing↗

PMU Data Quality and Sensor Health Monitoring

Phasor Measurement Units (PMUs) play a critical role in the evolution of the electric power industry by providing high-precision, real-time monitoring of essential power system metrics. However, effectively detecting abnormalities and critical events from PMU data is a complex task, complicated by intricate temporal patterns, a scarcity of labeled data for training algo- rithms, and constraints on online computational power. In this study, we apply TranAD, an innovative algorithm that combines transformer architectures with the refinement of adversarial learning, to both synthetic and real-world PMU datasets for developing a data quality and sensor online health monitoring platform for utilities. Our findings reveal that TranAD not only provides efficient detection and localization but also enhances the detail with which abnormalities are detected, marking a a significant step forward in the field of clean data acquisition processes for power system monitoring

deep neural network, machine learning (ML)↗

Enhanced Machine-Learning Flow for Microwave-Sensing Systems for Contaminant Detection in Food

The presence of foreign bodies in packaged food is a serious concern for both fnal consumers (allergies, injuries, choking) and food manufacturers (reputation and economic losses). In particular, low-density plastics, glass and wood splinters are hard to detect even by the most advanced X-ray imagers. One solution is Machine-Learning-based Microwave Sensing (MLMWS): a non-invasive, contactless, and real-time method which uses a machine-learning (ML) classifer to analyze the scattered microwaves from the irradiated target object. In this paper, we want to extend our previous work about contaminant detection in cocoa-hazelnut spread jars by proposing an enhanced ML flow to increase the accuracy of the ML classifier. For the first time in this case study, we use a multi-class classifier, we train it with scattering parameters measured at multiple microwave frequencies, with a new pre-processing scaler, data augmentation, quantization-aware training and a pruning schedule. The results show a contaminant detection multi-class accuracy of 94.167% with a latency of 26 µs when targeting an AMD/Xilinx Kria K26 FPGA. Finally, we released our datasets publicly to OpenML.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗