Search NASA⌕ Search

SEARCH · Search NASA

Results for “data sets”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

AERO-MAP: a data compilation and modeling approach to understand spatial variability in fine- and coarse-mode aerosol composition

Abstract. Aerosol particles are an important part of the Earth climate system, and their concentrations are spatially and temporally heterogeneous, as well as being variable in size and composition. Particles can interact with incoming solar radiation and outgoing longwave radiation, change cloud properties, affect photochemistry, impact surface air quality, change the albedo of snow and ice, and modulate carbon dioxide uptake by the land and ocean. High particulate matter concentrations at the surface represent an important public health hazard. There are substantial data sets describing aerosol particles in the literature or in public health databases, but they have not been compiled for easy use by the climate and air quality modeling community. Here, we present a new compilation of PM2.5 and PM10 surface observations, including measurements of aerosol composition, focusing on the spatial variability across different observational stations. Climate modelers are constantly looking for multiple independent lines of evidence to verify their models, and in situ surface concentration measurements, taken at the level of human settlement, present a valuable source of information about aerosols and their human impacts complementarily to the column averages or integrals often retrieved from satellites. We demonstrate a method for comparing the data sets to outputs from global climate models that are the basis for projections of future climate and large-scale aerosol transport patterns that influence local air quality. Annual trends and seasonal cycles are discussed briefly and are included in the compilation. Overall, most of the planet or even the land fraction does not have sufficient observations of surface concentrations – and, especially, particle composition – to characterize and understand the current distribution of particles. Climate models without ammonium nitrate aerosols omit ∼ 10 % of the globally averaged surface concentration of aerosol particles in both PM2.5 and PM10 size fractions, with up to 50 % of the surface concentrations not being included in some regions. In these regions, climate model aerosol forcing projections are likely to be incorrect as they do not include important trends in short-lived climate forcers.

Mahowald, Natalie M. (ORCID:000000022873997X)↗

Analyzing the impact of design factors on solar module thermomechanical durability using interpretable machine learning techniques

Solar modules in utility-scale systems are expected to maintain decades of lifetime to rival conventional energy sources. However, cyclic thermomechanical loading often degrades their long-term performance, highlighting the importance of effective design to mitigate thermal expansion mismatches between module materials. Given the complex composition of solar modules, isolating the impact of individual components on overall durability remains a challenging task. In this work, we analyze a comprehensive data set that comprises bill-of-materials (BOM) and thermal cycling power loss from 251 distinct module designs to identify the predominant design factors and their impacts on the thermomechanical durability of modules. The methodology of our analysis combines machine learning modeling (random forest) and Shapley additive explanation (SHAP) to correlate design factors with power loss and interpret the model’s decision-making. The interpretation reveals that silicon type (monocrystalline or polycrystalline), encapsulant thickness, busbar numbers, and wafer thickness predominantly influence the degradation. With lower power loss of around 0.6% on average in the SHAP analysis, monocrystalline cells present better durability than polycrystalline cells. This finding is further substantiated by statistical testing on our raw data set. The SHAP analysis also demonstrates that while thicker encapsulants lead to reduced power loss, further increasing their thickness over around 0.6 to 0.7 mm does not yield additional benefits, particularly for the front side one. In addition, other important BOM features such as the number of busbars are analyzed. This study provides a blueprint for utilizing explainable machine learning techniques in a complex material system and can potentially guide future research on optimizing the design of solar modules.

14 SOLAR ENERGY↗

Quantifying motion blur by imaging shock front propagation with broadband and narrowband X-ray sources

Time-integrated radiography using MeV Bremsstrahlung X-ray sources is the norm for imaging during system-level testing of components and structures under dynamic condition. One source of error in the analysis of the time-integrated radiography data sets stems from motion blur which smears out sharp interfaces to a greater degree with longer exposure times, which become necessary to provide sufficient signal-to-noise with low X-ray penetration of objects of interest. To quantify motion blur, a 1D shock wave through PMMA was investigated experimentally at The Dynamic Compression Sector at The Advanced Photon Source (DCS@APS) with tapered broadband and 25.46 ± 1.06 keV narrowband X-rays. Four cameras with different exposure times were used for each experiment to compare the effect that exposure time has on motion blur. In addition, our methodology to accurately simulate motion blur in terms of transmission and shape is presented and compared to our experimental results and quantified. There is a high level of agreement between the experimental and simulation results across the range of data sets investigated in this study with a percent difference range of 0.29–1.31% for the four shots. The methodology of this work serves as a steppingstone towards a physically validated model that could be used in conjunction with experimental results to deconvolve physical parameters, densities, and interfaces of interest in a way that would not be possible with experimental results alone.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A Parameter-masked Mock Data Challenge for Beyond-two-point Galaxy Clustering Statistics

The past few years have seen the emergence of a wide array of novel techniques for analyzing high-precision data from upcoming galaxy surveys, which aim to extend the statistical analysis of galaxy clustering data beyond the linear regime and the canonical two-point (2pt) statistics. We test and benchmark some of these new techniques in a community data challenge named “Beyond-2pt,” initiated during the Aspen 2022 Summer Program “Large-Scale Structure Cosmology beyond 2-Point Statistics,” whose first round of results we present here. The challenge data set consists of high-precision mock galaxy catalogs for clustering in real space, in redshift space, and on a light cone. Participants in the challenge have developed end-to-end pipelines to analyze mock catalogs and extract unknown (“masked”) cosmological parameters of the underlying ΛCDM models with their methods. The methods represented are density-split clustering, nearest neighbor statistics, BACCO power spectrum emulator, void statistics, LEFTfield field-level inference using effective field theory (EFT), and joint power spectrum and bispectrum analyses using both EFT and simulation-based inference. In this work, we review the results of the challenge, focusing on problems solved, lessons learned, and future research needed to perfect the emerging beyond-2pt approaches. The unbiased parameter recovery demonstrated in this challenge by multiple statistics and the associated modeling and inference frameworks supports the credibility of cosmology constraints from these methods. The challenge data set is publicly available, and we welcome future submissions from methods that are not yet represented.

Krause, Elisabeth [Univ. of Arizona, Tucson, AZ (U↗

Source apportionment of aerosols at the White River IMPROVE site near the SAIL site

This data set contains source apportionment results at the White River IMPROVE site (39.1536, -106.8209), which is about 30 km north of the Surface Atmosphere Integrated Field Laboratory (SAIL) Campaign site. The IMPROVE network (Malm et al. 1994) collected 24-hour aerosol filter samples every three days over several decades at this site. Chemical concentrations in the PM2.5 fraction of 19 elements (Al, As, Br, Ca, Cl, Cr, Cu, Fe, K, Mg, Mn, Na, Ni, Pb, Se, Si, Ti, V, and Zn), along with nitrate, sulfate, elemental carbon (EC), organic carbon (OC), and calculated coarse mass concentrations (PM10−PM2.5 mass concentrations), from 2014 to 2023, were used as input for the PMF analysis. PMF was performed using EPA PMF 5.0 (Norris et al. 2014). A five-factor solution was chosen as the optimal solution. These factors were identified as coarse dust, fine dust, biomass burning, sulfate-dominated, and nitrate-dominated sources. This data set is useful for understanding aerosol sources and their long-term variability near this region.

biomass burning↗

3 He +𝛼 resonances in 7 Be

Resonances in 7 Be which decay into the 3 He+α exit channel have been measured with improved precision using preexisting data sets. The energy and width of the J π =7/2 - state have been extracted from an invariant-mass study of projectile-breakup products originating from interactions of an E/A=10.7-MeV 10 C beam on Be and C targets. The excitation energy of this state (from the pole of the S-matrix) is determined to be E*=4.545(6)~MeV, a factor of 8 improvement in precision as compared to the ENSDF value. This improvement is enabled by fine tuning the detector calibrations using calibration resonances in 6 Li, 7 Li, 6 Be, 9 B, and 12 C whose decay energies are known to high precision. The J π =7/2- resonance in 7 Be can now itself be used as a calibration resonance in invariant-mass experiments. This utility is particularly helpful for 3 He energy calibrations of CsI(Tl) detectors which are often used in detector arrays employed for measurements with fast beams. This utility is demonstrated with a data set associated with E/A=70 MeV 7 Be beams which are inelastically excited to the 7/2 - and 5/2 - 1 resonances. Again using the pole of the S-matrix as the definition of the resonance parameters, the fitted excitation energy of the 5/2 - 1 resonance is 6.376(17) MeV, approximately 300 keV lower than the ENSDF value. Finally, its width of 565(4) keV is roughly half of the ENSDF value.

energy levels↗

Bias Correction and Statistical Downscaling of Future Solar Irradiance Projections Using the NSRDB

Assessing renewable energy resources under future climate scenarios has been highlighted to understand potential impacts of future climate change in renewable generation on the power sector. Climate model projection has been recognized by the renewable energy community as a useful data set to analyze the impacts of future climate change on renewable resources. However, future climate projections generated from general circulation models (GCMs) contain inherent biases that need to be corrected for accurate analysis of future projections of climate variables. In addition, the coarse spatiotemporal resolution of GCMs needs to be improved for regional climate studies. In this work, we develop statistical methods to downscale future projections of global horizontal irradiance (GHI) in a computationally efficient way. Our approach builds statistical downscaling models that correct bias of climate projection of GHI and downscale the future GHI projection from daily-scale to hourly-scale. The National Solar Radiation Database (NSRDB) is used to calibrate the statistical models and validate the downscaled GHI projections across the contiguous United State (CONUS). Preliminary results show that the statistical approach efficiently downscales climate projections of GHI with a nBIAS of 3%, nMAE of 34 % and nRMSE of 46% calculated against NSRDB for CONUS. This study describes the implemented methodology and initial results as well as future research to create high-resolution climate data sets for solar energy applications.

analytical models↗

Thermal images collected during large-scale 3D printing with Ingersoll MasterPrint

This data set contains images produced to test the performance of anomaly and fault detection methods in the context of additive manufacturing. Two print jobs were executed using the Ingersoll MasterPrint with the Model 30 Strangpresse extruder. The material used was Techmer compounded polylactic acid (PLA) with wood flour as a filler (80/20 PLA/WF by weight). The first print job consists of the 3D printing of a hexagonal cylinder with a two-bead wall. This print was sliced at a gantry velocity of 3000 millimeter per minute, and an extruder screw speed of 68.14 rotations per minute. This are considered the normal operating conditions. A second hexagonal cylinder was printed using a reduced extruder screw speed, 15% lower than under normal operating conditions. The images were collected with a Teledyne FLIR Lepton 3.5 infra-red camera, a small form factor radiometric long-wave infrared camera with a spectral range of 8 µm to 14 µm. Sensor resolution was 160x120 pixels, with a pixel size of 12 µm, a temperature range of -10 - 450°C, and an accuracy of +/- 10°C in its low gain configuration. Images in this data set were collected with cameras oriented at the printer nozzle. The nozzle camera setup consisted of two cameras located 12.5cm from the nozzle center. These were mounted directly to the print head, so that the camera positions relative to the print direction would remain constant as it rotated around its C-axis to follow the print path. One camera was placed ahead of the nozzle to capture the previous layer immediately before being covered by the new layer of material, while the second camera was placed behind the nozzle and captured the freshly extruded bead. Images are collecting during each phase of the print: (a) idle (i.e., no material deposited), (b) extrusion (i.e., to prime the extruder), and (c) printing (deposition of material to manufacture the hexagonal cylinder).

36 MATERIALS SCIENCE↗

JOINT APPOINTEE: Evolution of ferroelectric properties in SmxBi1-xFeO3 via automated Piezoresponse Force Microscopy across combinatorial spread libraries

Combinatorial spread libraries offer a innovative approach to explore the evolution of material properties over broad concentration, temperature, and growth parameter spaces. However, traditional limitation of this approach is the requirement for the read-out of functional properties across the library. Here we develop automated Piezoresponse Force Microscopy (PFM) for the exploration of combinatorial spread libraries and demonstrate its application in the SmxBi1-xFeO3 system with the ferroelectric-antiferroelectric morphotropic phase boundary. This approach relies on the synergy of the quantitative nature of PFM and the implementation of automated experiments that allow PFM-based sampling over macroscopic samples. The concentration dependence of pertinent ferroelectric parameters has been determined and used to develop the mathematical framework based on Ginzburg-Landau theory describing the evolution of these properties across the concentration space. We pose that a combination of automated scanning probe microscope and combinatorial spread library approach will emerge as an efficient research paradigm to close the characterization gap in the high-throughput materials discovery. We make the data sets open to the community and hope that this will stimulate other efforts to interpret and understand the physics of these systems.

Automated Microscopy, Combinatorial Library, Ferro↗

Improving Bond Dissociations of Reactive Machine Learning Potentials through Physics-Constrained Data Augmentation

In the field of computational chemistry, predicting bond dissociation energies (BDEs) presents well-known challenges, particularly due to the multireference character of reactive systems. Many chemical reactions involve configurations where single-reference methods fall short, as the electronic structure can significantly change during bond breaking. As generating training data for partially broken bonds is a challenging task, even state-of-the-art reactive machine learning interatomic potentials (MLIPs) often fail to predict reliable BDEs and smooth dissociation curves. By contrast, simple and inexpensive physics-based models, such as the well-established Morse potential, do not suffer from any such limitations. This work leverages the Morse potential to improve reactive MLIPs by augmenting the training data set with inexpensive Morse data along the dissociation pathways. Further, this physics-constrained data augmentation (PCDA) approach results in MLIPs with smooth bond dissociation curves as well as near coupled-cluster level BDEs, all without requiring any expensive multireference quantum mechanical calculations. A case study for methane combustion demonstrates how the PCDA approach can improve an existing reactive MLIP, namely, ANI-1xnr. In conclusion, not only are the BDEs and bond dissociation curves for all radicals and molecules significantly improved compared to ANI-1xnr but the PCDA-trained MLIP retains the reliability of ANI-1xnr when performing reactive molecular dynamics simulations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modeling the Enceladus dust plume based on in situ measurements performed with the Cassini Cosmic Dust Analyzer

We analyzed data recorded by the Cosmic Dust Analyzer on board the Cassini spacecraft during Enceladus dust plume traversals. Our focus was on profiles of relative abundances of grains of different compositional types derived from mass spectra recorded with the Dust Analyzer subsystem during the Cassini flybys E5 and E17. The E5 profile, corresponding to a steep and fast traversal of the plume, has already been analyzed. In this paper, we included a second profile from the E17 flyby involving a nearly horizontal traversal of the south polar terrain at a significantly lower velocity. Additionally, we incorporated dust detection rates from the High Rate Detector subsystem during flybys E7 and E21. We derived grain size ranges in the different observational data sets and used these data to constrain parameters for a new dust plume model. This model was constructed using a mathematical description of dust ejection implemented in the software package DUDI. Further constraints included published velocities of gas ejection, positions of gas and dust jets, and the mass production rate of the plume. Our model employs two different types of sources: diffuse sources of dust ejected with a lower velocity and jets with a faster and more colimated emission. From our model, we derived dust mass production rates for different compositional grain types, amounting to at least 28 kg s –1 . Previously, salt-rich dust was believed to dominate the plume mass based on E5 data alone. The E17 profile shows a dominance of organic-enriched grains over the south polar terrain, a region not well constrained by E5 data. By including both E5 and E17 profiles, we find the salt-rich dust contribution to be at most 1% by mass. This revision also results from an improved understanding of grain masses of various compositional types that implies smaller sizes for salt-rich grains. Our new model can predict grain numbers and masses for future mission detectors during plume traversals.

79 ASTRONOMY AND ASTROPHYSICS↗

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul [National Renewable Energy Lab. (NREL↗

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING↗

Inter-Kingdom Viral Interactions

Please cite as : Josué A. Rodríguez-Ramos, Amy E. Zimmerman, Ruonan Wu, Sheryl Bell, Trinidad Alfaro, Kirsten Hofmockel, William C. Nelson. 2025. Inter-Kingdom Viral Interactions. [Data Set] PNNL DataHub. This data is published under a CC0 license. The authors encourage data reuse and request attribution by referencing the above citations for the data package and associated manuscript. Deciphering viral ecology in soils is challenging due to their high physiochemical and community complexity. To enhance detection of sub-communities of DNA and RNA viruses, we applied fractionation approaches to soils collected across a moisture gradient from a grassland field experiment. Analyses included metagenomics and metatranscriptomics of size-fractionated extracellular viruses (i.e., DNA and RNA viromes), metagenomics of bacteria/archaea- or eukaryote-enriched samples, and whole soil metatranscriptomes with rRNA-depletion or polyadenylation enrichment. While RNA virome and whole soil RNA methods captured similar viral diversity, RNA viromes identified longer, higher-quality genomes. Further, we showed that significantly more DNA viruses were active in higher moisture than lower moisture samples, whereas responses by overall diversity vary by genome type (DNA versus RNA genomes). Finally, we demonstrate the power of fractionation approaches for identifying distinct viral communities that infect unique hosts, which has significant implications for ecological investigations, particularly related to interkingdom interactions.

59 BASIC BIOLOGICAL SCIENCES↗

BSEC VPRM 10m Hourly Biogenic Fluxes in Baltimore (2021)

Model outputs from the Vegetation Photosynthesis and Respiration Model (VPRM: version from Horne et al. in prep). Model remote sensing inputs come from Sential 2-derived EVI and LSWI. Model meteorological inputs for two-meter air temperature and shortwave incoming come from the BSEC WRF 2021 Control Run (Foust, W. 2023). Plant functional Types (PFTs) are spatially classified using the Chesapeake Bay Program 2018 land use land cover product. The final biogenic flux (µmol CO2 m^-2 s^-1) outputs of NEE, RESP, and GEE are a weighted average based on the portion of PFTs within the cell. Individual PFT outputs are saved inside PFT directories (e.g., Crops, Grass, etc.) inside the specific month directory. Model outputs are denoted as a negative flux into the land system (i.e., photosynthesis) and a positive flux as a net release into the overlying atmosphere. Respiration (RESP) fluxes are positive and combine heterotrophic (only soil) and autotrophic sources. Gross ecosystem exchange (GEE) is a negative flux driven by only photosynthetic activity from vegetation, and the Net ecosystem exchange (NEE) is the sum of the two (i.e., NEE=RESP+GEE). Data Characteristics Spatial Resolution: 10m Temporal Resolution: Hourly File Format: VPRM_ _BSEC. .tif (Hour is in UTC) For more information on the model results, please email Jason Horne (jph6488@psu.edu). References: Foust, W. (2023). BSEC WRF 2021 Control Run Output (v0.1.0) [Data set]. MSD-LIVE Data Repository. https://data.msdlive.org/records/m0e6m-vvq17

Baltimore↗

SAXS Assistant: Automated SAXS analysis for structural discovery in biologics and polymeric nanoparticles

Small-angle x-ray scattering (SAXS) is a powerful technique for assessing macromolecular structure. High-throughput SAXS is limited by the time-consuming and, at times, subjective nature of SAXS data interpretation. Here, we present SAXS Assistant, a Python-based script that streamlines SAXS data analysis to extract features for machine learning (ML) and key structural parameters, including the Guinier radius of gyration (R g ), pair distance distribution function (PDDF)-derived R g , maximum particle dimension (D max ), and Kratky plots. The script builds upon BioXTAS RAW and validates reliability via Guinier/PDDF R g agreement, an important indicator of well-measured data sets. For assistance in D max estimation, a multilayer perceptron regressor was trained with 1940 data files from the Small Angle Scattering Biological Data Bank. The model achieved a test set performance R 2 = 0.90 and mean absolute error = 11.7 Å. Training exclusively with experimental data translates analyses from researchers, including experts in the field, to the ML model, which helps assess D max estimations from PDDF. Gaussian mixture model clustering was implemented to classify profiles into structural classes based on entries in the Small Angle Scattering Biological Data Bank. Users may therefore assess the similarity between experimental samples and known biomolecular shapes within the mapped repository entries. This probabilistic clustering aids in quantifying information from Kratky and generating shape-descriptive features. SAXS Assistant accelerates SAXS data analysis through enforced quality control, ML-ready outputs, and flags for low-confidence results. In addition to providing the ability to analyze large data sets at high throughput, this tool is versatile and may serve researchers in both biological and synthetic polymer research fields.

36 MATERIALS SCIENCE↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗