Search NASA⌕ Search

SEARCH · Search NASA

Results for “BINARY DATA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Barge Science Van IMU Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY↗

Raw Data

This dataset contains high-frequency (10Hz) data from the GX5-45 IMU on the Barge Science vans. The data are all raw binary files.

17 WIND ENERGY↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Shared Use Travel Behavior for Improving Rural Mobility: Insights from Greene County, Pennsylvania

Rural communities are considered disadvantaged communities as they suffer from a lack of transport options. Thus, rural regionsprovide less accessibility for commuters to reach their destination as opposed to urban regions. However, the issues of transport disadvantageand shared use mobility in rural areas within the United States (US) have not been well investigated. Furthermore, transport disadvantagediffers between communities and regions across the globe; thus, there is a need to study the behavioral choices of rural commuters within theUS context. This study contributes by analyzing the behavioral choices of rural communities within the US through a case study site ofWaynesburg, Pennsylvania, for adopting a shared use shuttle service. K-means clusters showed that trips from the survey data were a goodrepresentation of real trips from Ecolane. Furthermore, random parameter-based binary logit models were calibrated using data collected fromstudents, faculty, and residents in Waynesburg, Greene County, to study the behavioral choices of commuters. The findings for the faculty andstudents group revealed that prior experience with shared services increases the likelihood of using a shared shuttle. An important personalcharacteristic of inconvenience showed a higher propensity toward using existing modes as opposed to a shared shuttle. Such commutersvalue personal vehicles as more convenient as they have childcare responsibilities and varying schedules for work that require them to moveback and forth across locations, thus making a shared shuttle less attractive for them. The socioeconomic factors of age and gender show ahigher propensity for using shared shuttles. Furthermore, the findings from this study could be helpful for agencies in improving rural mobility andconsidering such shared mobility services for rural communities

42 ENGINEERING↗

Relativistic gas accretion onto supermassive black hole binaries from inspiral through merger

Accreting supermassive black hole binaries are powerful multimessenger sources emitting both gravitational and electromagnetic (EM) radiation. Understanding the accretion dynamics of these systems and predicting their distinctive EM signals is crucial to informing and guiding upcoming efforts aimed at detecting gravitational waves produced by these binaries. To this end, accurate numerical modeling is required to describe both the spacetime and the magnetized gas around the black holes. In this paper, we present two key advances in this field of research. First, we have developed a novel 3D general relativistic magnetohydrodynamics (GRMHD) framework that combines multiple numerical codes to simulate the inspiral and merger of supermassive black hole binaries starting from realistic initial data and running all the way through merger. Throughout the evolution, we adopt a simple but functional prescription to account for gas cooling through photon emission. Next, we have applied our new computational method to follow the time evolution of a circular, equal-mass, nonspinning black hole binary for ∼200 orbits, starting from a separation of 20⁢𝑟 𝑔 and reaching the postmerger evolutionary stage of the system. We have shown how mass continues to flow toward the binary even after the binary “decouples” from its surrounding disk, but the accretion rate onto the black holes diminishes. We have identified how the minidisks orbiting each black hole are slowly drained and eventually dissolve as the binary compresses. We confirm previous findings that the system’s luminosity decreases by a factor of a few during inspiral; however, we observe an abrupt increase by ∼50% in this quantity at the time of merger, likely accompanied by an equally abrupt change in spectrum. Lastly, we have demonstrated that during the inspiral, fluid ram pressure regulates the fraction of the magnetic flux transported to the binary that attaches to the black holes’ horizons.

Accretion disk & black-hole plasma↗

Periodicity significance testing with null-signal templates: reassessment of PTF’s SMBH binary candidates

Periodograms are widely employed for identifying periodicity in time series data, yet they often struggle to accurately quantify the statistical significance of detected periodic signals when the data complexity precludes reliable simulations. We develop a data-driven approach to address this challenge by introducing a null-signal template (NST). The NST is created by carefully randomizing the period of each cycle in the periodogram template, rendering it non-periodic. It has the same frequentist properties as a periodic signal template, and we show with simulations that the distribution of false positives is the same as with the original periodic template, regardless of the underlying data. Thus, performing a periodicity search with the NST acts as an effective simulation of the null (no-signal) hypothesis, without having to simulate the noise properties of the data. We apply the NST method to the supermassive black hole binaries (SMBHB) search in the Palomar Transient Factory (PTF), where Charisi et al. had previously proposed 33 high signal-to-noise candidates utilizing simulations to quantify their significance. Our approach reveals that these simulations do not capture the complexity of the real data. There are no statistically significant periodic signal detections above the non-periodic background. To improve the search sensitivity, we introduce a Gaussian quadrature based algorithm for the Bayes Factor with correlated noise as a test statistic. We show with simulations that this improves sensitivity to true signals by more than an order of magnitude. However, the Bayes Factor approach also results in no statistically significant detections in the PTF data.

79 ASTRONOMY AND ASTROPHYSICS↗

eDNAjoint: An R package for interpreting paired or semi‐paired environmental DNA and traditional survey data in a Bayesian framework

Abstract Environmental DNA (eDNA) sampling is increasingly used in surveys of species distribution as a potentially sensitive and efficient monitoring method. Yet access to modelling tools designed specifically for interpreting this new data type lags behind its ubiquity. While occupancy modelling software has dominated the analytical landscape for eDNA data analysis of single species, this type of model may not always be the most appropriate. The rate of eDNA detection often corresponds to species density, rather than just occupancy, and researchers often have access to observations from non‐genetic sampling methods at the same sites. To provide users access to a modelling framework designed to maximize the use of all available data, we developed an R package, eDNAjoint . The package provides an easy‐to‐use interface for fitting a ‘joint’ model that integrates data from paired or semi‐paired eDNA and traditional surveys in a Bayesian framework. The model can be used to estimate parameters like the probability of a false positive eDNA detection and mean catch rate at a site, and the package allows access to multiple model variations and Bayesian prior customization. Additional functionality can be used for model selection, summarising posteriors and comparing the relative sensitivities of the two survey methods. We demonstrate the use of eDNAjoint by fitting a variation of the model with site‐level covariates that scale the sensitivity of eDNA sampling relative to traditional sampling. The example workflow uses binary eDNA and seine count data for the endangered tidewater goby ( Eucyclogobius newberryi ) from a study by Schmelzle and Kinziger (2016). This use case includes a prior sensitivity analysis and an evaluation of the relationship between detection rates and environmental variables. eDNAjoint has the potential to greatly increase the range of users who will be able to rigorously analyse eDNA and traditional survey data in a Bayesian framework, understand if and how eDNA can improve monitoring practices, and gain confidence in the interpretability of eDNA data.

Keller, Abigail G. [Department of Environment Scie↗

Do black holes remember what they are made of?

We study the ringdown signal of black holes formed in prompt-collapse binary neutron star mergers. We analyze data from 47 numerical relativity simulations. We show that the ($l$ = 2, $m$ = 2) and ($l$ = 2, $m$ = 1) multipoles of the gravitational wave signal are well fitted by decaying damped exponentials, as predicted by black-hole perturbation theory. We show that the ratio of the amplitude in the two modes depends on the progenitor binary mass ratio q and reduced tidal parameter $\overline{Λ}$. Unfortunately, the numerical uncertainty in our data is too large to fully quantify this dependency. If confirmed, these results will enable novel tests of general relativity in the presence of matter with next-generation gravitational-wave observatories.

79 ASTRONOMY AND ASTROPHYSICS↗

Chandra Discovery of a Candidate Hyperluminous X-Ray Source in MCG+11-11-032

We present a multiwavelength analysis of MCG+11-11-032, a nearby active galactic nucleus (AGN), with a unique classification as being both a binary and a dual AGN candidate. With new Chandra observations, we aim to resolve any dual AGN system via imaging data and search for signs of a binary AGN via analysis of the X-ray spectrum. Analyzing the Chandra spectrum, we find no evidence of the previously suggested double-peaked Fe Kα lines; the spectrum is instead best fit by an absorbed power law with a single Fe Kα line, as well as an additional line centered at ≈7.5 keV. The Chandra observation reveals faint, soft, and extended X-ray emission, possibly linked to low-level nuclear outflows. Further analysis shows evidence for a compact hard source—MCG+11-11-032 X2—located 3.″3 from the primary AGN. Modeling MCG+11-11-032 X2 as a compact source, we find that it is relatively luminous (L2–10 keV=1.5−0.5+0.9×1041erg s$^{−1}$), and the location is coincident with a compact and off-nuclear source resolved in Hubble Space Telescope infrared (F105W) and optical (F621M, F547M) bands. Pairing our X-ray results with a 144 MHz radio detection at the host galaxy location, we observe X-ray and radio properties similar to those of ESO 243-49 HLX-1, suggesting that MCG+11-11-032 X2 may be a hyperluminous X-ray source. This detection with Chandra highlights the importance of a high-resolution X-ray imager as well as how previous binary AGN candidates detected with large-aperture instruments can benefit from high-resolution follow-up. Future spatially resolved optical spectra, and deeper X-ray observations, can better constrain the origin of MCG+11-11-032 X2.

79 ASTRONOMY AND ASTROPHYSICS↗

The influence of cloud cover on the reliability of satellite-based solar resource data

Satellite-based solar resource data are often developed and validated by using binary cloudiness categories: clear sky or overcast cloudy sky. To investigate the reliability of solar resource data in partially cloudy conditions, we estimate cloud fraction using two distinct algorithms: a physical retrieval model using surface observed global horizontal irradiance (GHI) and direct normal irradiance (DNI) and a temporal average of cloud mask data estimated by the observed DNI. Our analysis reveals a significant presence of scattered clouds, broken clouds, and mismatches between satellite- and surface-based cloud data at 17 surface sites across the contiguous United States, though confidently clear and cloudy conditions collectively account for more than 70 % of the data. Solar radiation is computed using the National Solar Radiation Database (NSRDB) algorithm and validated using surface observations. Here, our findings suggest that, in the presence of scattered clouds, NSRDB data for clear-sky conditions can be subject to significant overestimation. In cloudy-sky conditions classified by satellite data, DNI computed by the Fast All-sky Radiation Model for Solar applications with DNI (FARMS-DNI) can be underestimated when limited clouds are detected by surface observations. The bias observed in several cloudiness categories indicates that the NSRDB is exceptionally accurate in confidently clear conditions. However, clear-sky conditions with scattered clouds and mismatched cloud data contribute significantly to the overall uncertainties in the NSRDB. Therefore, future improvements in solar resource data should involve development and implementation of satellite-derived cloud fraction and should consider a novel radiative transfer model accounting for amplified cloud reflection. The evaluation within cloudiness categories also provides a physical rationale for the superior performance of FARMS-DNI compared to the Direct Insolation Simulation Code (DISC) in both cloudy-sky and all-sky conditions.

14 SOLAR ENERGY↗

Quantifying Epistemic Uncertainty in Binary Classification via Accuracy Gain

ABSTRACT Recently, a surge of interest has been given to quantifying epistemic uncertainty (EU), the reducible portion of uncertainty due to lack of data. We propose a novel EU estimator in the binary classification setting, as the posterior expected value of the empirical gain in accuracy between the current prediction and the optimal prediction. In order to validate the performance of our EU estimator, we introduce an experimental procedure where we take an existing dataset, remove a set of points, and compare the estimated EU with the observed change in accuracy. Through real and simulated data experiments, we demonstrate the effectiveness of our proposed EU estimator.

97 MATHEMATICS AND COMPUTING↗

Forensic Analysis of SOHO Router Binaries

Small Office/Home Office (SOHO) routers are used by millions of consumers across the United States, and are commensurately vulnerable. Forensic analysis of SOHO router firmware helps to understand and mitigate those vulnerabilities. This poster focused particularly on analysis of BusyBox executables, a software suite that provides several Unix utilities in a single file. Three main tools were used to analyze the binaries. BinWalk was used to extract the files, but also to build entropy graphs, extract Linux kernel images, and identify CPU architectures; WiiBin processed the binaries to find endianness, architecture, the percent compressed/encrypted, and compiler data; and @DisCo, a machine learning tool used to determine function similarity in disassembled binaries, analyzed similarities and determined versions of extracted BusyBox files from each router. These tools found that venders from all five routers utilized the same version of the BusyBox software across different firmware updates, demonstrating the importance of constant firmware scrutiny to protect against security vulnerabilities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

System for controller area network payload decoding

A system for decoding an unknown automotive controller area network (“CAN”) message definitions. CAN data vehicle signal mappings are typically held in secret and varied by automotive model and year. Without knowledge of the mappings, the wealth of real-time vehicle data hidden in the automotive CAN packets is uninterpretable—impeding research, after-market tuning, efficiency and performance monitoring, fault diagnosis, and privacy-related technologies. This system can ascertain the CAN signals' boundaries (start bit and length), endianness (byte ordering), signedness (binary-to-integer encoding) from raw CAN data. This allows conversion of CAN data to time series. Interpreting the translated CAN data's physical meaning and finding a linear mapping to standard units (e.g., knowing the signal is speed and scaling values to represent units of miles per hour) can be achieved for many signals by leveraging diagnostic standards to obtain real-time measurements of in-vehicle systems. The system can be integrated into lightweight hardware enabling an OBD-II plugin for real-time in-vehicle CAN decoding or run on standard computers. The system can output a standard DBC file with the signal definition information.

Verma, Kiren E.↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗

NUM-DAT File Format Specification: Used in M-9 Gun Experiment Data Archiving

The M-9 Shock and Detonation Physics group executes experiments on gun and explosive platforms with large numbers of oscilloscopes used for data acquisition. The data acquisition from these oscilloscopes was automated many years ago using a custom piece of software called RunDig . The default save format from this software is a custom structure referred to as "NUM-DAT" format. This file format includes a text ".DAT" file which is a header file used to interpret the binary ".NUM" file which contains the oscilloscope data. The data save format was originally developed by John Vorthman and has been in use by M-9 personnel for over 20 years. This data format has been used for archiving data from experiments performed by M-9 personnel at TA-40, TA-39, and the TA-55 Impact Test Facility. Numerous custom analysis and visualization programs have also been developed, and continue to be used, that utilize this data format. This document describes the NUM-DAT format and provides code examples for reading the format and converting it to other formats.

47 OTHER INSTRUMENTATION↗

Predictive analytics of selections of russet potatoes

We explore the application of machine learning algorithms specifically to enhance the selection process of Russet potato (Solanum tuberosum L.) clones in breeding trials by predicting their suitability for advancement. This study addresses the challenge of efficiently identifying high-yield, disease-resistant, and climate-resilient potato varieties that meet processing industry standards. Leveraging manually collected data from trials in the state of Oregon, we investigate the potential of a wide variety of state-of-the-art binary classification models. The dataset includes 1086 clones, with data on 38 attributes recorded for each clone, focusing on yield, size, appearance, and frying characteristics, with several control varieties planted consistently across four Oregon regions from 2013 to 2021. We conduct a comprehensive analysis of the dataset that includes preprocessing, feature engineering, and imputation to address missing values. We focus on several key metrics such as accuracy, F1-score, and Matthews correlation coefficient (MCC) for model evaluation. The top-performing models, namely a feedforward neural network classifier (Neural Net), a histogram-based gradient boosting classifier (HGBC), and a support vector machine classifier (SVM), demonstrate consistent and significant results. To further validate our findings, we conducted a simulation study using the aims, data-generating mechanisms, estimands, methods, and performance measures (ADEMP) framework, simulating different data-generating scenarios to assess model robustness and performance through true positive, true negative, false positive, and false negative distributions, area under the receiver operating characteristic curve (AUC-ROC) and MCC. The simulation results highlight that non-linear models like SVM and HGBC consistently show higher AUC-ROC and MCC than logistic regression, thus outperforming the traditional linear model across various distributions, and emphasizing the importance of model selection and tuning in agricultural trials. Variable selection further enhances model performance and identifies influential features in predicting trial outcomes. The findings emphasize the potential of machine learning in streamlining the selection process for potato varieties, offering benefits such as increased efficiency, substantial cost savings, and judicious resource utilization. Our study contributes insights into precision agriculture and showcases the relevance of advanced technologies for informed decision-making in breeding programs.

60 APPLIED LIFE SCIENCES↗

SESAME: ASCII2 File Format

A new ASCII format for SESAME data is explicitly defined and dubbed ASCII2. This format fixes some of the onerous limitations of the legacy ASCII-styled format in addition to relaxing the FORTRAN style fixed-format layout of data tables. With the exception of the multi-tiered indexes of the legacy binary format, the new ASCII2 is more compatible with the binary in that it does not restrict the extent of various integer data (e.g., SESAME material and table numbers) to six digits nor does it restrict the extent or precision of the various floating point data. There will be very limited support of the legacy ASCII-styled formats in the future.

36 MATERIALS SCIENCE↗