Search NASA⌕ Search

SEARCH · Search NASA

Results for “data services”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

A Privacy First Path Analysis using Clickstream Data

In the modern digital economy, data-driven decision making is crucial for effectively meeting the ever-evolving demands of consumer engagement and satisfaction. Clickstream data has become invaluable for understanding customer behavior, yet concerns over privacy and security persist, especially with some internet service providers profiting from its sale. This article introduces an innovative methodology that blends experiential learning with advanced cryptographic techniques, including differential privacy and graph analytics. The core objective of this methodology is to estimate Customer Lifetime Value (CLV) by analyzing clickstream data, achieving an average prediction accuracy of 92.4% in user engagement levels while ensuring user anonymity through Recency, Frequency, and Monetary (RFM) analysis. Our study introduces the concept of a “data depositor” and a privacy manager, employing the composition theorem to merge non-adaptive queries effectively. Privacy budgets (? = 1.0, d = 10-5), sensitivity-specific techniques, and data partitioning were applied. Randomization and noise addition protect data integrity, with special handling for categorical values. This approach, differing from prior studies, offers a 12.6% improvement in privacy-preserving targeting accuracy while maintaining strict confidentiality, presenting a novel path forward in data-driven decision-making.

Frequency and Monetary (RFM) analysis↗

Expanding Access to Science Participation: A FAIR Framework for Petascale Data Visualization and Analytics

The massive data generated by scientists daily serve as both a major catalyst for new discoveries and innovations, as well as a significant roadblock that restricts access to the data. Here, our paper introduces a new approach to removing Big Data barriers and democratizing access to petascale data for the broader scientific community. Our novel data fabric abstraction layer allows user-friendly querying of scientific information while hiding the complexities of dealing with file systems or cloud services. We enable FAIR (Findable, Accessible, Interoperable, and Reusable) access to datasets such as NASA’s petascale climate datasets. Our paper presents an approach to managing, visualizing, and analyzing petabytes of data within a browser on equipment ranging from the top NASA supercomputer to commodity hardware like a laptop. Our novel data fabric abstraction utilizes state-of-the art progressive compression algorithms and machine-learning insights to power scalable visualization dashboards for petascale data. The result provides users with the ability to identify extreme events or trends dynamically, expanding access to scientific data and further enabling discoveries. We validate our approach by improving the ability of climate scientists to visually explore their data via three fully interactive dashboards. We further validate our approach by deploying the dashboards and simplified training materials in the classroom at a minority-serving institution. These dashboards, released in simplified form to the general public, contribute significantly to a broader push to democratize the access and use of climate data.

Computer science↗

Globus service enhancements for exascale applications and facilities

Many extreme-scale applications require the movement of large quantities of data to, from, and among leadership computing facilities, as well as other scientific facilities and the home institutions of facility users. These applications, particularly when leadership computing facilities are involved, can touch upon edge cases (e.g., terabyte files) that had not been a focus of previous Globus optimization work, which had emphasized rather the movement of many smaller (megabyte to gigabyte) files. We report here on how automated client-driven chunking can be used to accelerate both the movement of large files and the integrity checking operations that have proven to be essential for large data transfers. In conclusion, we present detailed performance studies that provide insights into the benefits of these modifications in a range of file transfer scenarios.

97 MATHEMATICS AND COMPUTING↗

SOMA: Observability, monitoring, and in situ analytics for exascale applications

With the rise of exascale systems and large, data-centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service-based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application-specific diagnostic data and performance data addresses these needs. Furthermore, we present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi-GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.

97 MATHEMATICS AND COMPUTING↗

WRS Capabilities Booklet [Slides]

WRS is the digital backbone of the Weapons Program—delivering trusted data assets, cyber-assured software and systems, and AI-enabling software—that transform insights into decisive action. We empower physicists, engineers, researchers, and scientists to think faster, act strategically, and stay ahead in an ever-evolving threat landscape. Our efforts ensure critical nuclear weapons data remains secure, accessible, and usable—supporting mission-critical work, informed decision making, and scientific advancement at LANL and across the Nuclear Security Enterprise (NSE).

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Progress Update on the In Situ Load Retention Aging Vessel

To improve upon our conventional thermally accelerated aging study methods which require periodic interruption of aging to perform load testing in an Instron machine, a new aging system is being developed to: (1) automate/facilitate data acquisition/analysis, (2) improve data quality, and (3) enable uninterrupted aging conditions (compression, temperature, atmosphere) which represents the service condition. In FY24, a thermal aging vessel instrumented with load cells was fabricated and tested. That unit demonstrated the primary function of the vessel: to continuously monitor the in situ load retention of up to three compressed polymer coupons undergoing thermally accelerated aging under nitrogen. A secondary objective, not realized in the unit due to inadequate sealing, was to enable a single initial nitrogen backfill (as opposed to a continuous purge) and gas sampling of the vessel headspace during thermal aging. Heating of the vessel was achieved using a custom heater jacket.

36 MATERIALS SCIENCE↗

Deny-by-Default Network Port Security: SPaRC Technical Bulletin #002

Operational Technology (OT) networks [e.g., industrial control systems (ICS) and supervisory control and data acquisition (SCADA) systems] have unique cyber security challenges due to their decades long service life, high availability requirements, and limited visibility. OT networks often take credit for being “air gapped” (i.e. disconnected from the Internet) and all devices within the OT network can “talk” to each other—even if they should not. This SPaRC Technical Bulletin describes how the unique limitations of OT networks can become strengths when it comes to cybersecurity.

Cybersecurity↗

A Framework to Demonstrate a DNP3 Interface With a CIM-Based Data Integration Platform: Preprint

The contemporary electrical grid is characterized by its complexity and abundance of data. A control-rich environment supported by information and communication technologies within an Advanced Distribution Management System (ADMS) presents a viable and cost-effective option for utility companies aiming to implement advanced real-time analytical schemes for monitoring and remotely controlling distribution feeders. Modular platform-based approaches to distribution operations require a structured framework for acquiring field device measurements, performing analytics, converting the setpoint to the correct protocol, and sending it on the appropriate communications network to the field devices. We present the development and deployment of an application service to integrate an open-source standardsbased platform with an ADMS test bed with field devices using the Distributed Network Protocol (DNP3) for data exchange. The step-by-step procedure for establishing the DNP3-Master service on an open-source distribution platform is outlined, comprehensively explaining the Master setup process. Moreover, sample use case results highlight the capabilities of the DNP3- Master service setup. Results demonstrate the scalability and configurability of the DNP3-Master service, making it adaptable for integration with other relevant applications, thus providing potential opportunities for real-world field trials and real-time assessments.

ADMS↗

An On-Demand Electric Transit Case Study of New Rochelle, New York

This work explores the extent to which an on-demand mobility service utilizing lightweight electric vehicles (EVs) provides community and sustainability benefits in New Rochelle, New York. Travel and survey data from September 2019 through 2023 are used to describe the system and estimate impacts on travelers. The system was found to be used more by women (nearly 60%) and younger demographics (>65% under the age of 42), with peak use in the middle of the day and a grocery store as a top origin and destination. The service is utilized primarily for short trips (86% under 2 miles), and the small, right-sized EVs have carbon dioxide emissions associated with charging the fleet that are roughly one-quarter of the fleet emissions of conventional hybrid vans and nearly 50 times less than a fleet of diesel buses. Mapping current socio-spatial dynamics of travel demand can inform equity performance, as well as assist future planning and service area development and possible extensions to similar smaller, lower-density environments that are nearby and connected to major metropolitan areas. The findings in this case study suggest on-demand electric transit may be a significant and growing space for advancing clean and highly valued public mobility services. Sustainable public transport interventions that consider right-sized, electric, on-demand vehicles can help achieve improved accessibility and reduce energy use and greenhouse gas emissions. This work was presented at the Transportation Research Board (TRB) 2025 Annual Meeting on January 7, 2025.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Plant Engineers Solar Energy Handbook: Southern California Region

Discussed in order after the introduction are solar components and systems (collectors, storage, service hot water systems, space heating with liquid and air systems, space cooling, heat pumps and controls); computer programs for system optimization; local solar and weather data; a description of buildings and plants in Southern California applying solar technology; current Federal and California solar legislation; standards, codes and performance testing information; a listing of manufacturers, distributors, and professional services available in Southern California region; and information access. Finally, solar design check lists for those engineers who wish to design their own systems. The program for the Solar Workshop for the Plant Engineer, March 30, 1978, Los Angeles, California is included.

14 SOLAR ENERGY↗

1995 OKI Household Activity and Travel Report Survey

The OKI Household Activity and Travel Report Survey 1995, conducted by Market Opinion Research, gathered household activity and travel data from 3,000 households within the OKI region (Ohio, Kentucky, and Indiana Council of Governments' service area). The region covers over 2,615 square miles and is home to 697,468 households (over 1.75 million people). The primary purpose of this study was to provide OKI with a new database of travel patterns and behavior to assist in updating the region's transportation models. Respondents were asked to report their activities for a 24-hour period from 3:00 a.m. to 3:00 a.m. on an assigned activity day. The survey included weekday assignments only from October 4, 1995, to November 30, 1995.

1Hz data↗

Investigating the Determinants of Household Capabilities Burden During Power Outages: The Case of Winter Storm Uri

Existing research primarily uses census data to identify the vulnerability of communities to hazards. These vulnerability indices provide aggregated data and are not hazard-specific nor well-validated with post-event data. In contrast, our study uses household survey data (n=1065) to understand which Texan households suffered the greatest loss of their capabilities due to power outages and other utility service disruptions during Winter Storm Uri. Inspired by the Capabilities Approach, our measures of burden include the number of household capability types disrupted during the outages (e.g., cooking, heating, refrigeration), the severity of impact for each disrupted capability, and the additional time and financial costs of coping with these disruptions. We perform a clustering analysis, and find two distinct groups in our data, consisting of ‘lesser burden' and ‘heavier burden' households. Results indicate that the households experiencing the heaviest capabilities burden were most likely to experience longer power outages and the loss of water services. They were also more likely to have a Hispanic-Latino household member, lack access to a generator, live in a rented home, have larger households with more young children, fewer adults over 65, lower household incomes, been impacted by the COVID-19 pandemic, and more family characteristics that made life harder. We also fit a logistic regression model to assess the role of outage, household, and community characteristics in predicting differences in capabilities burden. Our results offer insights into enumerating the consequences of utility service disruptions on households, which can inform more targeted and equitable resilience strategies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Foundational Industry Energy Dataset: Unit-level Characterization and Derived Energy Estimates for Industrial Facilities in 2017

The Foundational Industry Energy Dataset (FIED) addresses several of the areas of growing disconnect between the demands of industrial energy analysis and the state of industrial energy data by providing unit-level characterization by facility. Each facility is identified by a unique registryID, based on the U.S. Environmental Protection Agency (EPA) Facility Registry Service, and includes its coordinates and other geographic identifiers. Energy-using units are characterized by design capacity, as well as their estimated energy use, greenhouse gas emissions, and physical throughput using 2017 data from the EPA's National Emissions Inventory and Greenhouse Gas Reporting Program. An overview of the derivation methods is provided in a separate technical report which will be linked after publication. The Python code used to compile the dataset is available in a GitHub repository. An updated 2020 version is under development.

Array↗

G2PDeep-v2: A Web-Based Deep-Learning Framework for Phenotype Prediction and Biomarker Discovery for All Organisms Using Multi-Omics Data

Multi-omics data offers rich insights into complex traits across organisms, yet integrating and analyzing these datasets for phenotype prediction and marker discovery remains challenging. Researchers need accessible tools that combine deep learning, hyperparameter optimization, visualization, and downstream analysis in a unified web platform. To address this, we developed G2PDeep-v2, a web-based platform powered by deep learning for phenotype prediction and marker discovery from multi-omics data across a wide range of organisms, including humans and plants. The server provides multiple services for researchers to create deep-learning models through an interactive interface and train these models using an automated hyperparameter tuning algorithm on high-performance computing resources. Users can visualize the results of phenotype and markers predictions and perform Gene Set Enrichment Analysis for the significant markers to provide insights into the molecular mechanisms underlying complex diseases, conditions and other biological phenotypes being studied.

59 BASIC BIOLOGICAL SCIENCES↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗

CMIP7 data request: impacts and adaptation priorities and opportunities

The Coupled Model Intercomparison Project Phase 7 (CMIP7) undertook an extensive process to gather community input and refine data requests related to impacts and adaptation applications of Earth System Model (ESM) outputs. The Impacts and Adaptation (I&A) Data Request Team worked with CMIP7 leadership to distribute an open solicitation across many communities that use climate model outputs requesting inputs for new and existing variables, the most applicable temporal characteristics, and groupings of variables that together allow for specific application opportunities. This input was then collated and translated into CMIP7 standard templates for inclusion in the broader data request, leading to 13 I&A data request opportunities, 60 variable groups and 539 unique variables sought by vulnerability, impacts, adaptation, and climate services user communities. Here, we describe these opportunities and variable groups, as well as new insights into how ESM groups can prioritize outputs that set off a chain of further analyses, ultimately informing decisions impacting society and natural systems. These include an emphasis on high-resolution outputs to allow further modeling of climate impacts at regional and local scales, improved representation of extreme weather events, enhanced accuracy of downscaling and bias-adjustment techniques, and support for more detailed assessments for decision-making in adaptation and mitigation strategies. There is also broad interest in more extensive provisioning of two-dimensional variables at the Earth's surface, prioritizing experiments that enhance our understanding of both the recent past and future scenarios, and providing outputs that allow further downscaling and bias adjustment. We emphasize that variable groups are the fundamental level at which to engage with the I&A data request, matching the scale of input and the way output provision enables specific I&A applications. Given resource constraints, we applaud CMIP7 efforts to foster strong engagement and communication between ESM groups and the I&A team to build consensus around prudent compromises in priority variables, temporal resolutions, simulation experiments, time subsets, and ensemble members.

Ruane, Alex C. [NASA Goddard Inst. for Space Studi↗