Search NASA⌕ Search

SEARCH · Search NASA

Results for “Dataframes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

How R Developers explain their Package Choice: A Survey

Background: Contemporary software development relies heavily on reusing already implemented functionality, usually in the form of packages. Aims: We aim to shed light on developers' preferences when selecting packages in R language. Method: To do that, we create and administer a survey to over 1000 developers who have added one of two common dataframe enhancement libraries in R to their projects: data.table or tidyr. We design a questionnaire using the Social Contagion Theory (SCT) following prior work on technology adoption and ensure that key dimensions affecting developer choice are considered. Results: Of the 1085 developers we contacted, 803 completed the survey asking them to prioritize various factors known to affect developer perceptions of package quality and to provide their background. Most developers self-identified as data scientists with two to five years of work experience. We found significant differences between the preferences of developers who chose data.table and tidyr. Surprisingly, package reputation based on easy-to-see measures, such as the number of stars on GitHub, was not an important factor for either group. Conclusions: Our findings demonstrate the inherently social nature of package adoption. They can help design future studies on how different populations of developers make decisions on which software packages to use in their projects. Finally, package developers and maintainers can benefit by better understanding the prime concerns of the users of their packages.

Malviya Thakur, Addi↗

RouteE-Powertrain [SWR-19-19]

RouteE-Powertrain is a tool for predicting energy usage over a set of road links. RouteE-Powertrain is a Python package that allows users to work with a set of pre-trained mesoscopic vehicle energy prediction models for a varity of vehicle types. Additionally, users can train their own models if "ground truth" energy consumption and driving data are available. RouteE-Powertrain models predict vehicle energy consumption over links in a road network, so the features considered for prediction often include traffic speeds, road grade, turns, etc. The typical user will utilize RouteE's catalog of pre-trained models. Currently, the catalog consists of light-duty vehicle models, including conventional gasoline, diesel, hybrid electric (HEV), and battery electric (BEV). These models can be applied to link-level driving data (in the form of pandas dataframes) to output energy consumption predictions. Users that wish to train new RouteE models can do so. The model training function of RouteE enables users to use their own drive-cycle data, powertrain modeling system, and road network data to train custom models. https://pypi.org/project/nrel.routee.powertrain/ pip install nrel.routee.powertrain

Holden, Jacob↗

BuildStockQuery [SWR-23-58]

BuildStockQuery is a python library designed to simplify and streamline the process of querying massive, terabyte-scale datasets generated by ResStock(TM). ResStock (SWR-19-15) is a U.S. DOE-supported, NREL-built, national residential building energy stock model that enables a new approach to large-scale residential energy analysis across the U.S. by combining large public and private data sources, statistical sampling, detailed sub-hourly building simulations, and high-performance computing. BuildStockQuery offers an intuitive Object-Oriented Programming (OOP) interface to the ResStock output dataset allowing users to easily perform common queries and receive results in familiar pandas DataFrame format, abstracting away the need for complex SQL query. By initializing a query object with the pertinent Athena database and table names, users can easily query for various kinds of insights, for example, timeseries electricity for an end use for a given state grouped by building types.

Adhikari, Rajendra↗

polars-dovmed (dovmed) v0.1.0

polars-dovmed is a python package for text search and extraction from NCBI's PubMed Central Open Access subset. It is powered by the polars dataframe library and leverages modern file formats (parquet) to efficiently scan public literature.

Roux, Simon [Lawrence Berkeley National Laboratory↗

GRIDAPPSD/distopf (33583-E)

DistOPF is an open-source Python package providing a three-phase, asymmetric optimal power flow (OPF) tool specifically designed for distribution systems. The key inventive features include: - Asymmetrical 3-phase OPF modeling for distribution systems with unbalanced phases - Comprehensive control optimization supporting both active (P) and reactive (Q) power control variables - Built-in visualization and validation tools - Standard test system benchmarking platform for algorithm development and comparison - Modular CSV-based input system using Pandas DataFrames for flexible model specification - Standard power distribution model importer enabling direct conversion from CIM and OpenDSS format to optimization-ready models - Multiple solve interface compatibility (PYOMO, CVXPY, SciPy) with automatic solver selection based on problem type

Gray, Nathan [Pacific Northwest National Laborator↗

Hydrologic Model Data for the East Fork Poplar Creek Watershed Simulated with the Advanced Terrestrial Simulator (ATS): Streamflow and Network Expansion–Contraction Dynamics

This dataset supports hydrologic modeling and stream network expansion–contraction analysis for the East Fork Poplar Creek (EFPC) Watershed in Tennessee. It includes a Jupyter notebook for model setup, model configuration files, simulation outputs, and derived products used to evaluate model performance and investigate stream dynamics under varying hydrologic conditions. The dataset was generated using the Watershed Workflow Python package and the Advanced Terrestrial Simulator (ATS), enabling integrated surface–subsurface hydrologic simulations using a stream-aligned mesh. Outputs include high-resolution time series of streamflow, active network length, water table depth, and related hydrologic variables. Also included are spatially explicit stream persistency indices and classifications of reaches as perennial or non-perennial. These data facilitate reproducibility and support further research on stream intermittency and variability in network extent.The model data archive is organized in following directories:1) model_setup_inputsContains the Watershed Workflow Jupyter notebooks (accessed through any open source code editor), selected input datasets, and resulting ATS input files, including XML files (access through any open source code editor), computational mesh (.exo files can be viewed using Paraview), and meteorological forcing files (.h5 files can be accessed through h5py python package and HDFView open source software). 2) model_outputsIncludes ATS simulation outputs relevant to this study. Time series of spatially integrated or averaged variables (e.g., streamflow, water table depth) are provided as CSV files. Select spatial fields (e.g., ponded depth and water table depth) are saved as pickled Python objects to reduce file size, and can be accessed through pickle package in Python. Key geometry objects from Watershed Workflow—such as the surface mesh and river tree—are also included to support analysis of streamflow persistency and expansion–contraction dynamics. These files can also be accessed through Watershed Workflow Python package.3) model_evaluationProvides observed streamflow time series and field survey-based flow regime classifications used to evaluate model performance. Jupyter notebooks for processing ATS outputs and comparing model predictions with observations to build confidence in the model prior to scientific analysis are also included.4) Q_L_relationshipsContains workflows for generating time series of discharge, active network length, and related hydrologic variables used in the stream network expansion–contraction analysis. Includes routines for delineating baseflow-dominated periods. For each catchment, notebooks and processed data (as pickled DataFrames accessed through Pandas Python package) are provided. 5) figure_scriptsProvides the Jupyter notebooks used to generate the figures presented in the paper.

54 ENVIRONMENTAL SCIENCES↗

Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model: Modeling Archive

This archive contains the data and Python scripts required to reproduce the analyses and figures in the study: Gomez-Velez, J. D., Rathore, S. S., Cohen, M. J., & Painter, S. L. (2025). Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model. Submitted to Water Resources Research. The analysis utilizes the subgrid model Advection Dispersion Equation with Lagrangian Subgrids (ADELS) implemented in the Advanced Terrestrial Simulator (ATS; https://amanzi.github.io/ats/stable/). In this case, the ATS and Amanzi versions are (1) ATS version 1.5.1_f5ba18f8 and (2) Amanzi version 1.6-dev_53444cca4. The repository includes a Jupyter Notebook and the necessary data (Pandas DataFrames stored as pickle files) to generate the figures for the manuscript. Additionally, it contains Python scripts to create ATS input files, run the ATS simulations, and post-process the results. Finally, it provides routines for parameter estimation using the Single-Station Metabolism (SSM) model with the Differential Evolution Adaptive Metropolis (DREAM) Markov Chain Monte Carlo (MCMC) algorithm with ZS enhancements (DREAM-ZS).

54 ENVIRONMENTAL SCIENCES↗

Chemical Blast Standard (1 kg)

Chemical explosions create blast waves with large overpressure disturbances. It is important to develop a standard blast model based on data to accurately predict acoustic blast-wave amplitudes near detonations and invert for explosion energy from distant observations of blast-wave signals. However, open data from large, controlled chemical explosions with reliable ground truth can be challenging to find. The lack of access to such data could limit the number of contributions to related research and potentially stifle the rate of discoveries or validation of existing models. Here, to address these data scarcity problem, we have curated and compiled a standardized set of 817 blast-wave waveforms from 19 distinct high-explosive events. The blast-wave waveforms are standardized to a 1 kg trinitrotoluene explosion using scaling laws and corrections for location effects. A brief overview of the dataset is presented along with explosion feature models as well as recommendations for extracting explosion features. The resulting dataset is distributed to an open repository in both Seismic Analysis Code and pandas DataFrame formats containing the waveforms, the scaled distances, and the sample rates.

58 GEOSCIENCES↗

PowerAnalytics.jl: User-Centric Power Systems Analysis in Julia

The National Laboratory of the Rockies recently released version 1 of PowerAnalytics.jl, an analysis module for the outputs of its popular open-source electrical power systems modeling platform Sienna. It features an extensible framework - based on the flexible selecting of components, the execution of arbitrary metrics on them, and a familiar DataFrames-based output interface with embedded metadata - to process results in the Sienna style while keeping the interface as simple as possible for non-Julia experts. Here, I describe the package and where it fits into the Sienna ecosystem, how I harnessed user-centered design and Julia features to achieve beginner friendliness without sacrificing performance and expressibility, and what lessons might be drawn from the package's design and implementation.

97 MATHEMATICS AND COMPUTING↗

A Modular Framework for Integrating and Visualizing Telemetry for Mars 2020 Rover Mechanism Operations

The analysis of mechanism telemetry requires a wide variety of tools to quickly and effectively assess spacecraft state, capture long-term trends in system performance, and identify and track anomalous events. Such analysis often requires spacecraft telemetry to first be transformed into derived fields and aggregated statistics before operators can begin their analysis. In past missions, aspects of this process have been automated, but operators were expected to use their own tools and procedures to understand and visualize the data, which led to redundant and inconsistent tools and processes. The Mech Data Tools Python library (MDT) was developed to provide a flexible, unified tool set for operators to extract and analyze mechanism telemetry over the life of the Mars 2020 surface mission. MDT consists of a set of configurable components that implement standard interfaces for ingesting input and producing output. Components can be chained together to form a data processing pipeline. Data are ingested from several sources within the greater Mars 2020 cloud infrastructure and stored in pandas DataFrames, which allows users to leverage the data manipulation capabilities present within the widely-used pandas library. Visualization capabilities are provided through the Plotly library, which generates interactive plots for users to interpret. Following the beginning of Mars 2020 surface operations, usage of MDT has spread to all mechanism-focused subsystems and has demonstrated great utility in analyzing early surface activities. This paper describes MDT’s evolution from heritage mechanism telemetry tools, the critical architecture decisions and challenges faced over MDT’s two years of development, and current applications of MDT in support of mechanism operations.

Wolsieffer, Ben↗