Search NASA⌕ Search

SEARCH · Search NASA

Results for “data enhancement”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enhancing Data Quality Monitoring at CMS with Interactive Visualization Tools and Automated Reference Run Selection

Current data quality monitoring (DQM) tools at CMS offer granularity limited to per-run analysis. Consequently, issues manifesting at the per-lumisection level can go unnoticed or, even if detectable, often lead to the classification of the whole run as bad, resulting in unnecessary data loss. Additionally, shifters have to evaluate a large set of monitoring elements during their long shifts, increasing the probability of human errors or overlooked problems. In this contribution, we present ongoing work on the development of tools that will provide shifters with an accessible, granularity-enhanced view of DQM data through interactive and dynamic visualizations. Furthermore, we introduce a reference run selection tool currently under development, which will automate the selection based on data-taking conditions and will offer a curated set of training data for machine learning models that will be used for the partial automation of the offline data certification process. These endeavors will be integrated into the DIALS website, enabling enhancements in data certification accuracy and improving the accessibility of DQM at CMS.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Physical Model Enhanced Data Driven Method for High-Resolution Residential Load Profile Generation

Residential buildings account for significant energy consumption, creating opportunities to offer grid services. As electric utilities seek to implement effective system operation strategies, understanding residential energy consumption patterns becomes essential; However, the time intervals of load profiles measured by utilities' smart meters are typically from 15 minutes to 60 minutes. The low-resolution data make it hard to extract appliance-level load information, which is critical for providing grid services. This paper presents a load profile generator designed to produce synthetic load profiles for residential buildings that emphasizes the importance of accurate representations of realistic energy consumption patterns. The generator takes realistic low-resolution residential load measurements and weather data as inputs, producing 1-minute interval profiles that match the characteristics of the original profiles. Further, this generator can be used to populate load profiles in areas where actual measurements are limited to improve the ability of utilities to analyze their distribution systems. By providing more high-resolution residential building load profiles, this tool supports electric utilities to enhance their residential building load control strategies and improve overall grid stability.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Phosphoproteomics Modifications in Women with Rheumatoid Arthritis─Application of Web-Based Software to Enhance Data Visualization

Individuals with rheumatoid arthritis (RA) are at increased risk of functional disability, cardiovascular disease, and obesity, all of which are influenced by dysregulated skeletal muscle. Here, this pilot study aims to identify phosphoproteomics changes in RA skeletal muscle and visualize modifications through development of a web-based app designed to promote user-friendly data interpretation and visualization. NanoLC–MS/MS analysis was performed on vastus lateralis biopsies from three women with RA and matched healthy controls. Differential analysis was performed using the Limma R package. Kinase substrate enrichment analysis (KSEA) predicted changes in kinase activity. RA muscle displayed 35 upregulated and 60 downregulated phosphosites, including the cytoskeletal proteins TTN (Ser33201, Ser33013, Ser20925), NEB (Ser2219, Thr254, Ser33013, Ser20925), FLNA (Ser1459), and LASP1 (Ser146). Compared to healthy controls, KSEA predicted decreased activity of several kinases in RA muscle, including PRKACA and CDKs. All such changes were visualized by use of our web-based app. Overall, phosphoproteome analysis reveals signaling alterations in RA skeletal muscle linked to cytoskeletal proteins, representing candidate disease biomarkers; these modifications can be explored through use of our web-based software.

phosphoproteomics↗

An Autonomous MCP Bridge to Rucio: Enhancing Data Management Accessibility for High Energy Physics

The Rucio Data Management System [1] is an important tool used by High Energy Physics experiments, including those at Fermi National Accelerator Laboratory, to store and manage exabyte-scale scientific datasets. Despite its central role in coordinating data across globally distributed storage sites, Rucio's command line interface (CLI) presents a steep learning curve, and makes it difficult for scientists to navigate through. To solve this issue, a containerized Model Context Protocol (MCP) [2] server was built that connects Large Language Models directly to Rucio, allowing AI agents to handle data tasks by using simple, natural language rather than memorized terminal commands. The core engineering focus of this project was moving the server away from slow terminal commands that require text parsing and replacing them with a native Python Client API toolset and a planned REST API framework. Moving to the Python API handles data operations directly in memory, which helps clear up formatting errors, provides the AI with clean, structured JSON data and speeds up tool execution. To prove that the system actually works, a benchmarking pipeline was also built with various questions to test the AI across four different model configurations. The questions included finding data scopes, tracking down specific datasets, and checking replication rules. Through benchmarking, early runs showed that with raw terminal text, the model would get confused and stuck, whereas switching to the Python API to feed the AI clean, structured data yielded massive improvement. By creating an intelligent and autonomous bridge to a storage network, this project shows how AI can be implemented in scientific data management, which ultimately helps scientists at Fermilab spend less time sorting through data and more time focusing on their experiments and analysis.

Akella, Kashyap [William Rainey Harper Coll.]↗

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence↗

Leveraging Inequality-Constrained Data for Enhanced Liquidus Temperature Prediction in Nuclear Waste Glass Melts

Inequality-constrained data are frequently discarded in engineering, leading to significant information loss in data-scarce domains like glass characterization in nuclear waste vitrification. This paper presents a nonparametric censored-data regression framework based on an l1-norm optimization criterion that leverages slack variables to integrate left-, right-, and interval-constrained observations into training without distributional assumptions. Validated on synthetic data and a Physics-Informed Neural Network (PINN) for predicting liquidus temperature (TL), the method improved R2 from 0.60 to 0.89 and reduced Mean Absolute Error (MAE) by 48% (51.46 to 26.89?rC) on deterministic values. The traditional models failed to satisfy any inequality constraints while the proposed l1-norm PINN satisfies 81.25% of the constraints. The proposed framework effectively extracts actionable information from previously unusable data to enhance predictive accuracy, reduce epistemic uncertainty, and ensure physical consistency in complex industrial applications.

Garcia-Morado, Erick↗

Notes on Real-Beam Ground Mapping with Monopulse Radar

The spatial awareness required of modern flight systems is facilitated by images generated by ground-mapping radar. In the forward and aft directions, synthetic aperture techniques are not viable, leaving us with enhancing real aperture radar data. Enhancing real aperture radar data with monopulse information can be achieved with any of several monopulse beam sharpening techniques. Two such algorithms are discussed.

47 OTHER INSTRUMENTATION↗

Open data sets for assessing photovoltaic system reliability

Photovoltaic (PV) systems have become a cornerstone of renewable energy strategies, particularly due to the significant reduction in solar power costs over the past decade. However, the long-term reliability of PV installations presents a persistent challenge, requiring the development of advanced monitoring and predictive maintenance strategies. A wide range of data types is used to evaluate the health of PV systems, including environmental conditions, electrical performance, and inspection imagery. These data enable methodologies such as machine learning (ML) models for lifetime prediction and computer vision techniques for defect detection. However, the acquisition of high-quality and comprehensive data is difficult, particularly in terms of long-term consistency and data variety. Publicly available data sets serve as valuable resources for addressing these challenges, but they often suffer from fragmentation and are difficult to access. This paper presents a comprehensive review of existing open-source data sets related to PV degradation, analyzing their features, functionalities, and potential applications. We categorize these data sets based on the specific aspects of PV system information they cover, such as environmental conditions, operational monitoring, image inspection and module materials, and propose relevant tools and ML models for processing them. In addition, we propose practices for future data collection and usage, while also discussing potential directions in data-driven research. Our aim is to enhance data utilization and publication among researchers and industry professionals, promoting a deeper understanding of the role of data in enhancing the performance and durability of PV systems.

14 SOLAR ENERGY↗

Persistent global greening over the last four decades using novel long-term vegetation index data with enhanced temporal consistency

Advanced Very High-Resolution Radiometer (AVHRR) satellite observations have provided the longest global daily records from 1980s, but the remaining temporal inconsistency in vegetation index datasets has hindered reliable assessment of vegetation greenness trends. To tackle this, we generated novel global long-term Normalized Difference Vegetation Index (NDVI) and Near-Infrared Reflectance of vegetation (NIRv) datasets derived from AVHRR and Moderate Resolution Imaging Spectroradiometer (MODIS). We addressed residual temporal inconsistency through three-step post processing including cross-sensor calibration among AVHRR sensors, orbital drifting correction for AVHRR sensors, and machine learning-based harmonization between AVHRR and MODIS. After applying each processing step, we confirmed the enhanced temporal consistency in terms of detrended anomaly, trend and interannual variability of NDVI and NIRv at calibration sites. Our refined NDVI and NIRv datasets showed a persistent global greening trend over the last four decades (NDVI: 0.0008 yr -1 ; NIRv: 0.0003 yr -1 ), contrasting with those without the three processing steps that showed rapid greening trends before 2000 (NDVI: 0.0017 yr -1 ; NIRv: 0.0008 yr -1 ) and weakened greening trends after 2000 (NDVI: 0.0004 yr -1 ; NIRv: 0.0001 yr -1 ). These findings highlight the importance of minimizing temporal inconsistency in long-term vegetation index datasets, which can support more reliable trend analysis in global vegetation response to climate changes.

54 ENVIRONMENTAL SCIENCES↗

Data for "Enhancing Lipid Production in Plant Cells through Automated High-Throughput Genome Engineering and Phenotyping"

Plant bioengineering is a time-consuming and labor-intensive process with no guarantee of achieving desired traits. Here, we present a fast, automated, scalable, high-throughput pipeline for plant bioengineering (FAST-PB) in maize (Zea mays) and Nicotiana benthamiana. FAST-PB enables genome editing and product characterization by integrating automated biofoundry engineering of callus and protoplast cells with single-cell matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS). We first demonstrated that FAST-PB could streamline Golden Gate cloning, with the capacity to construct 96 vectors in parallel. Using FAST-PB in protoplasts, we found that PEG2050 increased transfection efficiency by over 45%. For proof-of-concept, we established a reporter-gene-free method for CRISPR editing and phenotyping via mutation of high chlorophyll fluorescence 136. We show that diverse lipids were enhanced up to 6-fold using CRISPR activation of lipid controlling genes. In callus cells, an automated transformation platform was employed to regenerate plants with enhanced lipid traits through introducing multigene cassettes. Lastly, FAST-PB enabled high-throughput single-cell lipid profiling by integrating MALDI-MS with the biofoundry, protoplast, and callus cells, differentiating engineered and unengineered cells using single-cell lipidomics. These innovations massively increase the throughput of synthetic biology, genome editing, and metabolic engineering and change what is possible using single-cell metabolomics in plants.

AI/ML↗

Dynamic CCS-EJ-SJ Database and Web Application - What's New

At the 2024 FECM/NETL Carbon Management Research Project Review Meeting, within the Carbon Transport and Storage Breakout Session 3, the presentation "Dynamic CCS-EJ-SJ Database and Web Application - What's New" highlights the critical tool designed to integrate environmental and social justice considerations into Carbon Capture and Storage (CCS) projects. Key features include an interactive dashboard for data access and visualization, which supports stakeholders in making informed decisions regarding CCS implementation, and updated data layers. The latest version enhances data integration and usability, providing a comprehensive resource for assessing the social and environmental impacts of CCS projects. There are 7 categories in the CCS EJSJ v2 database (released 03/31/2024): environmental justice, energy justice, economic justice, social justice, ecosystem assets, clean energy, and infrastructure. Most of the layers within each category have been updated in this version. As compared to the old database, there are 3 new categories in the v2 database: ecosystem assets, clean energy, and infrastructure.

Sharma, Maneesh↗