Search NASASearch

SEARCH · Search NASA

Results for “Informatics”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

66 records · Page 4

photoD with Rubin ’s Data Preview 1: First stellar photometric distances and faint blue star deficits

Aims. We investigate the utility of Rubin’s Data Preview 1 (DP1) for estimating stellar number density profiles across the Milky Way halo. Methods. We used stellar broad-band near-UV to near-IR ugrizy photometry released in Rubin’s DP1 to estimate distance and metallicity for blue main sequence stars brighter than r = 24 in three ~1.1 sq. deg. fields at southern Galactic latitudes. Results. Compared to TRILEGAL simulations of the Galaxy’s stellar content, we found a likely deficit of blue main sequence turn-off stars with 22 < r < 24. We interpreted this discrepancy as a signature of a steeper halo number density profile at galactocentric distances 10–50 kpc than the canonical ~1/r 3 profile assumed in TRILEGAL simulations. Conclusions. This interpretation is consistent with earlier suggestions based on observations of more luminous, but much less numerous, evolved stellar populations, along with a few pencil beam surveys of blue main sequence stars in the northern sky. These results bode well for the future Galactic halo exploration with Rubin’s Legacy Survey of Space and Time (LSST).

Galaxy: fundamental parameters

pyRMG: A framework for high-throughput, large-cell DFT calculations on supercomputers

Exascale computing delivers the raw power to simulate ever larger and more chemically realistic systems, but realizing this potential requires codes that can efficiently use thousands of processors. Our real-space multigrid (RMG) density functional theory (DFT) code’s grid-decomposition approach scales nearly linearly with the number of graphics processing units (GPUs), even for simulations exceeding thousands of atoms. This scalability makes RMG a compelling tool for high-throughput DFT studies of materials that would otherwise be bottlenecked in other codes (for example, by global fast Fourier transforms in plane-wave DFT). However, the limited workflow infrastructure for RMG has thus far constrained its adoption to a small user community. In this work, we present pyRMG, a Python package designed to streamline the setup and execution of RMG DFT calculations. Built on the pymatgen and ASE (Atomic Simulation Environment) computational materials science Python packages, pyRMG automates input generation and convergence checking, and it integrates with modern job schedulers (e.g., Flux) on leadership-class platforms such as Frontier and Perlmutter. Here, we demonstrate pyRMG for a high-throughput study of strain effects in 2D 2L-Bi 2 Se 3 /2L-NbSe 2 heterostructures, which offers chemical insights into this system and shows that RMG-based workflows can converge with limited user intervention.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

A path to intelligent watersheds: coordinating the data to decision pipeline

Operations of multi-reservoir systems are challenged in-part by the interplay of complex physical processes functioning within the watershed. The employment of intelligent systems can be of aid by linking environmental sensing, information technology, data analytics, simulation and decision support to achieve a data-to-decision flow of information. A further challenge is that watershed resources are managed for multiple purposes requiring some level of coordination among numerous resource managers, asset operators and users. System intelligence in this context relies on shared community platforms (data portals, community models), and coordinated communication between decision makers. Opportunities to enrich watershed intelligence has been the subject of a roadmapping exercise for the Department of Energy’s Water Power Technologies Office which has relied on broad stakeholder engagement. Initial phases of engagement involved personal interviews and a series of virtual group meetings, which focused on identifying opportunities to improve the intelligence of the physical infrastructure within our watersheds—examples of feedback include improved sensing of snowpack and runoff, data standards for facilitated data sharing, and better forecasting tools. The latter phase of engagement involved the conduct of a case study in the Upper Colorado River basin where key stakeholders were interviewed to map how their decisions are informed by intelligence from other basin stakeholders. Our presentation will highlight the interdisciplinary flow of information in complex watershed systems and identify physical and institutional opportunities toward the strategic operation of water infrastructure.

Colorado River

The search for high-entropy fuel-cell catalysts using disorder descriptors

The transition to a hydrogen economy depends on efficient, affordable catalysts for fuel cells. Platinum—the industry standard for fuel-cell electrodes—is costly and scarce, highlighting the need for practical alternatives. High-entropy alloys offer vast compositional diversity and tunable properties that can mitigate these issues, yet their chemical complexity and configurational disorder have hindered rational discovery. Here, we introduce a data-driven framework that couples machine learning with first-principles disorder descriptors—including the entropy forming ability, disordered enthalpy-entropy descriptor, and electronic-structure similarity metrics to platinum—to predict alloy synthesizability and catalytic performance. These descriptors are applied for the first time in the context of fuel-cell catalyst discovery. The workflow rapidly screens more than 20 000 compositions and identifies several platinum-free candidates that are economically viable, readily scalable, and exhibit promising predicted activity. These results demonstrate that disorder descriptors are reliably predicted by machine learning models and can be effectively integrated into materials-discovery pipelines, accelerating innovation across complex compositional spaces.

fuel-cell catalysts

MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration

Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in agriculture, engineering and medicine. Usually, the biosynthesis of these natural products is governed by sets of co-regulated and physically clustered genes known as biosynthetic gene clusters (BGCs). To share information about BGCs in a standardized and machine-readable way, the Minimum Information about a Biosynthetic Gene cluster (MIBiG) data standard and repository was initiated in 2015. Since its conception, MIBiG has been regularly updated to expand data coverage and remain up to date with innovations in natural product research. Here, we describe MIBiG version 4.0, an extensive update to the data repository and the underlying data standard. In a massive community annotation effort, 267 contributors performed 8304 edits, creating 557 new entries and modifying 590 existing entries, resulting in a new total of 3059 curated entries in MIBiG. Particular attention was paid to ensuring high data quality, with automated data validation using a newly developed custom submission portal prototype, paired with a novel peer-reviewing model. MIBiG 4.0 also takes steps towards a rolling release model and a broader involvement of the scientific community. MIBiG 4.0 is accessible online at https://mibig.secondarymetabolites.org/.

59 BASIC BIOLOGICAL SCIENCES

In-orbit performance of the soft X-ray imaging telescope Xtend aboard XRISM

Here, we present a summary of the in-orbit performance of the soft X-ray imaging telescope Xtend onboard the X-Ray Imaging and Spectroscopy Mission (XRISM), based on in-flight observation data, including first-light celestial objects, calibration sources, and results from the cross-calibration campaign with other currently operating X-ray observatories. XRISM/Xtend has a large field of view of ${38{^{\prime }_{.}}5}$ $\times$ ${38{^{\prime }_{.}}5}$, covering an energy range of 0.4–13 keV, as demonstrated by the first-light observation of the galaxy cluster Abell 2319. It also features an energy resolution of 170–180 eV at 6 keV, which meets the mission requirement and enables us to resolve He-like and H-like Fe K$\alpha$ lines. Throughout the observation during the performance verification phase, we confirm that two issues identified in the Soft X-ray Imager (SXI) onboard the previous Hitomi mission—light leakage and crosstalk events—are addressed and suppressed in the case of Xtend. A joint cross-calibration observation of the bright quasar 3C 273 results in an effective area measured to be $\sim$420 cm$^{2}$ at1.5 keV and $\sim$310 cm$^{2}$ at 6.0 keV, which matches values obtained in ground tests. We also continuously monitor the health of Xtend by analyzing overclocking data, calibration source spectra, and day-Earth observations; the readout noise is stable and low, and contamination is negligible even one year after launch. A low background level compared with other major X-ray instruments onboard satellites, combined with the largest grasp ($\Omega _{\rm eff}\sim 60$ cm$^2$ deg$^2$) of Xtend, will not only support Resolve analysis, but also enable significant scientific results on its own. This includes near-future follow-up observations and transient searches in the context of time-domain and multi-messenger astrophysics.

instrumentation: detectors: Xtend

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure

The Arctic

The Arctic environment in 2024 continued on a trajectory that has put it in a state far different from that of the twentieth century. Ongoing accumulation of greenhouse gases in the atmosphere continues to quickly warm the Arctic, resulting in rapid changes in the cryosphere that are driving cascading impacts to climate, ecological, and societal systems. Many weather- and climate-related impacts in the Arctic are the result of compounding change, such as increased riverbank erosion, which is proximately due to increased river discharge from higher seasonal precipitation, yet is also exacerbated by thawing permafrost. However, even individual storms occur within very different ocean and ice conditions than were typically present in the late twentieth century. As a result, the impacts, including high winds, excessive precipitation, and coastal inundation, may be quite different nowadays, as exemplified by the October 2024 storm in northwest Alaska that produced severe coastal flooding in several communities. To share some of these impacts with a wider audience, select extreme weather impacts around the greater Arctic have been highlighted through the inclusion of sidebars in recent State of the Climate Arctic chapters (e.g., Benestad et al. 2023; Thoman et al. 2024).

Thoman, Richard L. [Univ. of Alaska, Fairbanks, AK

Assessing the reliability of medical resource demand models in the context of COVID-19

Abstract Background Numerous medical resource demand models have been created as tools for governments or hospitals, aiming to predict the need for crucial resources like ventilators, hospital beds, personal protective equipment (PPE), and diagnostic kits during crises such as the COVID-19 pandemic. However, the reliability of these demand models remains uncertain. Methods Demand models typically consist of two main components: hospital use epidemiological models that predict hospitalizations or daily admissions, and a demand calculator that translates the outputs of the epidemiological model into predictions for resource usage. We conducted separate analyses to evaluate each of these components. In the first analysis, we validated various hospital use epidemiological models using a recent validation framework designed for epidemiological models. This allowed us to quantify the accuracy of the models in predicting critical aspects such as the date and magnitude of local COVID-19 peaks, among other factors. In the second analysis, we evaluated a range of demand calculators for ventilators, medical gowns, and COVID-19 test kits. To achieve this, we decoupled these demand calculators from the underlying epidemiological models and provided ground truth data for their inputs. This approach enabled a direct comparison of the demand calculators, comparing them against each other and actual usage data when available. The code is available athttps://doi.org/10.5281/zenodo.13712387. Results Performance varied greatly across the epidemiological models, with greater variability in COVID-19 hospital use predictions than for COVID-19 deaths as analyzed previously. Some models did not have any peaks. Among those that did, the models under-estimated date of peak approximately as often as they over-estimated, but were more likely to under-estimate magnitude of peak, with typical relative errors around 50%. Regarding demand calculator predictions, there was significant variability, including five-fold differences in predictions for gown models. Validation against actual or surrogate usage data illustrated the potential value of demand models while demonstrating their limitations. Conclusions The emerging field of demand modeling holds promise in averting medical resource shortages during future public health emergencies. However, achieving this potential necessitates focused efforts on standardization, transparency, and rigorous model validation before placing reliance on demand models in critical public health decision-making.

Medical Informatics

Leveraging Large Language Models for Real-World Data Evidence: A Framework for Automated Treatment Extraction and Data Harmonization

Background: The ability to comprehensively collect treatment information from cancer patient medical records would enable studies to evaluate real-world benefits and risks tied to specific treatments. Currently, it is difficult to system- atically collect high-quality treatment information because it is often stored in unstructured text. Manually extracting and standardizing drug and regimen data is time-intensive. Recent advances in large language models (LLMs) offer a potential solution for automated extraction of structured treatment information from clinical text. Objective: This study systematically evaluates the utility of four LLMs from the Llama family for automated extraction of oncology treatment information from clinical text. This information can guide researchers using cancer registry data to provide insights into cancer care and outcomes beyond clinical trials. Methods: Four instruction-tuned Llama models with varying parameter counts (1B, 3B, 8B, and 70B) were evaluated for their ability to extract treatment information from clinical documents. A unified oncology knowledge base integrating seven major public data sources was developed to standardize and normalize extracted entities—a critical step for harmonizing data from diverse sources. Extracted treatment data were compared against expert-annotated ground truth. Model performance was assessed using accuracy metrics (Precision, Recall, F1-Score) and opera- tional feasibility metrics, including processing speed and structural compliance of the output. Results: A strong positive correlation was observed between model size and extraction accuracy. F1-score improved from 0.609 for the 1B model to 0.710 (3B), 0.807 (8B), and 0.828 (70B). While larger models demonstrated superior accuracy and compliance, they incurred higher computational costs. The modest performance difference between 8B and 70B suggests diminishing returns with increasing model size. Conclusions: LLMs represent a viable technology for automating oncology treatment extraction. The 8B-parameter model emerged as a highly effective option, balancing high accuracy and computational efficiency. Selecting an appropriate LLM for deployment in cancer registries involves a trade-off between desired accuracy and available operational resources. Harmonizing extracted entities with the oncology knowledge base facilitates standardized integration into common data models, enhancing data quality for real-world evidence analyses.

artificial intelligence