Search NASA⌕ Search

SEARCH · Search NASA

Results for “Data Reasoning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

VISIONARY: Virtual Intelligence System for Optimizing Novel Analytical Research Yields

VISIONARY is an AI system that accelerates energy materials discovery by automatically generating hypotheses about structure-property relationships. It analyzes patterns in materials data, identifies promising correlations, and proposes testable scientific hypotheses without human intervention. By streamlining this reasoning process, VISIONARY helps researchers efficiently identify candidate materials with desired properties, significantly speeding up the materials development pipeline for energy applications. During the project, we developed a standalone application. The application uses a combination of papers provided by the user and data collected from FutureHouse’s dataset to build an understanding of the background that the user wants to explore for the hypothesis.

36 MATERIALS SCIENCE↗

A roadmap toward scaling, reasoning and self-evolving foundation models for nuclear and particle physics

Foundation models have revolutionized artificial intelligence, with Large Language Models demonstrating unprecedented capabilities in multimodal understanding, reasoning and tool use. Nuclear and particle physics stands at a critical juncture where similar transformative potential awaits realization. The field generates exabytes of experimental data, exascale simulations, and decades of theoretical insights — yet these remain largely disconnected from modern Artifical Intelligence (AI) capabilities, with most physics AI applications confined to narrow, task-specific models that suffer from domain shifting when applied to real experimental data. We present a roadmap for FM4NPP (Foundation Model for Nuclear and Particle Physics), systematically scaling from current proof-of-concept models to trillion-parameter architectures capable of autonomous discovery. Our approach advances three critical frontiers: unified data infrastructure integrating detector data, scientific knowledge and computational tools across global facilities; multi-facility foundation models enabling cross-experiment knowledge transfer and accelerated discovery; and agentic AI capabilities for reasoning and autonomous tool use. The resulting self-evolving FM4NPP will transform physics research by converting time-intensive data analysis, theory derivation and computational bottlenecks into rapid AI–human collaborative discovery. This paradigm shift promises to fundamentally accelerate scientific progress in nuclear and particle physics, enabling researchers to focus on high-level insights while AI handles routine analysis and explores vast parameter spaces beyond human capacity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Hybrid Storage Solution

With the rise of artificial intelligence and machine learning, data sets used to train models have become increasingly large. The availability, accessibility and integrity of large data sets has become important to the research conducted at Los Alamos National Laboratory. Ceph is a storage solution suitable for use with critical data because of its distributed nature and ability to keep multiple copies of a file in different locations. The amount of data means that bandwidth, latency, and cost are important factors and the reason most storage solutions are on-premises. However, there are distinct advantages to hosting services in the cloud, namely scalability and ease-of-use. In this paper, we explore the possibility of provisioning a hybrid Ceph cluster that leverages the benefits of both cloud architectures and on-premise performance.

97 MATHEMATICS AND COMPUTING↗

GOLEM: GOld standard for Learning and Evaluation of Motifs

Motifs are distinctive, recurring, widely used idiom-like words or phrases, often originating from folklore, whose meaning is anchored in a narrative and have a significance as communicative devices across a wide range of media, including news, literature, and propaganda. Many motifs concisely imply a large constellation of culturally relevant information, and their broad usage suggests their cognitive importance as touchstones of cultural knowledge. As such, their detection is a step towards culturally aware natural language processing. We present GOLEM (GOld standard for Learning and Evaluation of Motifs) a dataset of English news articles, opinion pieces, and broadcast transcripts annotated for motific information. The dataset identifies 25,737 motif candidates across 34 motif types drawn from three cultural or national groups: Jewish, Irish, and Puerto Rican. The dataset contains 2,024,141 words split into 25,737 text snippets drawn from 8,073 articles. Each motif candidate is labeled according to a scheme which identifies the type of usage (motific, referential, eponymic, or unrelated), resulting in 1,743 actual motific instances in the data. Annotation was performed by individuals identifying as members of each group and achieved a Fleiss’ kappa (?) of > 0.55. In addition to the data, we demonstrate that classification of the candidate type is a challenging task for Large Language Models (LLMs) using a few-shot approach; recent models such as T5, FLAN-T5, GPT-2, and Llama 2 (7B) achieved a performance of 41% accuracy at best, where the majority class accuracy is 41% and the average chance accuracy is 27%. These data will support development of new models and approaches for detecting (and reasoning about) motific information in text.

motif, culture, natural language, artificial intel↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Energy Materials Chemistry Integrating Theory, Experiment and Data Science (Final Report)

The Energy Materials Chemistry Integrating Theory, Experiment and Data Science (EM-CITED) project is a multidisciplinary research effort focused on accelerating discovery of scientific knowledge via incorporation of data science and artificial intelligence in materials chemistry research. The project aims to advance materials chemistry-aware data science to unify theory and experiment knowledge streams. The work resulted in foundational AI frameworks for materials chemistry – Deep Reasoning Networks (DRNets), Hierarchical Correlation Learning for Multi-property Prediction (H-CLMP), and Material-to-Spectrum (Mat2Spec) prediction – as well as a host of strategies for accelerated scientific discoveries through principled incorporation of data science in computational and experimental research.

36 MATERIALS SCIENCE↗

Impact of various DIII-D diagnostics on the accuracy of neural network surrogates for kinetic EFIT reconstructions

Abstract Kinetic equilibrium reconstructions make use of profile information such as particle density and temperature measurements in addition to magnetics data to compute a self-consistent equilibrium. They are used in a multitude of physics-based modeling. This work develops a multi-layer perceptron (MLP) neural network (NN) model as a surrogate for kinetic Equilibrium Fitting (EFITs) and trains on the 2019 DIII-D discharge campaign database of kinetic equilibrium reconstructions. We investigate the impact of including various diagnostic data and machine actuator controls as input into the NN. When giving various categories of data as input into NN models that have been trained using those same categories of data, the predictions on multiple equilibrium reconstruction solutions (poloidal magnetic flux, global scalars, pressure profile, current profile) are highly accurate. When comparing different models with different diagnostics as input, the magnetics-only model outputs accurate kinetic profiles and the inclusion of additional data does not significantly impact the accuracy. When the NN is tasked with inferring only a single target such as the EFIT pressure profile or EFIT current profile, we see a large increase in the accuracy of the prediction of the kinetic profiles as more data is included. These results indicate that certain MLP NN configurations can be reasonably robust to different burning-plasma-relevant diagnostics depending on the accuracy requirements for equilibrium reconstruction tasks.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

Coupled Aerodynamic and Hydrodynamic Hybrid Simulation of Floating Offshore Wind Turbines

The development and innovation of floating offshore wind energy in the U.S. requires detailed high-fidelity observations and measurements of turbine and platform loading due to wind, waves, and currents. However, full-scale and quasi-full-scale experiments require significant financial and temporal investments for construction, experimental testing, and long-term field campaigns. To support the commercial advancement of the offshore wind energy industry, specialized wind tunnel and wave basin experimental facilities are critical to be able to test FOWT designs at small scale under controlled conditions prior to full-scale deployment. Oregon State University (OSU) is internationally known as a leader in water and energy research, development, and testing. The O.H. Hinsdale Wave Research Laboratory (HWRL) and the Wallace Energy Systems and Renewables Facility (WESRF) at OSU have extensive experience building, modeling, monitoring, controlling, and actuating scaled systems. Experiments on wave-structure interaction have been performed at the HWRL since its establishment in 1972. Studies have included the interaction of waves with coastal structures (breakwaters, seawalls, buildings, cylinders, bridges, fixed foundations of offshore wind turbines, etc.) and with floating structures (e.g., wave energy converters, maneuvering of vessels, etc.). Hinsdale is actively used by marine energy technology developers, both for private testing and OSU-collaborative research projects. However, despite the availability of several large-scale facilities for hydrodynamic testing (at OSU and elsewhere in the U.S.), existing experimental laboratories are generally limited in their ability to accurately generate combined wind and wave conditions. The simulation of both wind and waves in experimental testing is complicated due to a number of constraints, including: [i] incompatible similitude laws governing the wind and waves for scaled experiments, [ii] producing accurate wind over a large enough control volume via fans, and [iii] generating wind that reasonably represents the atmospheric boundary layer in existing wave basins/flumes. Hence, physical test data providing insight into the simultaneous wave- and wind-structure response of floating offshore wind components can be difficult to generate. Given the aforementioned challenges in classic hydrodynamic experiments, the motivation of this project is to establish a real-time hybrid simulation (RTHS) approach that can apply aero- and hydro-dynamic loading by augmenting wave-only experimental facilities with virtual aerodynamic forces through numerical models representing the remaining dynamic forces. RTHS is a physical-numerical approach that partitions a prototype system into physical and numerical sub-assemblies that interact with each other through actuators and sensors in real time. In coupling physical and numerical models, the hybrid simulation approach applied herein is ideal for problems with: (1) structures subjected to different scaling laws, such as floating offshore wind turbines subjected to combined aero/hydro-dynamic loading, (2) structures that are too large or complex to be tested entirely in a laboratory setting, such as deep-water mooring applications, and (3) component testing, where the behavior of a portion of the assembly is uncertain but still interacts with other portions of the structure, such as testing the fatigue life of turbine blades. Few U.S. experimental facilities are able to test simultaneous aero- and hydro-dynamic loading and none can accurately produce aero/hydro-dynamic response on scaled FOWT models due to conflicting similitude laws between the wind (commonly Reynolds) and the waves (commonly Froude). To aid in accelerating the development of the U.S. floating offshore industry, there is a significant need to develop a flexible, modular framework that can expand the capacities of existing wave-only laboratories. The project goal is to demonstrate a hydrodynamic real-time hybrid simulation (hydro-RTHS) framework that couples numerical wind and physical waves acting on a FOWT, thus representing simultaneous aero/hydro-dynamic loading. The FOWT is partitioned into a full-scale numerical sub-assembly associated with the aerodynamics and a model-scale physical sub-assembly associated with the hydrodynamics. The numerical-physical partition associated with hydro-RTHS mitigates scaling constraints by supplying different scaling laws to the physical and numerical sub-assemblies. Herein, length, force, and time are scaled and exchanged between the sub-assemblies using Froude scaling to represent the open-channel flow in the physical sub-assembly. Other similitude laws could also be utilized depending on the problem definition. It is envisioned that the ability to model FOWTs under waves and wind, with mitigation of similitude distortions, would result in reduced development costs (currently, FOWT concept development is performed with full-size pro- totypes at enormous expense and risk) and increase the reliability of the FOWT industry (since extreme wave and wind conditions and contingency events can be tested safely in a controlled environment).

16 TIDAL AND WAVE POWER↗

Evaluating High-Halide Waste Form Options for Salt-Based Nuclear Waste Simulants

In this study, waste forms were being explored for an electrochemical salt simulant, , referred to as ERV3, which is a high-LiCl/KCl salt containing simulated fission products (i.e., Sr, Cs, Nd) and Na to represent bond sodium from Experimental Breeder Reactor-II metallic fast reactor fuel. The goal was to find glassy systems that could be used to immobilize the salt in a single-step process. The envisionment of this process would be find a frit glass that could be added to the salt waste, heat treated, poured into waste canisters, and then stored for disposal. To perform this study, a literature review was conducted, the most promising seven systems were fabricated without the salt, and then mixed with salt simulant and heat treated under different processes. The criteria that were used to screen potential compositions included demonstrated alkali incorporation, could be melted at reasonably low temperature ( T ≤ 1000°C), and if the compositions had some demonstrated data for waste-form-related properties, that was a benefit. High marks were given for compositions that showed amorphous nature after slow cooling of samples containing ERV3.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

REFSafE: A RAG-Enabled Framework for Predictive Risk Analysis and Automated Safety Report Generation in Mission-Critical Environments

Operational safety in mission-critical environments requires AI systems that are accurate, interpretable, and resistant to hallucination. We present an agentic Retrieval-Augmented Generation (RAG) framework, REFSafe, for grounded hazard analysis and automated safety report generation. The system integrates Large Language Models (LLMs) with structured operational data, historical incident repositories, policy documents, and external authoritative sources. Through iterative agentic reasoning, the framework retrieves, verifies, and synthesizes evidence prior to generation, enforcing citation-backed outputs with explicit source attribution (documents, links, and prior events) to ensure traceability and trust. To mitigate hallucinations and unsupported claims, all risk assessments and forecasts are constrained to retrieved evidence, with confidence signals derived from retrieval relevance and source consistency. A transparent pipeline enables subject matter experts (SMEs) to validate predictions, and provide structured feedback, forming a continuous performance calibration loop. Preliminary deployment demonstrates improved reliability in hazard detection and safety/vulnerability report generation. This work advances trustworthy, evidence-grounded AI for predictive safety intelligence in mission-critical operations.

Das, Sanjay [ORNL] (ORCID:0009000542591915)↗

Denoising Autoencoder for Reconstructing Sensor Observation Data and Predicting Evapotranspiration: Noisy and Missing Values Repair and Uncertainty Quantification

Abstract Machine learning (ML) methods applied in scientific research often deal with interrelated features in high‐dimensional data. Reducing data noise and redundancy is needed to increase prediction accuracy and efficiency especially when dealing with data from field sensors. We explored an unsupervised learning method, the denoising autoencoder (DAE), to extract the underlying data structure from noisy raw data in the context of predicting hydrologic quantities from multiple field sensors. These sensors have intrinsic instrumental noise and occasional malfunctions that cause missing values. Our DAE neural network reconstructed meteorological sensor data containing noise and missing values to predict evapotranspiration in a mountainous watershed. The DAE reconstructed the sensor variables with a mean coefficient of determination value of 0.77 across 15 dimensions representing individual sensors. It reduced variance and bias uncertainties compared to a classical autoencoder model. The reconstruction quality varied across dimensions depending on their cross‐correlation and alignment with the underlying data structure. Uncertainties arising from the model structure were overall higher than those resulting from data corruption. We attached the DAE structure to a downstream ET‐prediction neural network in three formats and achieved reasonably accurate ET predictions . The use of the DAE notably reduced variance uncertainty in ET prediction. However, excessive variance reduction may be accompanied by an increase in bias due to the intrinsic bias‐variance tradeoff. Our method of evaluating and reducing uncertainties in aggregated data from different sources can be used to improve predictive models, process understanding, and uncertainty quantification for better water resource management. Plain Language Summary We present a machine learning method, namely the denoising autoencoder, which reduces the effects of data noise and missing values typically present in scientific data sets collected through sensor measurements. This method selects the most relevant information from noisy raw data collected by the instruments and fills in missing values. To demonstrate the effectiveness of our method, we applied it to predict evapotranspiration, a hydrologic variable that represents the water moved from the land surface to the atmosphere through a combination of evaporation and plant water use (transpiration). We also used a random sampling technique (the Monte Carlo method) to compare the uncertainty in the predictions when using the raw and noisy data versus the reconstructed data. The denoising process produced more accurate predictions of evapotranspiration with less uncertainty. Improved predictions of evapotranspiration can lead to a better understanding and accounting of water budgets. This ML approach is broadly suitable for a wide variety of applications that involve noisy sensor data with missing values. Key Points We used a denoising autoencoder (DAE) neural network to reduce noise in meteorological and soil sensor observations by on average We used Monte Carlo sampling to estimate the bias and variance of all model outputs, including uncertainty sources from data and the model We attached the DAE component to a downstream neural network to predict ET with the variance reduced by , compared to that without the DAE

denoising autoencoder↗

Comparison of Results between the Legacy and Refined RELAP5-3D Models of the High Temperature Test Facility in Exercises 1 and 2 of the HTTF Benchmark

Work conducted in FY23 identified that RELAP5-3D was capable of reproducing trends in HTTF data during experiment PG-27 but was incapable of reproducing measured values. The primary cause of this discrepancy between RELAP5-3D results and experimental data was hypothesized to be a distortion in power density that was introduced by the radial nodalization of the model. We further hypothesized that a new model would provide better results when compared to the experiments PG-27 and PG-29. Work this FY developed a new model that is better capable of capturing local heat generation rates and contains a representation of each 1/6 azimuthal sector of the core. We used this model to develop a new set of solutions to Exercises 1 and 2 of Problems 2 and 3 in the benchmark. In this report, we present the first comprehensive comparison of the results between the two models. We see that in Exercise 1A, which is common between problems 2 and 3, the results are similar, though the results from the new model show greater detail than those from the legacy model. In Problem 2 Exercise 1B and Problem 3 Exercise 1B, we see that heat removal is slower in the new model than the legacy model. Problem 3 Exercise 1C shows temperatures that are lower in most places in the new model than the legacy model, but the area with active heat generation has higher block temperatures in the new model than the legacy model. Problem 3 Exercise 1D further shows that long-term heat removal is lower in the new model. Problem 2 Exercise 1C demonstrated that the new model observes higher temperatures in the core regions than the legacy model, justifying the need to preserve the power density in HTTF. The validation of PG-27 and PG-29 also demonstrated the improved temperature agreement in the core regions, particularly with a calibrated model that implements an effective thermal conductivity for the core material. Overall, PG-27 models show reasonable to excellent agreement for steady-state temperatures and minimal to reasonable agreement for transients. PG-29 models showed minimal to insufficient agreement with the data.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

An improved dataset for predicting mammal infecting viruses from genetic sequence information

There have been several attempts to develop machine learning (ML) models to identify human infecting viruses from their genomic sequences, with varying degrees of success. Direct comparison between models is problematic, because these models are typically trained and evaluated on different datasets with alternative data splitting schemes, features, and model performance metrics. In this paper we present a standardized dataset of mammal infecting and non-infecting viral pathogens, refined from the previous work of Mollentze et al. to include the latest literature evidence, roughly doubling the number of curated host-virus records available to the community, and new host target labels, primate and mammal. The new host labels were included for several reasons, including previous reports that classification performance is better at broader taxonomic ranks and the idea that there may be more data for primate infection that might serve as a suitable proxy for zoonotic potential and avoidance of false positives for human infection due to absence of evidence. On this dataset, we report the performance of eight machine learning models for predicting mammal-infecting viruses from their genomic sequences. We find that randomly assigning cases in our improved dataset to training/testing sets, when compared to the original assignments into training/testing in Mollentze et al., increases the overall average ROC AUC of prediction of human infection from 0.663 ± 0.070 to 0.784 ± 0.013, consistent with the reduction in phylogenetic distance between train and test sets (relative entropy change from 3.00 to 0.08). The broadest host category of mammal infection can be predicted most reliably at 0.850 ± 0.020. We share our improved dataset and code to enable standardized comparisons of machine learning methods to predict human host infections. Overall, we have presented preliminary evidence that classification of virus host infection is more tractable at higher taxonomic ranks, that unsurprisingly reducing the phylogenetic distance between training and test sets can improve predictive performance, that peptide kmer features appear to be harmful to out of sample model performance, and we are left with the question of whether models for virus host prediction can reasonably be expected to perform well in out of sample scenarios given the likelihood that viruses do not share a common ancestor. Consistent with this concern, when the data is resampled such that there is no overlap between viral families in training and test sets (relative entropy > 24), models perform no better than random chance at prediction of human infection regardless of whether kmers are included (ROC AUC 0.50 ± 0.08) or not (ROC AUC 0.50 ± 0.04).

59 BASIC BIOLOGICAL SCIENCES↗

Expansion-history preferences of DESI DR2 and external data

We explore the origin of the preference of Dark Energy Spectroscopic Instrument (DESI) Data Release 2 (DR2) baryon acoustic oscillation measurements and external data from cosmic microwave background (CMB) and type Ia supernovae (SNIa) that dark energy behavior departs from that expected in the standard cosmological model with vacuum energy (Λ ⁢CDM). In our analysis, we allow a flexible scaling of the expansion rate with redshift that nevertheless allows reasonably tight constraints on the quantities of interest, and adopt and validate a simple yet accurate compression of the CMB data that allows us to constrain our phenomenological model of the expansion history. We find that data consistently show a preference for a 3%–4% increase in the expansion rate at 𝑧 ≃ 0.7 relative to that predicted by the standard Λ⁢ CDM model, in excellent agreement with results from the less flexible (𝑤 0 ,𝑤 𝑎 ) parametrization which was used in previous analyses. Even though our model allows a departure from the best-fit Λ⁢ CDM model at zero redshift, we find no evidence for such a signal. We also find no evidence (at greater than 1⁢𝜎 significance) for a departure of the expansion rate from the Λ ⁢CDM predictions at higher redshifts for any of the data combinations that we consider. Altogether, our results strengthen the robustness of the findings using the combination of DESI, CMB, and SNIa data to dark-energy modeling assumptions.

Cosmological parameters↗

Initial PIP-II Beam Current Monitor Fault Case Analyses & Beam Position Monitor Linearity Studies in CST Studio Suite

The use of non-invasive sensors & systems to measure particle beam characteristics is a crucial part of modern accelerator control systems due to their ability to return real time beam data while minimizing negative effects on beam quality. To ensure that one can be reasonably confident these sensors will behave as desired upon be-ing implemented within the beamline, simulations pre-dicting the performance of these sensors under beamline conditions can be used as a valuable tool for checking sensor functionality without a physical test bench. This paper details the design, testing, and results of two sensor models developed using CST Studio Suite soft-ware designed to mimic two sensors to be implemented within the PIP-III beamline: an elliptical, large-aperture beam position monitor (BPM) for which vertical & hori-zontal position signal linearity was analyzed, and an AC current transformer (ACCT) beam current monitor (BCM) used to search for potential fault cases within the BCM and beam pipe flange gaps. Special focus is given to the discovery of linearity variations within the BPM and the use of frequency domain techniques in the BCM fault case analyses.

Rouzky, A. R.↗

Measurement of Charged-current Muon Neutrino–argon Interactions Without Final-state Pions Using the MicroBooNE Detector

This thesis presents a new high-statistics measurement of flux-integrated single- and double-differential cross sections for charged-current muon neutrino interactions on argon nuclei without final-state pions. The analysis utilizes the full 1.3 $\times$ $10^{21}$ protons-on-target dataset collected by the MicroBooNE liquid argon time projection chamber between 2015 and 2020 at Fermilab's Booster Neutrino Beam. The results of this study are reported with respect to final-state muon kinematic variables and compared with predictions from commonly used neutrino event generators. In one-dimensional distributions, all generators perform reasonably well. However, in two dimensions, only a few demonstrate good agreement with the data. These findings provide valuable insight into characterizing neutrino-nucleon interactions. Such advancements are vital for future long-baseline neutrino experiments, which aim to accurately measure neutrino oscillation and investigate beyond-the-Standard-Model physics with high sensitivity.

Englezos, Panagiotis [Rutgers U., Piscataway (main↗

Evaluating the Effectiveness of Retrieval-Augmented Large Language Models in Scientific Document Reasoning

Despite the dramatic progress in Large Language Model (LLM) development, LLMs often provide seemingly plausible but not factual information, often referred as hallucinations. Retrieval-augmented LLMs provide a non-parametric approach to solve these issues by retrieving relevant information from external data sources and augment the training process. These models helps to trace evidence from an externally provided knowledge base allowing the model predictions to be better interpreted and verified. In this work, we critically evaluate these models in their ability to perform in scientific document reasoning tasks. To this end, we tuned multiple such model variants with science-focused instructions and evaluated them on a scientific document reasoning benchmark for the usefulness of the retrieved document passages. Our findings suggest that models justify predictions in science tasks with fabricated evidence and leveraging scientific corpus as pretraining data does not alleviate the risk of evidence fabrication.

• Artificial intelligence (AI) / machine learning ↗