Search NASA⌕ Search

SEARCH · Search NASA

Results for “task tuning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Graph Embeddings for CEBAF Operations: Progress and Future Plans

We describe research towards leveraging deep learning on graph representations of the injector beamline at the Continuous Electron Beam Accelerator Facility (CEBAF) in order to create a tool for improving the efficiency of beam tuning tasks. Specifically, we use graphs to represent the injector beamline at any arbitrary date and time and invoke a graph neural network to extract a low-dimensional, informative representation that can be visualized in two-dimensions. By analyzing years of operational data from the CEBAF archiver, good and bad regions of parameter space can be identified. The goal is to exercise this framework as a real-time tool to aid beam tuning, which represents the dominant source of machine downtime.

Tennant, C.↗

Graph Analytics for CEBAF Operations

We report on the progress achieved during a 2-year Laboratory Directed Research and Development (LDRD) project titled “Graph Analytics for CEBAF Operations”. The objective of this project is to leverage deep learning on graph representations of CEBAF’s injector beamline in order to create a tool for improving the efficiency of beam tuning tasks. Specifically, we use graphs to represent the injector beamline at any arbitrary date and time and invoke a graph neural network (GNN) to extract a low-dimensional, informative representation that can be visualized in two-dimensions. By analyzing years of operational data from the CEBAF archiver, good and bad regions of parameter space can be identified. The goal is to exercise this framework as a real-time tool to aid beam tuning, which represents the dominant source of machine downtime.

43 PARTICLE ACCELERATORS↗

MATEY: multiscale adaptive transformer models for spatiotemporal physical systems

Accurate representation of the multiscale features in spatiotemporal physical systems using vision transformer architectures requires extremely long, computationally prohibitive token sequences. To address this issue, we propose two novel adaptive tokenization schemes that dynamically adjust patch sizes based on local features: one ensures convergent behavior to uniform patch refinement, while the other offers better computational efficiency. Moreover, we present a set of spatiotemporal attention schemes, where the temporal or axial spatial dimensions are decoupled, to evaluate their baseline computational and data efficiencies and to determine whether adaptive tokenization can improve this performance. We assess the performance of the proposed multiscale adaptive model, MATEY, in a sequence of experiments. Compared to a full spatiotemporal attention scheme or a scheme that decouples only the temporal dimension, we find that fully decoupled axial attention is less efficient and expressive, requiring more training time and model parameters to achieve the same accuracy. The experiments on the adaptive tokenization schemes show that, compared to a uniformly refined model, the proposed schemes achieve comparable or improved accuracy at a much lower cost in the tested two-dimensional settings. While the asymptotic analysis suggests the potential for favorable scaling, empirical validation at substantially longer sequence lengths remains to be performed in future work. Finally, we demonstrate in two fine-tuning tasks featuring different physics that models pretrained on PDEBench data outperform the ones trained from scratch, especially in the low data regime with frozen attention.

adaptive tokenization↗

Deep transfer operator learning for partial differential equations under conditional shift

Transfer learning enables the transfer of knowledge gained while learning to perform one task (source) to a related but different task (target), hence addressing the expense of data acquisition and labelling, potential computational power limitations and dataset distribution mismatches. Here, we propose a new transfer learning framework for task-specific learning (functional regression in partial differential equations) under conditional shift based on the deep operator network (DeepONet). Task-specific operator learning is accomplished by fine-tuning task-specific layers of the target DeepONet using a hybrid loss function that allows for the matching of individual target samples while also preserving the global properties of the conditional distribution of the target data. Inspired by conditional embedding operator theory, we minimize the statistical distance between labelled target data and the surrogate prediction on unlabelled target data by embedding conditional distributions onto a reproducing kernel Hilbert space. We demonstrate the advantages of our approach for various transfer learning scenarios involving nonlinear partial differential equations under diverse conditions due to shifts in the geometric domain and model dynamics. Our transfer learning framework enables fast and efficient learning of heterogeneous tasks despite considerable differences between the source and target domains.

42 ENGINEERING↗

Integrated Controls Package for High Performance Interior Retrofit [Slides]

For the last few years, networked lighting control (NLCs) have promised significant energy savings beyond what is achieved through a basic light-emitting diode (LED) lighting retrofit. At the same time, NLCs can substantially increase the cost and complexity of the lighting retrofit. And as lighting system wattage declines because of the increasing efficiency of LEDs, advanced controls have less lighting energy to save and the cost-effectiveness of the NLC investment decreases. But NLCs can be leveraged to achieve significant energy savings and value by enhancing control of other building systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Optimization and stabilization of Fermilab Booster using hybrid Bayesian/RL framework

PIPII project will raise Fermilab Booster intensity and ramp rate. Beam losses will limit average power and are hard to simulate. Presently, Booster uses operator-guided empirical tuning. This task is challenging due to high dimensionality, multiple objectives, critical safety constraints, and drifts. We developed a synergistic suite of Bayesian optimization (BO) and reinforcement learning (RL) tools to optimize and stabilize beam losses. First, active learning was used to build a rough model. Data was collected parasitically using two novel safety constraint types – nonlinear input space restrictions (based on optics model), and uncertainty constraints (to stop bad steps/beam aborts). We then applied online multi-objective BO with scalarized objectives and fitting to improve/rebalance losses, increasing safety margins by 25%. Using BO model as a safety veto, we tried several on/off-policy RL agents for long term stabilization; SAC had best performance. We found that adding contextual (state) information further improved performance, eventually integrating key knobs like linac phase and temperature into the parameter space. Long term testing is ongoing to enable operational use.

Kuklev, Nikita [Fermilab]↗

Defects go green: using defects in nanomaterials for renewable energy and environmental sustainability

Induction of point defects in nanomaterials can bestow upon them entirely new physics or augment their pre-existing physical properties, thereby expanding their potential use in green energy technology. Predicting structure-property relationships for defects a priori is challenging, and developing methods for precise control of defect type, density, or structural distribution during synthesis is an even more formidable task. Hence, tuning the defect structure to tailor nanomaterials for enhanced device performance remains an underutilized tool in materials design. We review here the state of nanomaterial design through the lens of computational prediction of defect properties for green energy technology, and synthesis methods to control defect formation for optimal performance. We illustrate the efficacy of defect-focused approaches for refining nanomaterial physics by describing several specific applications where these techniques hold potential. Most notably, we focus on quantum dots for reabsorption-free solar windows and net-zero emission buildings, oxide cathodes for high energy density lithium-ion batteries and electric vehicles, and transition metal dichalcogenides for electrocatalytic green hydrogen production and carbon-free fuels.

14 SOLAR ENERGY↗

Transformer quantum state: A multipurpose model for quantum many-body problems

Here, inspired by the advancements in large language models based on transformers, we introduce the transformer quantum state (TQS): a versatile machine learning model for quantum many-body problems. In sharp contrast to Hamiltonian/task specific models, TQS can generate the entire phase diagram, predict field strengths with experimental measurements, and transfer such a knowledge to new systems it has never been trained on before, all within a single model. With specific tasks, fine-tuning the TQS produces accurate results with small computational cost. Versatile by design, TQS can be easily adapted to new tasks, thereby pointing towards a general-purpose model for various challenging quantum problems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

SetBERT: the deep learning platform for contextualized embeddings and explainable predictions from high-throughput sequencing

MOTIVATION: High-throughput sequencing (HTS) is a modern sequencing technology used to profile microbiomes by sequencing thousands of short genomic fragments from the microorganisms within a given sample. This technology presents a unique opportunity for artificial intelligence to comprehend the underlying functional relationships of microbial communities. However, due to the unstructured nature of HTS data, nearly all computational models are limited to processing DNA sequences individually. This limitation causes them to miss out on key interactions between microorganisms, significantly hindering our understanding of how these interactions influence the microbial communities as a whole. Furthermore, most computational methods rely on post-processing of samples which could inadvertently introduce unintentional protocol-specific bias. RESULTS: Addressing these concerns, we present SetBERT, a robust pre-training methodology for creating generalized deep learning models for processing HTS data to produce contextualized embeddings and be fine-tuned for downstream tasks with explainable predictions. By leveraging sequence interactions, we show that SetBERT significantly outperforms other models in taxonomic classification with genus-level classification accuracy of 95%. Furthermore, we demonstrate that SetBERT is able to accurately explain its predictions autonomously by confirming the biological-relevance of taxa identified by the model. AVAILABILITY AND IMPLEMENTATION: All source code is available at https://github.com/DLii-Research/setbert. SetBERT may be used through the q2-deepdna QIIME 2 plugin whose source code is available at https://github.com/DLii-Research/q2-deepdna.

Ludwig, David W↗

Bayesian optimization of the beam injection process into a storage ring

We have evaluated the data-efficient Bayesian optimization method for the specific task of injection tuning in a circular accelerator. In this paper, we describe the implementation of this method at the Karlsruhe Research Accelerator with up to nine tuning parameters, including the determination of the associated hyperparameters. We show that the Bayesian optimization method outperforms manual tuning and the commonly used Nelder-Mead optimization algorithm in both simulation and experiment. The algorithm was also successfully used to ease the commissioning phase after the installation of new injection magnets and is regularly used during accelerator operations. We demonstrate that the introduction of context variables that include intrabunch scattering effects, such as the Touschek effect, further improves the control and robustness of the injection process.

43 PARTICLE ACCELERATORS↗

Foundation Models of Scientific Knowledge for Chemistry: Opportunities, Challenges and Lessons Learned

Foundation models pre-trained on large corpora demonstrate significant gains across many natural language processing tasks and domains e.g., law, healthcare, education, etc. However, only limited efforts have investigated the opportunities and limitations of applying these powerful models to science and security applications. In this work we develop foundation models of scientific knowledge for chemistry to augment scientists with the advanced ability to perceive and reason at scale previously unimagined. Specifically, we build large-scale (1.47B parameter) general-purpose models for chemistry that can be effectively used to perform a wide range of in-domain and out-of-domain tasks. Evaluating these models in a zero-shot setting, we analyze the effect of model and data scaling, knowledge depth, and temporality on model performance in context of model training efficiency. Our novel findings demonstrate that (1) model size significantly contributes to the task performance when evaluated in a zero-shot setting; (2) data quality (aka diversity) affects model performance more than data quantity; (3) similarly, unlike previous work (Luu et al., 2021) temporal order of the documents in the corpus boosts model performance only for specific tasks, e.g., SciQ; and (4) models pre-trained from scratch perform better on in-domain tasks than those tuned from general-purpose models like Open AI’s GPT-2.

Foundation Models, Chemistry↗

Towards an astronomical foundation model for stars with a transformer-based model

ABSTRACT Rapid strides are currently being made in the field of artificial intelligence using transformer-based models like Large Language Models (LLMs). The potential of these methods for creating a single, large, versatile model in astronomy has not yet been explored. In this work, we propose a framework for data-driven astronomy that uses the same core techniques and architecture as used by LLMs. Using a variety of observations and labels of stars as an example, we build a transformer-based model and train it in a self-supervised manner with cross-survey data sets to perform a variety of inference tasks. In particular, we demonstrate that a single model can perform both discriminative and generative tasks even if the model was not trained or fine-tuned to do any specific task. For example, on the discriminative task of deriving stellar parameters from Gaia XP spectra, we achieve an accuracy of 47 K in Teff, 0.11 dex in log g, and 0.07 dex in [M/H], outperforming an expert XGBoost model in the same setting. But the same model can also generate XP spectra from stellar parameters, inpaint unobserved spectral regions, extract empirical stellar loci, and even determine the interstellar extinction curve. Our framework demonstrates that building and training a single foundation model without fine-tuning using data and parameters from multiple surveys to predict unmeasured observations and parameters is well within reach. Such ‘Large Astronomy Models’ trained on large quantities of observational data will play a large role in the analysis of current and future large surveys.

Leung, Henry W. (ORCID:0000000200362752)↗

HCPerf: Driving Performance-Directed Hierarchical Coordination for Autonomous Vehicles

The rapid development of autonomous driving poses new research challenges to the on-vehicle computing system. In particular, the execution time of autonomous driving tasks highly depends on the specific driving environment. For instance, the execution time of configurable sensor fusion increases significantly as the scene becomes complex, which leads to end-to-end deadline misses from sensing to control and may cause accidents. Thus, a framework that can effectively utilize the system resources to guarantee the end-to-end deadlines of autonomous driving tasks as well as effectively prioritize the responsiveness and throughput of the control commands is crucial for autonomous driving. In this paper, we propose HCPerf, a performance-directed hierarchical coordination framework that intelligently coordinates the autonomous driving tasks with high execution time variation and complex dependencies according to the driving performance in real-time. Specifically, HCPerf mainly consists of two coordinators. The internal coordinator intelligently schedules the tasks according to the driving performance of the vehicle in order to help them meet the end-to-end deadlines while well prioritizing the responsiveness and throughput of the control commands. At the same time, the external coordinator dynamically tunes the rates of tasks according to the schedulability in order to efficiently utilize the system resource. We conduct extensive experiments on both simulation and hardware testbeds with the representative autonomous driving application. The results show that HCPerf can effectively improve the driving performance by 7.69%-45.94% in different driving scenarios.

Ma, Jialiang↗

Customized Bayesian optimization for efficient beam tuning at the facility for rare isotope beams

Bayesian optimization (BO) has recently emerged as a powerful approach for on-line beam tuning, and it is rapidly gaining adoption across accelerator facilities due to its flexibility and efficiency in handling complex optimization tasks. At the Facility for Rare Isotope Beams, rapid and reliable tuning is essential to support the delivery of diverse ion species. To improve the practicality of BO in this setting, we implemented several enhancements, including scalarized composite objective construction for multicriteria optimization, asynchronous evaluation for better resource utilization, prior-mean-assisted optimization to accelerate convergence, and a local search strategy for rapid completion of the task. We present the details of these methods, discuss challenges-encountered, and share our experience applying them to specific beam-tuning tasks.

Accelerators & storage rings↗

Covariance-Free Bifidelity Control Variates Importance Sampling for Rare Event Reliability Analysis

Multifidelity modeling has been steadily gaining attention as a tool to address the problem of exorbitant model evaluation costs that makes the estimation of failure probabilities a significant computational challenge for complex real-world problems, particularly when failure is a rare event. To implement multifidelity modeling, estimators that efficiently combine information from multiple models/sources are necessary. In past works, the variance reduction techniques of control variates (CV) and importance sampling (IS) have been leveraged for this task. In this paper, we present the CVIS framework—a creative take on a coupled CV and IS estimator for bifidelity reliability analysis. The framework addresses some of the practical challenges of the CV method by using an estimator for the control variate mean and sidestepping the need to estimate the covariance between the original estimator and the control variate through a clever choice for the tuning constant. Furthermore, the task of selecting an efficient IS distribution is also considered, with a view towards maximally leveraging the bifidelity structure and maintaining expressivity. Additionally, a diagnostic is provided that indicates both the efficiency of the algorithm as well as the relative predictive quality of the models utilized. Finally, the behavior and performance of the framework is explored through analytical and numerical examples.

Markov chain Monte Carlo↗

KEBLM: Knowledge-Enhanced Biomedical Language Models

Pretrained language models (PLMs) have demonstrated strong performance on many natural language processing (NLP) tasks. Despite their great success, these PLMs are typically pretrained only on unstructured free texts without leveraging existing structured knowledge bases that are readily available for many domains, especially scientific domains. As a result, these PLMs may not achieve satisfactory performance on knowledge-intensive tasks such as biomedical NLP. Comprehending a complex biomedical document without domain-specific knowledge is challenging, even for humans. Inspired by this observation, we propose a general framework for incorporating various types of domain knowledge from multiple sources into biomedical PLMs. We encode domain knowledge using lightweight adapter modules, bottleneck feed-forward networks that are inserted into different locations of a backbone PLM. For each knowledge source of interest, we pretrain an adapter module to capture the knowledge in a self-supervised way. We design a wide range of self-supervised objectives to accommodate diverse types of knowledge, ranging from entity relations to description sentences. Once a set of pretrained adapters is available, we employ fusion layers to combine the knowledge encoded within these adapters for downstream tasks. Each fusion layer is a parameterized mixer of the available trained adapters that can identify and activate the most useful adapters for a given input. Our method diverges from prior work by including a knowledge consolidation phase, during which we teach the fusion layers to effectively combine knowledge from both the original PLM and newly-acquired external knowledge using a large collection of unannotated texts. After the consolidation phase, the complete knowledge-enhanced model can be fine-tuned for any downstream task of interest to achieve optimal performance. Extensive experiments on many biomedical NLP datasets show that our proposed framework consistently improves the performance of the underlying PLMs on various downstream tasks such as natural language inference, question answering, and entity linking. These results demonstrate the benefits of using multiple sources of external knowledge to enhance PLMs and the effectiveness of the framework for incorporating knowledge into PLMs. Finally, while primarily focused on the biomedical domain in this work, our framework is highly adaptable and can be easily applied to other domains, such as the bioenergy sector.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

LLaMP v0.1.0

Reducing hallucination of Large Language Models (LLMs) is imperative for use in the sciences, where reliability and reproducibility are crucial. However, LLMs inherently lack long-term memory, making it a nontrivial, ad hoc, and inevitably biased task to fine-tune them on domain-specific literature and data. LLaMP is a multimodal retrieval-augmented generation (RAG) framework of hierarchical reasoning and acting (ReAct) agents that can dynamically and recursively interact with Materials Project to ground large language models on high-fidelity materials informatics.

Riebesell, Janosh [Lawrence Berkeley National Labo↗

ChatPORT: Fine-Tuned LLM for Easy Code {PORT}ing

Fine-tuning existing LLMs for specialized tasks has become a very attractive alternative due to its low cost and quick development cycle. With many pre-trained LLMs available, it is an increasingly complex task to choose the correct model as the starting point or base model. In this work we discuss ChatPORT - a specialized fine-tuned LLM geared towards providing correctly translated codes from one programming model to another. We evaluate a number of base models and compare and contrast their features and characteristics that make them a viable starting point. In this paper, we focus on the OpenMP offload porting capabilities of ChatPORT. We build our training data using kernels from the Heterogeneous Computing Benchmarks (HeCBench) [12] and the OpenMP Validation and Verification suite [5] to fine-tune the base models. We then test the model using unseen kernels extracted from the HeCBench benchmark suite. Our results show that: (1) not all open LLMs geared towards HPC are aware of programming models like OpenMP, (2) although all base models benefit from fine-tuning they learn differently and produce different correctness rates, (3) depending on the memory size and compute resource available, different base models can be used for fine-tuning without significantly affecting the quality of transpiled code they generate, (4) fine-tuning improved the correctness rate of the LLM by an average of 43.2%, and (5) feedback-based training data further increased the correctness rate by an average of 6% over the LLMs tested.

Pophale, Swaroop [ORNL] (ORCID:0000000185446367)↗