Search NASA⌕ Search

SEARCH · Search NASA

Results for “HERS Raters”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Calibration of multisite raters for prospective visual reads of amyloid PET scans

Abstract INTRODUCTION In multicenter Alzheimer's disease studies, amyloid positron emission tomography (PET) visual reads are typically performed centrally by a few experts. Incorporating a broader reader network enhances scalability and generalizability. METHODS Ten neuroimaging experts from eight Alzheimer's Disease Research Centers (ADRCs) visually read 180 amyloid PET scans (30 scans and 15 duplicate scans for each of four tracers, imaged across a wide variety of scanners), using preferred reading software without anatomical imaging or quantitation. Scans were classified as elevated or non‐elevated per tracer‐specific criteria. Inter‐ and intra‐rater agreement was assessed. RESULTS Inter‐rater agreement was substantial (Fleiss’κ = 0.78), with full consensus on 69% of scans. Inter‐rater reliability was substantial to perfect across tracers (Fleiss’κ = 0.70–0.87). Intra‐rater agreement was substantial to perfect (Cohen'sκ = 0.79‐1). Scans with intermediate (10–40 Centiloid) quantitation had lower reader agreement. DISCUSSION A multicenter expert network achieved substantial agreement classifying amyloid PET scans. These scans provide a standard for reader training and reliability assurance in future studies. Highlights Calibration methods ensure reliable amyloid positron emission tomography (PET) visual reads across multiple raters. Substantial agreement is possible across readers using their preferred tools. Agreement is also substantial regardless of the amyloid PET tracer used. Scans with intermediate (10–40 Centiloid) quantitation have lower reader agreement. The calibration set will become a training tool for amyloid PET visual read studies.

Neurosciences & Neurology↗

CRM Assessment: Determining the Generalization of Rater Calibration Training. Summary of Research Report: Gold Standards Training

The extent to which pilot instructors are trained to assess crew resource management (CRM) skills accurately during Line-Oriented Flight Training (LOFT) and Line Operational Evaluation (LOE) scenarios is critical. Pilot instructors must make accurate performance ratings to ensure that proper feedback is provided to flight crews and appropriate decisions are made regarding certification to fly the line. Furthermore, the Federal Aviation Administration's (FAA) Advanced Qualification Program (AQP) requires that instructors be trained explicitly to evaluate both technical and CRM performance (i.e., rater training) and also requires that proficiency and standardization of instructors be verified periodically. To address the critical need for effective pilot instructor training, the American Institutes for Research (AIR) reviewed the relevant research on rater training and, based on "best practices" from this research, developed a new strategy for training pilot instructors to assess crew performance. In addition, we explored new statistical techniques for assessing the effectiveness of pilot instructor training. The results of our research are briefly summarized below. This summary is followed by abstracts of articles and book chapters published under this grant.

Baker, David P.↗

Exploring new frontiers in multi-rater feedback

The presentation provides an overview of the Upward Feedback Program used at JPL to provide feedback to managers on their employees' perceptions of their management effectiveness.

upward feedback leadership employee assessment↗

Final Technical Report

Statement of the problem or situation that is being addressed in your application. The DOE and its national laboratories developed the Home Energy Score™ (HES) to encourage homeowners to improve their energy performance, lower costs and to share energy information through the MLS listing, appraisal, and financing channels. While the HES is an instrumental tool, it is currently underutilized and consists of technical, structural and sector barriers which need to be addressed in order to scale and many energy efficiency contractors are understandably overwhelmed by the added time and effort and lack of incentive to sell and deliver deep retrofit projects while simultaneously meeting the DOE HES program requirements; consequently, contractors may decide to forgo participation. Home Energy Rating System (HERS) Raters have the opportunity to play the critical Assessor role in producing a Home Energy Score (HES); this role has immense potential but currently is unfulfilled. Lastly, while utilities are interested in their customer base achieving greater energy efficiency, especially to help offset growing residential loads in states like California that are accelerating electrification, utilities do not have access to the market actors who are on the front line of influence to homeowners or review and approve their permits: HERS Raters, assessors, contractors and building departments. General statement of how this problem is being addressed: ConSol will integrate the Home Energy Score™ (HES) to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California (CHEERS+HES). The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the utilities in supporting existing homes in their jurisdictions with HES and develop measures to improve energy efficiency and reduce emissions. How is this problem being addressed? What is the overall project approach? In effort to expand the Home Energy Score™ (HES) by increasing the use of aggregable home energy asset data, ConSol proposes to integrate the DOE HES via Application Programming Interface (API) to its State of California approved home energy rating services platform (CHEERS). Once the CHEERS platform and HES are integrated (CHEERS+HES), this enhanced platform will be instantly available and actively deployed via Phase 1 pilot to HERS Raters, assessors and contractors in California to market-test the solution, understand the rate of adoption and identify opportunities for improvement prior to scaling nationally. The CHEERS high fidelity energy code permitting data will be integrated with HES for simple, accurate, easy-to-use home energy estimation and analysis and will directly gain access to the retrofit and renovations markets with the same upgraded platform. This innovative project will assist the building industry and homeowners with an easy-to-use assessment if energy and carbon impacts of existing homes, and assist the utilities in supporting existing homes in their jurisdictions with HES to improve energy efficiency and reduce emissions. What is to be done in Phase I? During Phase I of this proposed project, ConSol will (1) design software architecture that links CHEERS to the Home Energy ScoreTM via API, (2) solicit partnership from one or more California utilities for a regional pilot, (3) test the new software with its HERS Raters and contractor network in the partnership utility jurisdiction, (4) launch a pilot version of the newly developed software with HERS Raters and contractors in the utility territory, and (5) explore California’s GoGreen energy efficiency homeowner lending program in parallel with the pilot. Commercial Applications and Other Benefits. Summarize the future applications or public benefits if the project is carried over into Phase II or Phase III and beyond. The CHEERS+HES commercialized product will be ready for national market scale following a successful Phase 1 performance. The CHEERS+HES adoption is estimated to reach a 5% adoption growth rate versus the 110,000 baseline, starting in Year 1 after Phase I completion, and continuing each year. As a direct benefit to the DOE, CHEERS will set a goal of 100,000 Home Energy Score assessments for existing home alterations within the first 10 years following Phase 1 performance. The technical benefits of this proposed project include the harmonized, automated, and seamless integration of the DOE HES into the widely used and market leading California energy registry, CHEERS. The social benefits include the aggregate energy, cost and GHG savings by allowing the broader public streamlined access to the CHEERS+HES measurement and the energy efficiency recommended measures that may result. Key Words: Home Energy ScoreTM (HES); Application Programming Interface (API); Home Energy Rating Services (HERS); HERS Raters; contractors; assessors; existing homes, energy asset data; cost, energy, and emissions saving measures; energy code (Title 24) compliance; document repository; utilities; pilot; newly developed software; energy efficiency; homeowner. Summary for Members of Congress: The DOE Home Energy Score™ (HES) is a tool to encourage homeowners to improve their energy performance, lower costs and share energy information but is underutilized and consists of barriers which need to be addressed in order to scale. In effort to expand the HES, CHEERS, Inc. will integrate the HES to its State of California, approved home energy rating services (HERS) platform (CHEERS) to develop a single tool for contractors nationwide to assess, record and install recommended cost, energy, and emissions saving measures to the 140 million single-family homes throughout the U.S. and 14 million homes in California.

Application Programming Interface (API)↗

The Role of Spatial Disorientation in Fatal General Aviation Accidents

In-flight Spatial Disorientation (SD) in pilots is a serious threat to aviation safety. Indeed, SD may play a much larger role in aviation accidents than the approximate 6-8% reported by the National Transportation Safety Board (NTSB) each year, because some accidents coded by the NTSB as aircraft control-not maintained (ACNM) may actually result from SD. The purpose of this study is to determine whether SD is underestimated as a cause of fatal general aviation (GA) accidents in the NTSB database. Fatal GA airplane accidents occurring between January 1995 and December 1999 were reviewed from the NTSB aviation accident database. Cases coded as ACNM or SD as the probable cause were selected for review by a panel of aerospace medicine specialists. Using a rating scale, each rater was instructed to determine if SD was the probable cause of the accident. Agreement between the raters and agreement between the raters and the NTSB were evaluated by Kappa statistics. The raters agreed that 11 out of 20 (55%) accidents coded by the NTSB as ACNM were probably caused by SD (p less than 0.05). Agreement between the raters and the NTSB did not reach significance (p greater than 0.05). The 95% C.I. for the sampling population estimated that between 33-77% of cases that the NTSB identified as ACNM could be identified by aerospace medicine experts as SD. Aerospace medicine specialists agreed that some cases coded by the NTSB as ACNM were probably caused by SD. Consequently, a larger number of accidents may be caused by the pilot succumbing to SD than indicated in the NTSB database. This new information should encourage regulating agencies to insure that pilots receive SD recognition training, enabling them to take appropriate corrective actions during flight. This could lead to new training standards, ultimately saving lives among GA airplane pilots.

Scheuring, RIchard↗

Development and Validation of a Scoring System for Abnormalities in the Gopher Frog (Rana capito)

Headstarting efforts are thought to be critical in supplementing populations of the at-risk Gopher Frog (Rana capito); however, recent efforts have occasionally resulted in juveniles with developmental abnormalities. In response, we developed a scoring system to collect quantitative data on the presence and severity of these developmental abnormalities. Our objective was to describe and validate the abnormality scoring system so that it can be used by all Gopher Frog headstarting facilities. The scoring system covers five primary conditions encompassing commonly observed abnormalities. Two groups of participants with different levels of prior experience working with Gopher Frogs assigned scores to a set of images presented to them for each condition. We used intra-class correlation coefficients (ICC) to test the scoring system for inter- and intra-rater agreement as well as agreement with the benchmark standard (established by the authors). We found high ICC values for inter-rater agreement, intra-rater agreement, and agreement to the benchmark standard indicating either excellent or good reliability for all five conditions and for all raters when grouped together. These findings support the reliability and validity of the proposed developmental abnormality scoring system. Gopher Frog headstarting facilities can implement this scoring system to assist in tracking the frequency and severity of abnormalities observed in future headstarting efforts. We hope that by creating a reliable scoring system for Gopher Frogs, it can provide an overall framework and serve as a valuable resource to evaluate abnormalities across any amphibian species.

abnormality↗

Integration of Optical Coherence Tomography Scan Patterns to Augment Clinical Data Suite

Vision changes identified in long duration spaceflight astronauts has led Space Medicine at NASA to adopt a more comprehensive clinical monitoring protocol. Optical Coherence Tomography (OCT) was recently implemented at NASA, including on board the International Space Station in 2013. NASA is collaborating with Heidelberg Engineering to increase the fidelity of the current OCT data set by integrating the traditional circumpapillary OCT image with radial and horizontal block images at the optic nerve head. The retinal nerve fiber layer was segmented by two experienced individuals. Intra-rater (N=4 subjects and 70 images) and inter-rater (N=4 subjects and 221 images) agreement was performed. The results of this analysis and the potential benefits will be presented.

Mason, S.↗

Evaluation of Saccadic Component Measure on Smooth Pursuit Tests

ABSTRACT Introduction Despite the advancement of eye-tracking technology for smooth pursuit (SP) eye movement evaluation, qualitative observation offers much information that is not captured by computers; hence, both objective and qualitative information should be utilized to evaluate SP. This study examined the consistency among our clinicians when evaluating SP using normal (N), grossly normal (GN), mildly abnormal (MA), and abnormal (AB) as classifications. We then evaluated the effect of combining GN and MA into a single subclinical (SUBC) category. We also evaluated the computerized percent saccade (PS) metric by determining its sensitivity and specificity in classifying SP. Materials and Methods Retrospective horizontal and vertical SP test videos and numerical data for 70 participants were obtained from the Neuro Kinetics Neuro-Otologic Test Center and de-identified. From this, eye-tracking videos, time plots of eye-tracking positional data, and tables of SP eye-tracking performance data were generated for 0.1, 0.3, and 0.5 Hz in both horizontal and vertical planes, totaling 6 tests per subject. Three clinicians rated each subject’s SP performance as N, GN, MA, or AB for a total of 6 ratings (3 frequencies, horizontal and vertical). This process was repeated using N, SUBC, and AB as rating categories. Clinicians also provided an overall SP rating for each plane as follows: AB if the results were abnormal for 2 or more frequencies tested. Alternatively, if fewer than 2 frequencies presented with a rating of AB, then an overall rating of MA, GN, or N was determined at the respective clinician’s discretion. Results When the 3 clinicians were tasked with classifying SP videos using 4 clinical categories, fair overall agreement was demonstrated. However, when MA and GN categories were combined into an SUBC category, the overall agreement for the 3 clinicians improved slightly for both horizontal SP (HSP) and vertical SP (VSP). This pattern of agreement did not differ considerably when comparing HSP versus VSP, and good consistency and reliability was observed across clinicians. Again, inter-rater consistency was smaller for VSP versus HSP despite the reduction in clinical categories. Cut-off values were generated for the PS metric and demonstrated good specificity and sensitivity when they were exceeded for 2 or more frequencies in a particular plane when evaluating a subject’s SP test. Conclusions

General & Internal Medicine↗

Specification-based software sizing: An empirical investigation of function metrics

For some time the software industry has espoused the need for improved specification-based software size metrics. This paper reports on a study of nineteen recently developed systems in a variety of application domains. The systems were developed by a single software services corporation using a variety of languages. The study investigated several metric characteristics. It shows that: earlier research into inter-item correlation within the overall function count is partially supported; a priori function counts, in themself, do not explain the majority of the effort variation in software development in the organization studied; documentation quality is critical to accurate function identification; and rater error is substantial in manual function counting. The implication of these findings for organizations using function based metrics are explored.

Jeffery, Ross↗

LOFT Debriefings: An Analysis of Instructor Techniques and Crew Participation

This study analyzes techniques instructors use to facilitate crew analysis and evaluation of their Line-Oriented Flight Training (LOFT) performance. A rating instrument called the Debriefing Assessment Battery (DAB) was developed which enables raters to reliably assess instructor facilitation techniques and characterize crew participation. Thirty-six debriefing sessions conducted at five U.S. airlines were analyzed to determine the nature of instructor facilitation and crew participation. Ratings obtained using the DAB corresponded closely with descriptive measures of instructor and crew performance. The data provide empirical evidence that facilitation can be an effective tool for increasing the depth of crew participation and self-analysis of CRM performance. Instructor facilitation skill varied dramatically, suggesting a need for more concrete hands-on training in facilitation techniques. Crews were responsive but fell short of actively leading their own debriefings. Ways to improve debriefing effectiveness are suggested.

Dismukes, R. Key↗

A Gold Standards Approach to Training Instructors to Evaluate Crew Performance

The Advanced Qualification Program requires that airlines evaluate crew performance in Line Oriented Simulation. For this evaluation to be meaningful, instructors must observe relevant crew behaviors and evaluate those behaviors consistently and accurately against standards established by the airline. The airline industry has largely settled on an approach in which instructors evaluate crew performance on a series of event sets, using standardized grade sheets on which behaviors specific to event set are listed. Typically, new instructors are given a class in which they learn to use the grade sheets and practice evaluating crew performance observed on videotapes. These classes emphasize reliability, providing detailed instruction and practice in scoring so that all instructors within a given class will give similar scores to similar performance. This approach has value but also has important limitations; (1) ratings within one class of new instructors may differ from those of other classes; (2) ratings may not be driven primarily by the specific behaviors on which the company wanted the crews to be scored; and (3) ratings may not be calibrated to company standards for level of performance skill required. In this paper we provide a method to extend the existing method of training instructors to address these three limitations. We call this method the "gold standards" approach because it uses ratings from the company's most experienced instructors as the basis for training rater accuracy. This approach ties the training to the specific behaviors on which the experienced instructors based their ratings.

Baker, David P.↗

Evidence of Participation in the Advancement of Knowledge from 25 Years of Publications for the GLOBE Program

The Global Learning and Observations to Benefit the Environment (GLOBE) Program was announced on Earth Day 1994 and began on Earth Day 1995. Over the first 25 years of its operation, there have been more than 600 formal publications and 340 conference presentations about the program, in at least 15 languages. As the 25th anniversary of the program approached, we began a focused effort to find and characterize those publications in order to better understand GLOBE’s impact to date as well as its future potential. GLOBE was envisioned from the beginning as a science and education program, so GLOBE publications cover a wide range of topics across Earth system science as well as in education and evaluation topics. To characterize the collection, a large set of dimensions was first identified from the literature, across science and education. These include codes related to: motivation and conception of the research investigation, educator practices and training, education outcomes & demonstration that these outcomes were achieved, data and study scope, characteristics of the investigation or its findings. Application of these codes was first tested by three raters (the first three authors) on a subset of the publications, then refined to capture missing aspects. The full set of publications was then rated on the full set of criteria. As a second step, subsets of the collection with particularly in depth discussion of The GLOBE Program are now being analyzed in more detail. This poster will summarize various attributes of the GLOBE bibliography and identify areas of potential for further geoscience investigations.

Lin Chambers↗

Defining Change Thresholds: What Change Is Outside Typical Sources of Variation?

Researchers often have a difficult time defining meaningful thresholds for change. We sometimes identify subtle changes but what amount of change is beyond typical sources of variation? This is especially complicated when trying to understand new disease pathogenesis like the constellation of eye changes leading to Spaceflight-associated Neuro-ocular Syndrome (SANS). To support decision makers in defining minimal meaningful change, we used a Bayesian hierarchical model to estimate innate sources of variability such as natural day to day variation. Healthy subjects were recruited and imaged with MRI, OCT, and US on separate days and measured by several technicians. Models were developed specifying random effects for the sources of variation – between left and right eyes, within-individuals over time, between raters, and finally between individuals. This allowed us to find the posterior distribution for the total typical variation, within an eye, which we use to define a threshold where change beyond typical sources of variation is likely. This threshold is now used as our earliest indicator of systematic increase in Total Retinal Thickness (a precursor to optic disc edema).

Millennia Young↗

Results from a GPS Shuttle Training Aircraft flight test

A series of Global Positioning System (GPS) flight tests were performed on a National Aeronautics and Space Administration's (NASA's) Shuttle Training Aircraft (STA). The objective of the tests was to evaluate the performance of GPS-based navigation during simulated Shuttle approach and landings for possible replacement of the current Shuttle landing navigation aid, the Microwave Scanning Beam Landing System (MSBLS). In particular, varying levels of sensor data integration would be evaluated to determine the minimum amount of integration required to meet the navigation accuracy requirements for a Shuttle landing. Four flight tests consisting of 8 to 9 simulation runs per flight test were performed at White Sands Space Harbor in April 1991. Three different GPS receivers were tested. The STA inertial navigation, tactical air navigation, and MSBLS sensor data were also recorded during each run. C-band radar aided laser trackers were utilized to provide the STA 'truth' trajectory.

Saunders, Penny E.↗