Search NASA⌕ Search

SEARCH · Search NASA

Results for “validation dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A Hyperspectral Inversion Framework for Estimating Absorbing Inherent Optical Properties and Biogeochemical Parameters in Inland and Coastal Waters

The simultaneous remote estimation of biogeochemical parameters (BPs) and inherent optical properties (IOPs) from hyperspectral satellite imagery of globally distributed optically distinct inland and coastal waters is a complex, unsolved, non-unique inverse problem. To tackle this problem, we leverage a machine-learning model termed Mixture Density Networks (MDNs). MDNs outperform operational algorithms by calculating the covariance between the simultaneously estimated products. We train the MDNs on a large ( N = 8237) dataset of co-aligned, in situ measured, hyperspectral remote sensing reflectance (R rs ), BPs, and absorbing IOPs from globally representative optically distinct inland and coastal waters. The estimated IOPs include absorption due to phytoplankton (a ph ), chromophoric dissolved organic matter (a cdom ), and non-algal particles (a nap ). The estimated BPs include chlorophyll-a, total suspended solids, and phycocyanin (PC). MDNs dramatically reduce uncertainty in the retrievals, relative to operational algorithms, when using a 50/50 dataset split, where the MDNs are trained on a randomly selected half of the in situ dataset and validated on the other half. Our model is shown to have higher, or equivalent, generalization performance than the calculated operational algorithms available for all BPs and IOPs (except PC) via a leave-one-out cross-validation assessment. The MDNs are sensitive to uncertainties in the hyperspectral satellite R rs , resulting from instrument noise and atmospheric correction; there is a difference of ~37.4–62.8% (using median symmetric accuracy) between the MDNs’ estimates derived from co-located satellite-derived R rs and in situ R rs . Of the IOPs, a cdom and a nap are less sensitive to uncertainties in hyperspectral satellite imagery relative to a ph , with remote estimates of a ph exhibiting incorrect spectral shape and magnitude relative to in situ measured IOPs. Despite the uncertainties in satellite derived R rs , the spatial distributions of BPs and IOPs in MDN-derived product maps of Lake Erie and the Curonian Lagoon, based on imagery taken with the Hyperspectral Imager for the Coastal Ocean (HICO) and PRecursore Iper-Spettrale della Missione Applicativa (PRISMA), are confirmed via co-aligned in situ measurements and agree with the literature’s understanding of these well-studied regions. The consistency and accuracy of the model on HICO and PRISMA imagery, despite radiometric uncertainties, demonstrate its applicability to future hyperspectral missions, such as the Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission, where the simultaneous estimation model will serve as a key part of phytoplankton community composition analysis.

Ryan E. O'Shea↗

Do Better Satellite Precipitation Algorithms Improve Landslide Hazard Assessment?

Satellites make it possible to estimate precipitation in near real time. Given the challenges of achieving global coverage by other means, these data are used widely. However, few systems for landslide hazard assessment rely on satellite precipitation estimates. This could be due in part to perceptions of accuracy, although latency, spatial resolution, and other factors may also be important. We test whether recent changes to data streams from the Global Precipitation Measurement mission (GPM) have improved its potential for use in landslide prediction. Specifically, we examine data produced by the Integrated Multi-satellitERetrievals for the GPM (IMERG) algorithm, which was upgraded to version 7 this year. IMERG relies upon other algorithms, including the Goddard Profiling Algorithm (GPROF) and the GPM Combined Radar-Radiometer Algorithm (CORRA). Many changes have been made during the switch from IMERG version 6 to version 7. These include upgrading CORRA and GPROF to version 7, to improve the accuracy of precipitation in frozen, mountainous, and coastal areas. The measured intensity of some storms has been enhanced with a new algorithm, the Scheme for Histogram Adjustment with Ranked Precipitation Estimates in the Neighborhood. Combined with many others, these changes to IMERG should improve its utility for landslide hazard assessment in a variety of contexts. To test this idea, we retrain the global Landslide Hazard Assessment for Situational Awareness (LHASA) model twice—first with data from IMERG version 6B and second with 7B. Since current daily rainfall is the most important variable in determining outcomes predicted by LHASA, it should reflect changes made to that input. First, we grid the landslides at a daily, thirty-arcsecond resolution. This serves as the response variable. At each of these sites current and antecedent rainfall are extracted, along with antecedent snow mass and soil moisture, slope, and PGA. In addition, one million grid cells are selected at random points to represent conditions under which landslides (probably) do not occur. After merging these data, we hold back 20% of the dataset for validation purposes and train a machine-learning model with the rest. We assess both the model’s overall ability to identify landslides and its ability to predict specific large landslide disasters.

Thomas A Stanley↗

Deducing Land-Atmosphere Coupling Regimes from SMAP Soil Moisture

In recent years, there has been a growing recognition of the significance of Land-Atmosphere (L-A) interactions and feedback mechanisms in understanding and predicting Earth’s water and energy cycles. Soil moisture plays a critical role in mediating the strength of L-A interactions and is important for understanding the complex and governing processes across this interface. This study aims to identify the significance of soil moisture in identifying L-A coupling strength within the Convective Triggering Potential (CTP) and Humidity Index (HI) framework. To address this, a consistent and reliable dataset of atmospheric profiles is created by merging CTP and HI using Triple Collocation (TC) with three reanalysis datasets. The merged CTP and HI product demonstrates enhanced performance globally as compared to the individual datasets when validated with radiosonde and satellite observations. This merged product of CTP and HI is then used to compare the L-A coupling strength based on Soil Moisture Active Passive Level 3 (SMAPL3) and SMAP Level 4 (SMAPL4) over two decades (2003-2022) where L-A coupling strength is defined as the persistence probability within the dry and wet coupling regimes. Results indicate that the persistency-based coupling strength is related to the ability of soil moisture to predict future atmospheric humidity and dry vs. wet coupling state. The coupling strength in SMAPL4 is consistently stronger than in SMAPL3 and is likely due to its reliance on a land surface model and reduced susceptibility to random noise. The difference in coupling strength based on the same CTP-HI underscores the importance of soil moisture data in estimating coupling strength within the CTP-HI framework. These findings lay the groundwork for understanding the role of L-A interactions and drought evolution due to soil moisture variations, by providing insight into the quantification of coupling strength and its role in drought monitoring and forecast efforts.

Land-atmosphere coupling↗

Dataset for ASME VVUQ Symposium Workshop on Regression of Validation Data to an Application Point

This dataset consists of a collection of Excel spreadsheets that contain output from analysis specified in the workshop. The analysis involves ASME V&V 20-style validation as well as the application of a supplement methodology for regression of validation comparison error and validation uncertainty to application points where experimental data does not exist for comparison. The simulation results and experimental data are provided by the workshop organizers and a NASA report, respectively.

Kirsch, Jared Roelof [Sandia National Laboratories↗

Spectrally Simplified Approach for Leveraging Legacy Geostationary Oceanic Observations

The use of multispectral geostationary satellites to study aquatic ecosystems improves the temporal frequency of observations and mitigates cloud obstruction, but no operational capability presently exists for the coastal and inland waters of the United States. The Advanced Baseline Imager (ABI) on the current iteration of the Geostationary Operational Environmental Satellites, termed the R Series (GOES-R), however, provides sub-hourly imagery and the opportunity to overcome this deficit and to leverage a large repository of existing GOES-R aquatic observations. The fulfillment of this opportunity is assessed herein using a spectrally simplified, two-channel aquatic algorithm consistent with ABI wave bands to estimate the diffuse attenuation coefficient for photosynthetically available radiation, K(d)(PAR). First, an in situ ABI dataset was synthesized using a globally representative dataset of above- and in-water radiometric data products. Values of K(d)(PAR) were estimated by fitting the ratio of the shortest and longest visible wave bands from the in situ ABI dataset to coincident, in situ K(d)(PAR) data products. The algorithm was evaluated based on an iterative cross-validation analysis in which 80% of the dataset was randomly partitioned for fitting and the remaining 20% was used for validation. The iteration producing the median coefficient of determination (R2) value (0.88) resulted in a root mean square difference of 0.319 m−1, or 8.5% of the range in the validation dataset. Second, coincident mid-day images of central and southern California from ABI and from the Moderate Resolution Imaging Spectroradiometer (MODIS) were compared using Google Earth Engine (GEE). GEE default ABI reflectance values were adjusted based on a near infrared signal. Matchups between the ABI and MODIS imagery indicated similar spatial variability (R2 = 0.60) between ABI adjusted blue-to-red reflectance ratio values and MODIS default diffuse attenuation coefficient for spectral downward irradiance at 490 nm, K(d)(490), values. This work demonstrates that if an operational capability to provide- ABI aquatic data products was realized, the spectral configuration of ABI would potentially support a sub-hourly, visible aquatic data product that is applicable to water-mass tracing and physical oceanography research.

Advanced Baseline Imager↗

The Salinity Pilot-Mission Exploitation Platform (Pi-MEP): A Hub for Validation and Exploitation of Satellite Sea Surface Salinity Data

The Pilot-Mission Exploitation Platform (Pi-MEP) for salinity is an ESA initiative originally meant to support and widen the uptake of Soil Moisture and Ocean Salinity (SMOS) mission data over the ocean. Starting in 2017, the project aims at setting up a computational web-based platform focusing on satellite sea surface salinity data, supporting studies on enhanced validation and scientific process over the ocean. It has been designed in close collaboration with a dedicated science advisory group in order to achieve three main objectives: gathering all the data required to exploit satellite sea surface salinity data, systematically producing a wide range of metrics for comparing and monitoring sea surface salinity products’ quality, and providing user-friendly tools to explore, visualize and exploit both the collected products and the results of the automated analyses. The Salinity Pi-MEP is becoming a reference hub for the validation of satellite sea surface salinity missions by providing valuable information on satellite products (SMOS, Aquarius, SMAP), an extensive in situ database (e.g., Argo, thermosalinographs, moorings, drifters) and additional thematic datasets (precipitation, evaporation, currents, sea level anomalies, sea surface temperature, etc.). Co-localized databases between satellite products and in situ datasets are systematically generated together with validation analysis reports for 30 predefined regions. The data and reports are made fully accessible through the web interface of the platform. The datasets, validation metrics and tools (automatic, user-driven) of the platform are described in detail in this paper. Several dedicated scientific case studies involving satellite SSS data are also systematically monitored by the platform, including major river plumes, mesoscale signatures in boundary currents, high latitudes, semi-enclosed seas, and the high-precipitation region of the eastern tropical Pacific. Since 2019, a partnership in the Salinity Pi-MEP project has been agreed between ESA and NASA to enlarge focus to encompass the entire set of satellite salinity sensors. The two agencies are now working together to widen the platform features on several technical aspects, such as triple-collocation software implementation, additional match-up collocation criteria and sustained exploitation of data from the SPURS campaigns

ocean↗

Multi-Artifact Analysis of Self-Admitted Technical Debt in Scientific Software

Context: Self-admitted technical debt (SATD) occurs when developers acknowledge shortcuts in code. In scientific software (SSW), such debt poses unique risks to the validity and reproducibility of results. Objective: This study aims to identify, categorize, and evaluate scientific debt, a specialized form of SATD in SSW, and assess the extent to which traditional SATD categories capture these domain-specific issues. Method: We conduct a multi-artifact analysis across code comments, commit messages, pull requests, and issue trackers from 23 open-source SSW projects. We construct and validate a curated dataset of scientific debt, develop a multi-source SATD classifier to guide SATD management, and conduct a practitioner validation to assess the practical relevance of scientific debt. Results: Our classifier performs strongly across 900,358 artifacts from 23 SSW projects. SATD is most prevalent in pull requests and issue trackers, underscoring the value of multi-artifact analysis. Models trained on traditional SATD often miss scientific debt, emphasizing the need for its explicit detection in SSW. Practitioner validation confirmed that scientific debt is both recognizable and useful in practice. Conclusions: Scientific debt represents a unique form of SATD in SSW that that is not adequately captured by traditional categories and requires specialized identification and management. Our dataset, classification analysis, and practitioner validation results provide the first formal multi-artifact perspective on scientific debt, highlighting the need for tailored SATD detection approaches in SSW.

Melin, Eric [Boise State University]↗

Climate Fingerprinting Sounder Product (ClimFiSP) Skin Temperature Trends Analysis

Climate fingerprinting Sounder Product (ClimFiSP) has been developed at NASA Langley Research Center which includes daily skin temperature, surface emissivity, air temperature, H2O, trace gases, and cloud properties on 0.5x0.5 grid. Those properties are derived from IR hyper-spectral radiance measured by AIRS on Aqua and CrIS on SNPP and JPSS series. Global skin temperature trends have been derived using more than 20 years of monthly mean skin temperature data from ClimFiSP. The ClimFiSP algorithm use the spectral fingerprinting methodology that allows a low latency procession of more than two decades long satellite data record. The computational cost can be reduced by more than two orders of magnitude as compared with traditional Level-Level2-Level3 retrieval algorithms. In this presentation the global skin temperature trend from ClimFiSP will be compared with skin temperature trend derived from CLIMCAPS, ERA5, GISTEMP, HadCRUT5, and IASI data products, and the results show that pattern of ClimFiSP global skin temperature trends overall matches well with other datasets. The zonally averaged skin temperature anomaly will also be validated using those datasets. It is expected that ClimFiSP surface temperature data can serve as an important complement for surface-based estimates, especially in the regions where the spatial coverage of the surface-based observations is scarce.

Liqiao Lei↗

Toward Drilling the Perfect Geothermal Well: An International Research Coordination Network for Geothermal Drilling Optimization Supported by Deep Machine Learning and Cloud Based Data Aggregation

The EDGE project, supported by the U.S. Department of Energy Geothermal Technologies Office under award DE-EE0008793, established a data-driven framework for improving the efficiency, cost-effectiveness, and reliability of geothermal well drilling. The project focused on developing scalable data infrastructure, advanced machine learning and probabilistic models, and integrated analytics tools to support continuous drilling optimization. A central objective was to reduce geothermal drilling costs by up to seventy percent while minimizing the risk of well failure through predictive diagnostics and adaptive planning. Over the project period, a comprehensive data repository was designed and deployed, incorporating records from over one hundred geothermal wells across varied geological settings. This repository supported both structured and unstructured data and adhered to FAIR data principles, enabling provenance tracking, quality control, and standardized metadata. The project introduced automated ingestion pipelines and a cloud-hosted platform that facilitated access to raw, processed, and derived datasets. This infrastructure served as the foundation for model development and analysis. Machine learning workflows were developed to predict key drilling metrics including rate of penetration, non-productive time, and total drilling costs. Self-organizing maps and dimensionality reduction methods were used to uncover operational patterns and outliers, while supervised learning algorithms such as random forests and deep neural networks were applied to forecast performance outcomes. The models were validated on heterogeneous datasets from both U.S. and Icelandic fields, demonstrating variable but significant predictive accuracy. The results indicated that finer temporal resolution, inclusion of lithological data, and consistency in operational annotations could substantially improve model performance. The project also implemented process mining techniques to reconstruct state-transition models from drilling event logs. These models enabled the identification of deviations from optimal workflows and provided insights into recurring failure modes. Analysis of non-productive time highlighted the impact of equipment failures, geological challenges, and human factors, offering opportunities for targeted mitigation strategies. The EDGE Dashboard was developed as a web-based expert system integrating data visualization, model outputs, and user-driven queries. It provided an accessible interface for operators to explore historical data, evaluate predicted outcomes, and compare drilling scenarios. Initial feedback from project partners suggested that the dashboard could serve as a foundation for more advanced advisory and optimization tools. Overall, the EDGE project demonstrated the feasibility and value of applying modern data science techniques to geothermal drilling. It delivered a set of interoperable tools and models that can support more efficient, lower-risk well development. The findings point toward a viable path for transitioning from advisory analytics to semi-autonomous drilling systems, contingent on continued collaboration, expanded datasets, and field validation. The project results have immediate relevance for drilling operations, data management practices, and future geothermal R&D efforts aimed at achieving reliable, cost-competitive geothermal energy at scale.

15 GEOTHERMAL ENERGY↗

Open Source Synergy: Developing and Validating PMU Data Analysis Techniques Using Open Source Tools and Datasets

This paper presents an exploration into the development and validation of data analysis approaches for Phasor Measurement Units (PMUs) using open-source datasets and tools. Various methods for event detection, event classification, frequency response, and oscillation analysis were tested. We leverage the capabilities of Archive Walker (AW), the Frequency Response Analysis Tool (FRAT), and the Oscillation Baselining and Analysis Tool (OBAT), all open-source tools, for efficient processing and analysis of synchrophasor data. The open-source Transmission Signature Library (TSL) dataset was employed as a dataset for a comprehensive evaluation to assess the performance and reliability of the proposed methods.

PMU, event analysis, oscillation, Frequency Respon↗

Verified, Archived, Library of Inputs and Data (VALID) Supporting Files

This dataset contains input, output, and sensitivity data files for computational simulations with the SCALE code system as part of the Verified, Archived Library of Inputs and Data (VALID). The simulations cover critical benchmark experiments from the International Criticality Safety Benchmark Evaluation Project. The files are to be housed in a public directory for distribution. The information contained in the files have been approved for release by the Organisation for Economic Co-operation and Development Nuclear Energy Agency (NEA). Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases.

keff↗

Midwest Water Resources II: Evaluating Evapotranspiration with NASA Earth Observations and In Situ Observations to Understand Water Balance in Midwest Agriculture

Seasonal water variability in the midwestern United States extensively affects the agricultural community, as it impacts irrigation schedules, growing seasons, and overall ecosystem function. Evapotranspiration (ET) is a critical climatic variable in the water cycle and is used to evaluate spatiotemporal trends in drought and flood conditions. The NASA DEVELOP team partnered with the United States Department of Agriculture (USDA) Midwest Climate Hub, the Minnesota Department of Agriculture, the Illinois State Water Survey, and Michigan State University to compare remotely sensed ET products with in situ observations from January 2001 through December 2020. Remotely sensed actual ET (aET) data were sourced from NASA’s Terra Moderate Resolution Imaging Spectroradiometer (MODIS), and reference ET (refET) data were derived from the Gridded Surface Meteorological (gridMET) dataset. For in situ comparison, aET data were downloaded from the AmeriFlux database while refET data were collected from the Illinois Climate Network and Michigan State University’s Enviro-weather database. For a holistic assessment of ET, this project generated comparisons between remotely sensed and in situ observations, calculated descriptive statistics for validation between refET datasets, and spatially produced statistical validation maps regarding in situ sites. The temporal and spatial gaps of AmeriFlux data limited aET analysis. This comparative assessment of ET products across the Midwest can be used by project partners to assess regional water trends and guide future land management decisions.

Addison Pletcher↗

Atmospheric Correction Inter-comparison eXercise, ACIX-II Land: An Assessment of Amospheric Correction Processors for Landsat 8 and Sentinel-2 Over Land

The correction of the atmospheric effects on optical satellite images is essential for quantitative and multi-temporal remote sensing applications. In order to study the performance of the state-of-the-art methods in an integrated way, a voluntary and open-access benchmark Atmospheric Correction Inter-comparison eXercise (ACIX) was initiated in 2016 in the frame of Committee on Earth Observation Satellites (CEOS) Working Group on Calibration & Validation (WGCV). The first exercise was extended in a second edition wherein twelve atmospheric correction (AC) processors, a substantially larger testing dataset and additional validation metrics were involved. The sites for the inter-comparison analysis were defined by investigating the full catalogue of the Aerosol Robotic Network (AERONET) sites for coincident measurements with satellites' overpass. Although there were more than one hundred sites for Copernicus Sentinel-2 and Landsat 8 acquisitions, the analysis presented in this paper concerns only the common matchups amongst all processors, reducing the number to 79 and 62 sites respectively. Aerosol Optical Depth (AOD) and Water Vapour (WV) retrievals were consequently validated based on the available AERONET observations. The processors mostly succeeded in retrieving AOD for relatively light to medium aerosol loading (AOD < 0.2) with uncertainties <0.08, while the overall uncertainty values were typically 0.23 ± 0.15. Better performances were observed for WV retrievals with >90% of the results falling within the suggested empirical specifications and with the Root Mean Square Error (RMSE) being mostly <0.25 g/cm2. Regarding Surface Reflectance (SR) validation two main approaches were followed. For the first one, a simulated SR reference dataset was computed over all of the test sites by using the 6SV (Second Simulation of the Satellite Signal in the Solar Spectrum vector code) full radiative transfer modelling (RTM) and AERONET measurements for the required aerosol variables and water vapour content. The performance assessment demonstrated that the retrievals were not biased for most of the bands. The uncertainties ranged from approximately 0.003 to 0.01 (excluding B01) for the best performing processors in both sensors' analyses. For the second one, measurements from the radiometric calibration network RadCalNet over La Crau (France) and Gobabeb (Namibia) were involved in the validation. The performance of the processors was in general consistent across all bands for both sensors and with low standard deviations (<0.04) between on-site and estimated surface reflectance. Overall, our study provides a good insight of AC algorithms' performance to developers and users, pointing out similarities and differences for AOD, WV and SR retrievals. Such validation though still lacks of ground-based measurements of known uncertainty to better assess and characterize the uncertainties in SR retrievals.

Atmospheric correction↗

Characterizing Seasonal Variation of the Atmospheric Mixing Layer Height Using Machine Learning Approaches

As machine learning becomes more integrated into atmospheric science, XGBoost has gained popularity for its ability to assess the relative contributions of influencing factors in the atmospheric boundary layer height. To examine how these factors vary across seasons, a seasonal analysis is necessary. However, dividing data by season reduces the sample size, which can affect result reliability and complicate factor comparisons. To address these challenges, this study replaces default parameters with grid search optimization and incorporates cross-validation to mitigate dataset limitations. Using XGBoost with four years of data from the atmospheric radiation measurement (ARM) (Southern Great Plains (SGP) C1 site, cross-validation stabilizes correlation coefficient fluctuations from 0.3 to within 0.1. With optimized parameters, the R value can reach 0.81. Analysis of the C1 site reveals that the relative importance of different factors changes across seasons. Lower tropospheric stability (LTS, ~0.53) is the dominant factor at C1 throughout the year. However, during DJF, latent heat flux (LHF, 0.44) surpasses LTS (0.22). In SON, LTS (0.58) becomes more influential than LHF (0.18). Further comparisons among the four long-term SGP sites (C1, E32, E37, and E39) show seasonal variations in relative importance. Notably, during JJA, the differences in the relative importance of the three factors across all sites are lower than in other seasons. This suggests that boundary layer development in the summer is not dominated by a single factor, reflecting a more intricate process likely influenced by seasonal conditions such as enhanced convective activity, higher temperatures, and humidity, which collectively contribute to a balanced distribution of parameter impacts. Furthermore, the relative importance of LTS gradually increases from morning to noon, indicating that LTS becomes more significant as the boundary layer approaches its maximum height. Consequently, the LTS in the early morning in autumn exhibits greater relative importance compared to other seasons. This reflects a faster development of the mixing layer height (MLH) in autumn, suggesting that it is easier to retrieve the MLH from the previous day during this period. The findings enhance understanding of boundary layer evolution and contribute to improved boundary layer parameterization.

54 ENVIRONMENTAL SCIENCES↗

Retrievals of Aerosol Optical Depth Over the Western North Atlantic Ocean During ACTIVATE

Aerosol optical depth was retrieved from two airborne remote sensing instruments, the Research Scanning Polarimeter (RSP) and Second Generation High Spectral Resolution Lidar (HSRL-2), during the National Aeronautics and Space Administration (NASA) Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE). The field campaign offers a unique opportunity to evaluate an extensive 3-year dataset under a wide range of meteorological conditions from two instruments on the same platform. However, a long-standing issue in atmospheric field studies is that there is a lack of reference datasets for properly validating field measurements and estimating their uncertainties. Here we address this issue by using the triple collocation method, in which a third collocated satellite dataset from the Moderate Resolution Imaging Spectroradiometer (MODIS) is introduced for comparison. HSRL-2 is found to provide a more accurate retrieval than RSP over the study region. The error standard deviation of HSRL-2 with respect to the ground truth is 0.027. Moreover, this approach enables us to develop a simple, yet efficient, quality control criterion for RSP data. The physical reasons for the differences in two retrievals are determined to be cloud contamination, aerosols near the surface, multiple aerosol layers, absorbing aerosols, non-spherical aerosols, and simplified retrieval assumptions. These results demonstrate the pathway for optimal aerosol retrievals by combining information from both lidars and polarimeters for future airborne and satellite missions.

Aerosol optical depth↗

Enhancing Air Traffic Control Planning with Automatic Speech Recognition

The decisions made during the Federal Aviation Administration Air Traffic Control System Command Center's planning teleconferences hold significant sway over the National Airspace System. Held every two hours, these teleconferences convene air traffic managers and stakeholders from across the nation to discuss airspace conditions, weather, and constraints, leading to the formulation and adjustment of traffic management initiatives. Given the critical nature of these decisions, the need for accurate and efficient record-keeping is paramount. In recent years, the application of automatic speech recognition has gained popularity across diverse industries, including aviation. While traditional applications focus on transcribing air traffic control communication, this paper explores a unique application of automatic speech recognition by converting the audio from planning teleconferences into text transcriptions. This innovative approach addresses key challenges in the field, presenting potential benefits for quality assurance, real-time participation, and downstream natural language processing tasks. A notable breakthrough in the machine learning community, namely the transformer neural network architecture, forms the backbone of the proposed solution in this paper. The transformer architecture's role in this research represents a paradigm shift in the efficiency of automatic speech recognition models. By reducing the amount of in-domain training data required, this architecture allows for the fine-tuning of such models like Whisper, originally pretrained on vast English speech datasets. The adaptability of the transformer architecture proves invaluable in capturing the nuances of aviation terminology and specific language used in planning teleconferences. Leveraging the Whisper model as a baseline, our research details the fine-tuning and validation using a dataset comprising 20 hours of meticulously transcribed planning teleconferences. Notably, the baseline pretrained Whisper model exhibited a word error rate of 18.77%. Through the fine-tuning process, the model achieved a substantial improvement, demonstrating an impressive performance with a reduced word error rate of 6.82%. This substantial decrease in WER not only highlights the effectiveness of the transformer architecture but also emphasizes the practical advancements achieved through the application of automatic speech recognition in this specific domain. The utilization of automatic speech recognition in planning teleconferences in this work introduces several novelties. Firstly, the creation of text transcriptions offers a valuable tool for quality assurance and facilitates the efficient review of teleconferences. This is an important aspect of the proposed solution, given the time-sensitive and high-stakes nature of decisions made during these meetings. Furthermore, text-searchable transcriptions provide a streamlined approach for locating and validating critical information, potentially saving hours of manual effort in searching through audio recordings. Moreover, our research identifies a key use case for external facilities and stakeholders. In situations where attendance at the planning teleconference is not feasible, having access to text transcriptions in real-time or shortly after the teleconference ends, proves to be a time-saving and informative resource. This feature enhances collaboration and ensures that stakeholders can stay abreast of important discussions and decisions even in their absence. Despite the efficiency gains facilitated by the transformer architecture in automatic speech recognition technology, it is essential to acknowledge the human factors in data creation. Subject matter experts play a crucial role in accurately transcribing planning teleconferences due to the specificity and complexity of the information discussed. The research dataset, consisting of 20 hours of transcribed planning teleconferences, forms the foundation for fine-tuning and validating the Whisper model. The achieved word error rate of 6.82% demonstrates promising advancements, particularly in recognizing essential aviation terminology within the teleconferences. In conclusion, this paper presents a comprehensive exploration of the application of automatic speech recognition in Air Traffic Control System Command Center planning teleconferences, leveraging the transformer architecture for enhanced efficiency. The novel contributions lie in the improved accessibility of decision-making records, real-time participation opportunities for external stakeholders, and the potential for downstream natural language processing advancements. As the aviation industry continues to evolve, the integration of automatic speech recognition technologies holds the promise of revolutionizing decision-making processes and contributing to the overall safety and efficiency of air traffic management.

ATM↗

Heterogeneous Multi-Domain Dataset Synthesis to Facilitate Privacy and Risk Assessments in Smart City IoT

The emergence of the Smart Cities paradigm and the rapid expansion and integration of Internet of Things (IoT) technologies within this context have created unprecedented opportunities for high-resolution behavioral analytics, urban optimization, and context-aware services. However, this same proliferation intensifies privacy risks, particularly those arising from cross-modal data linkage across heterogeneous sensing platforms. To address these challenges, this paper introduces a comprehensive, statistically grounded framework for generating synthetic, multimodal IoT datasets tailored to Smart City research. The framework produces behaviorally plausible synthetic data suitable for preliminary privacy risk assessment and as a benchmark for future re-identification studies, as well as for evaluating algorithms in mobility modeling, urban informatics, and privacy-enhancing technologies. As part of our approach, we formalize probabilistic methods for synthesizing three heterogeneous and operationally relevant data streams—cellular mobility traces, payment terminal transaction logs, and Smart Retail nutrition records—capturing the behaviors of a large number of synthetically generated urban residents over a 12-week period. The framework integrates spatially explicit merchant selection using K-Dimensional (KD)-tree nearest-neighbor algorithms, temporally correlated anchor-based mobility simulation reflective of daily urban rhythms, and dietary-constraint filtering to preserve ecological validity in consumption patterns. In total, the system generates approximately 116 million mobility pings, 5.4 million transactions, and 1.9 million itemized purchases, yielding a reproducible benchmark for evaluating multimodal analytics, privacy-preserving computation, and secure IoT data-sharing protocols. To show the validity of this dataset, the underlying distributions of these residents were successfully validated against reported distributions in published research. We present preliminary uniqueness and cross-modal linkage indicators; comprehensive re-identification benchmarking against specific attack algorithms is planned as future work. This framework can be easily adapted to various scenarios of interest in Smart Cities and other IoT applications. By aligning methodological rigor with the operational needs of Smart City ecosystems, this work fills critical gaps in synthetic data generation for privacy-sensitive domains, including intelligent transportation systems, urban health informatics, and next-generation digital commerce infrastructures.

IoT↗