KASI AI Hub KASI AI Hub

Papers

Showing 114 of 114 KASI papers with AI/ML connections, curated from 2021–2026.

2026 (10 papers)

A new deep-learning approach to infer solar and geomagnetic parameters for the 1859 Carrington event

Lee et al. (2026)

KASI: Eunsu Park (#3)

In this study, we investigate the solar and geomagnetic parameters of the 1859 Carrington event using deep learning and empirical relationships. For this, we apply an image translation model, a popular deep learning method based on conditional Generative Adversarial Networks, to the generation of magnetograms from sunspot drawings. We train the model using pairs of sunspot data from Debrecen Photoheliographic Data and their corresponding Solar and Heliospheric Observatory/Michelson Doppler Imager (SOHO/MDI) and Solar Dynamics Observatory/Helioseismic and Magnetic Imager (SDO/HMI) magnetograms from 1996 to 2018, using data from January-July and December of each year for training and data from August and November for validation. To test the model, we compare actual magnetograms with artificial-intelligence-based (AI-based) ones for September and October. Our results show that the unsigned magnetic fluxes of AI-based magnetograms closely match those of the originals. Applying this model to Carrington’s full-disk sunspot drawing of 1 September 1859, we generate an AI-based magnetogram and estimate its unsigned magnetic flux. To estimate solar and geomagnetic parameters, we use the following empirical relationships: magnetic flux and flare peak flux, magnetic flux and coronal mass ejection (CME) speed, CME speed and transit time, CME speed and interplanetary coronal mass ejection (ICME) speed, and ICME speed and the Disturbance Storm Time (Dst) index to obtain upper-limit estimates for an extreme event. We find that the estimated Sun-Earth transit time is 16.7 h, consistent with the historical observations. The corresponding Dst value is about -1313 nT, which is broadly consistent with previous reconstruction-based estimates for the Carrington storm.

Tier 3solar/heliophysicsGAN

A Convolutional Neural Network-Transformer Denoiser for Low-signal-to-noise-ratio Galaxy Spectra: Stellar Population Recovery in Synthetic Tests

Kim et al. (2026)

KASI: Kim, Suk (#1), Lee, Joon Hyeop (#2)

Stellar population measurements in integral field unit surveys are often limited by low signal-to-noise ratios(S/Ns) in low-surface-brightness spaxels. Using controlled synthetic experiments, we investigate whether a deeplearning-based denoising can recover stellar population information from such spectra without requiring spatial binning. We introduce the Enhanced U-Net Transformer (EUT), a one-dimensional convolutional neural network-transformer model trained on 90,000 synthetic spectra constructed from MILES simple stellar population (SSP) models following J. H. Lee et al., with wavelength-dependent noise injected on the fly to emulate SAMI-like data (S/N - 5-20, measured in a 4484.77-4573.12 A continuum window). Utilizing an independent test set of 10,000 spectra, the EUT reduces the full-spectrum rms residual by -96.5% at S/N = 5(and by -94% at S/N = 20), achieving recovery rates of -99.8% (the Pearson correlation coefficient between the noise-free and comparison spectra expressed in percent). In fixed windows around Ca II H, Hδ, Hβ, Fe I 4383, Mg b, and Na D, residuals decrease by -88% while preserving line-profile structure. In downstream analysis withPPXF we assess parameter recovery using the Pearson correlation coefficient Rp and the rms scatter: the scatter in recovered mass-weighted age decreases from -0.41 to -0.25 dex at S/N = 5 and from -0.32 to -0.22 dex at S/N = 10; the corresponding mass-weighted global metallicity, [M/H], scatter decreases from -0.45 to-0.36 dex and from -0.32 to -0.28 dex. At S/N = 20, denoising yields results consistent with those from the noisy inputs within the synthetic-test uncertainties. These controlled experiments suggest that hybrid CNN-transformer denoisers can enhance the usable low-surface-brightness area for stellar population studies, although further validation with observed spectra will be needed before practical application.

Tier 3galaxiesCNNtransformer

The SPHEREx Ices Investigation: An Overview

Melnick et al. (2026)

KASI: 김재영 (#5), 이정은 (#9), 강미주 (#11) +6

SPHEREx is a NASA mission designed to perform an all-sky spectroscopic survey in the 0.75-5 μm wavelength range. Its primary science objectives are to investigate: (1) inflationary cosmology; (2) the history of galaxy formation; and (3) the abundance of molecular ices-critical for prebiotic chemistry-found on the surfaces of interstellar dust grains within planet-forming regions. This paper focuses on the third theme, the SPHEREx Ices Investigation, for which SPHEREx is conducting a spectroscopic survey of nearly 10 million preselected sources throughout the Milky Way and Magellanic Clouds, to characterize their ice absorption features. By selecting targets based on infrared color, spatial isolation, and brightness, the Ices Investigation secures high-signal-to-noise-ratio spectra across a broad range of astrophysical environments that are relatively free of spectral contamination. Rather than attempting to decompose each spectrum into its individual ice components, the Ices Investigation prioritizes accurate measurements of the integrated optical depths of key molecular ice absorption features. This approach enables statistically powerful correlation studies between ice abundances and environmental parameters-including extinction, temperature, gas composition, radiation field strength, cosmic-ray flux, and star formation activity. The data pipeline developed for this purpose incorporates machine learning for continuum estimation, drawing on both SPHEREx and ancillary data sets. Ultimately, the expansive spectral archive produced by SPHEREx, combined with targeted follow-up from facilities like JWST, will transform our understanding of Galactic ice formation, evolution, abundance, and their inheritance into planetary systems and prebiotic inventories.

Tier 1ISMclassical-ML

K-DRIFT Science Theme: Illuminating the Next Era of Galaxy Cluster Science

Yoo et al. (2026)

KASI: Jaewon Yoo (#1), Kyungwon Chun (#2), Ko, Jongwan (#3) +10

The KASI Deep Rolling Imaging Fast Telescope (K-DRIFT) is a pioneering instrument designed to explore low-surface-brightness (LSB) phenomena. This white paper presents a compelling set of science cases that showcase K-DRIFT’s unique capabilities in unraveling the mysteries of intracluster light (ICL) and other LSB components within galaxy clusters. Exploring the origin of ICL in galaxy clusters and comparing the spatial distributions of ICL and dark matter will offer new insights into galaxy cluster dynamics. Moreover, investigating LSB objects in galaxy clusters, such as LSB structures in the brightest cluster galaxies, ultra-diffuse galaxies, and tidal features, will enhance our understanding of galaxy evolution within the cluster environment. We present our strategies for addressing scientific queries, encompassing LSB observation and analysis techniques, specialized simulations, and machine-learning approaches. Additionally, we examine the potential synergies between K-DRIFT and other ongoing and forthcoming multi-wavelength surveys. This white paper advocates for the recognition and support of K-DRIFT as a dedicated tool for advancing our understanding of the universe’s subtlest phenomena.

Tier 1galaxiesclassical-ML

Deep Learning-Based Prediction of Tool Influence Function for Nanometric Control in Space Optical Material

Han et al. (2026)

KASI: Han Jeong-Yeol (#2), 이지우 (#3), 박은수 (#4)

Polishing is a critical process in fabricating space-telescope mirrors because it determines the surface figure and consequently optical performance. Deterministic polishing relies on the tool influence function (TIF), which describes the spatial materialremoval profile. At nanometric removal depths, the TIF becomes highly sensitive to process conditions, limiting the accuracy of analytic models such as Preston’s equation. In this study, we propose a deep learning-based approach to predict TIF depth for polishing a Silicon Carbide (SiC) mirror surface. To mitigate data scarcity, we augment 231 experimental measurements with Gaussian noise consistent with the repeatability observed in repeated trials (- 20 nm peak-to-peak). The resulting model achieves a validation mean absolute error (MAE) of 4.24 nm and a test MAE of 3.99 nm; on nine additional experimental cases, the MAE is 6.75 nm. These results indicate that the proposed augmentation improves robustness to experimental variability and supports the development of a data-driven, automated polishing workflow

Tier 2instrumentationCNN

Development Directions in Ground-Based Optical Tracking Sensors for Space Surveillance: Analysis and Recommendations for South Korea

Choi et al. (2026)

KASI: 최진 (#1), Dong-Goo Roh (#2), Myung-Jin Kim (#3) +5

Ground-based optical tracking for space situational awareness is evolving rapidly in response to the sharp growth of the number of resident space objects and the rise of commercial actors in the New Space era. We synthesize qualitative advances over the last five years and outline development priorities for Korea’s ground-based optical tracking sensor systems. Current situation feature commercial off-the-shelf (COTS)-driven miniaturization and networked operations via simplified observatory infrastructure; the emergence of short-wave infrared (SWIR)-enabled cathemeral (day-and-night) tracking as a core capability; and the adoption of artificial intelligence (AI) data processing and autonomous operations. These advances have reopened the low-Earth-orbit (LEO) regime to optical tracking and reduced warning latency through persistent monitoring. Based on these findings, we recommend prioritizing cathemeral performance-centered on SWIR sensing and end-to-end automation-to secure robust LEO tracking capability. We further assess the expected gains from cathemeral observation of national LEO satellites using an observation-opportunity simulation.

Tier 1otherclassical-ML

The redshifts from 122 bands: Comparative redshift forecast for low-resolution spectra from SPHEREx and the 7-Dimensional Sky Survey (7DS)

Bae et al. (2026)

KASI: Bomee Lee (#2), Jeong Woong-Seob (#18)

The recently initiated SPHEREx and 7DS surveys will deliver low-resolution spectra (R- 30-130) for hundreds of millions of galaxies over the optical to near-infrared range (0.4-5.0,μ m), covering a wide sky area without sample selection. These unique datasets will improve redshift estimation and provide a rich redshift catalog for the community. In this study, we forecast the performance of photometric redshift estimations using simulated SPHEREx and 7DS data. Four widely used template-fitting approaches and two machine-learning (ML) methods are used to derive photometric redshifts from low-resolution spectrophotometric data. We measured redshifts using mock catalogs based on the GAMA and COSMOS galaxy samples and achieved high precision for bright (13 < i < 18) galaxies, with NMAD łesssim 0.005, bias łesssim 0.005, and a catastrophic failure rate łesssim 0.005 for all methods employed. We find that the combined SPHEREx + 7DS dataset significantly improves redshift estimation compared to using either the SPHEREx or 7DS datasets alone, highlighting the synergy between the two surveys. Moreover, we compare the redshift estimation performance across magnitude ranges for the different methods and examine the probability distribution functions (PDFs) produced by the template-fitting approaches. As a result, we identify some factors that can affect the redshift measurements, for example, treatments on dust extinction or inclusion of flux uncertainty in the ML model. We also show that the PDFs are relatively well calibrated, although the confidence intervals are generally underestimated, particularly for bright galaxies in the template-fitting methods. This study demonstrates the strong potential of SPHEREx and 7DS to deliver improved redshift measurements from low-resolution spectrophotometric data, underscoring the scientific value of jointly utilizing both datasets.

Tier 1galaxiesclassical-ML

AI-Based Improvement of IRI-2020 Electron Density Profiles With COSMIC Radio Occultation Data

Ji et al. (2026)

KASI: Young-Sil Kwak (#3), Jeong-Heon Kim (#5), Woong Jeon (#6)

In this study, we propose an AI-based method to improve the electron density profiles generated by the International Reference Ionosphere (IRI)-2020 model using the Constellation Observing System for Meteorology, Ionosphere, and Climate (COSMIC) radio occultation (RO) data. Specifically, we employ a Multi-Layer Perceptron (MLP), a type of artificial neural network (ANN), which learns to transform IRI-2020 profiles into COSMIC-like profiles using paired data collected between 2007 and 2019. The data set is divided into training (2007-2013), validation (2014, 2019), and test (2015-2018) subsets. Using the test set, we evaluate the performance of our model by calculating the correlation coefficient (CC) and root mean square error (RMSE) between the model outputs and COSMIC electron density profiles. The results show that our model outperforms the IRI model, yielding higher CCs and lower RMSEs. The model demonstrates consistent improvements across various geomagnetic conditions and geographic regions, particularly in low- and mid- latitudes. Further evaluation using incoherent scatter radar (ISR) data from two stations indicates that both our model and the IRI model show comparable performance in capturing the vertical structure of the ionosphere. These findings demonstrate that machine learning techniques, such as MLPs, provide an effective means of leveraging satellite-based observational data to improve the performance of empirical models such as the IRI model.

Tier 1solar/heliophysicsclassical-ML

Forecasting Total Electron Content During Geomagnetic Storms Using Convolutional Long Short-Term Memory (ConvLSTM): Performance and Limitations

Jeong et al. (2026)

KASI: Se-Heon Jeong (#1), 길효섭 (#2), Woo Kyoung Lee (#3) +2

This study investigates the effects of quiet time ionospheric conditions and the number of storm events used for training on the prediction of ionospheric total electron content (TEC) during geomagnetic storms using a deep learning method. A Convolutional Long Short-Term Memory (ConvLSTM) model is employed for training and prediction of regional TEC maps around the Korean Peninsula. To ensure high-resolution, gap-free input data, TEC maps were reconstructed using a Deep Convolutional Generative Adversarial Network-Poisson Blending (DCGAN-PB) method. Geomagnetic storm days were selected based on Dst index values below - 50nT, and for each event, a 24-hr dataset was constructed starting from the minimum Dst time. To address the limited number of storm events, partially overlapping 24-hr segments were extracted from each storm using a sliding-window approach to augment the training data set. To further improve regional prediction accuracy, a region-weighted loss function was introduced, giving additional emphasis to the Korean Peninsula. Results show that the ConvLSTM outperforms both comparison models, achieving a root mean square error (RMSE) of 5.09 TECU compared with 6.23 TECU for the 24-hr-lag persistence model and 8.37 TECU for International Reference Ionosphere-2016. Adding quiet-day data to the training did not improve storm-time performance, suggesting that ionospheric responses during geomagnetic storms are independent of prior-day conditions. However, model's performance improved in proportion to the number of storm events used for training. This result indicates that the availability of storm data is a key factor in accurately predicting storm-time ionospheric plasma density.

Tier 3solar/heliophysicsCNNRNN/LSTM

Revealing Hidden Cosmic Flows through the Zone of Avoidance with Deep Learning

Dupuy et al. (2026)

KASI: Sungwook E. Hong (#3)

We present a refined deep-learning-based method to reconstruct the 3D dark matter density, gravitational potential, and peculiar velocity fields in the Zone of Avoidance (ZOA), a region near the Galactic plane with limited observational data. Using a convolutional neural network (V-Net) trained on A-SIM simulation data, our approach reconstructs density or potential fields from galaxy positions and radial peculiar velocities. The full 3D peculiar velocity field is then derived from the reconstructed potential. We validate the method with mocks that mimic the spatial distribution of the Cosmicflows-4 (CF4) catalog and apply it to actual data. Given CF4's significant observational uncertainties and since our model does not yet account for them, we use peculiar velocities corrected via an existing Hamiltonian Monte Carlo reconstruction, rather than raw catalog distances. Our results demonstrate that the reconstructed density field recovers known galaxy clusters detected in an H I survey of the ZOA, despite this dataset not being used in the reconstruction. This agreement underscores the potential of our method to reveal structures in data-sparse regions. Most notably, streamline convergence and watershed analysis identify a mass concentration consistent with the Great Attractor, at (l, b) = (308 .° 4 ± 2 .° 4, 29 .° 0 ± 1 .° 9) and cz = 4960.1 ± 404.4 km s-1, for 64% of realizations. Our method is particularly valuable as it does not rely on data point density, enabling accurate reconstruction in data-sparse regions and offering strong potential for future surveys with more extensive galaxy datasets.

Tier 2cosmologyCNN

2025 (19 papers)

Artificial-intelligence-based Reconstruction of Solar Farside Vector Magnetograms from Multispacecraft Extreme-ultraviolet Data

Jeong et al. (2025)

KASI: Eunsu Park (#2)

In this study, we generate full-disk vector magnetic field data of the solar farside, as viewed from the Solar Terrestrial Relations Observatory-Ahead (STEREO-A), Solar Terrestrial Relations Observatory-Behind (STEREO-B), and Solar Orbiter (SolO), using a deep learning model based on the Pix2PixCC architecture. Our model takes extreme-ultraviolet (EUV) 304 and 171 A images, together with reference magnetic field data from a surface flux transport (SFT) model, as inputs. To train and evaluate the model, we use EUV images from the Solar Dynamics Observatory (SDO)/Atmospheric Imaging Assembly and SFT-predicted magnetic field data from one solar rotation earlier as inputs, and we use vector magnetograms from the SDO/Helioseismic and Magnetic Imager (HMI) as targets. For frontside test datasets covering Solar Cycles 24 and 25, our model successfully generates all three vector magnetic field components consistent with those from SDO/HMI, showing improved performance compared with previous studies. We then generate vector magnetic field data by applying the trained model to EUV observations from STEREO-A and SolO. For the first time, we compare the artificial intelligence (AI)-generated results with corresponding SDO/HMI data obtained when STEREO-A and SolO were near inferior conjunction with SDO in 2023 and 2022, respectively. We also track active-region magnetic fields and derive vector magnetic parameters using AI-generated farside data from STEREO-A, STEREO-B, and SolO, along with frontside SDO/HMI data. These results demonstrate the potential for continuous monitoring of solar vector magnetic fields and derived parameters from the farside to the frontside using AI-generated and SDO/HMI data.

Tier 3solar/heliophysicsCNN

Extended dark energy analysis using DESI DR2 BAO measurements

Lodha et al. (2025)

KASI: Kushal Lodha (#1), Rodrigo Calderon Bruni (#2), William Luke Matthewson (#3) +2

We conduct an extended analysis of dark energy constraints, in support of the findings of the Dark Energy Spectroscopic Instrument (DESI) second data release cosmology key paper, including DESI data, Planck cosmic microwave background observations, and three different supernova compilations. Using a broad range of parametric and nonparametric methods, we explore the dark energy phenomenology and find consistent trends across all approaches, in good agreement with the w0waCDM (cold dark matter) key paper results. Even with the additional flexibility introduced by nonparametric approaches, such as binning and Gaussian processes, we find that extending ΛCDM to include a two-parameter wðzÞ is sufficient to capture the trends present in the data. Finally, we examine three dark energy classes with distinct dynamics, including quintessence scenarios satisfying w ≥ -1, to explore what underlying physics can explain such deviations. The current data indicate a clear preference for models that feature a phantom crossing; although alternatives lacking this feature are disfavored, they cannot yet be ruled out. Our analysis confirms that the evidence for dynamical dark energy, particularly at low redshift (z - 0.3), is robust and stable under different modeling choices.

Tier 1cosmologygaussian-processclassical-ML

Few EURONEAR NEA mini-surveys observed with the INT, KASI and T80S telescopes during the ParaSOL synthetic tracking project

Vaduvescu et al. (2025)

KASI: Lee, Chung-Uk (#20), Dong-jin Kim (#21)

The modern synthetic tracking technique (ST) can make use of small and medium-sized telescopes to detect asteroids fainter than the classic blinking methods, by arbitrary shifting and co-adding more images of the same survey field, if GPU computing resources are available. In the framework of the Romanian ParaSOL project, we developed and tested an innovative ST algorithm capable of detecting in near-real-time very faint near-Earth asteroids (NEAs), which likely became the first ST pipeline developed in Europe. To test our pipeline, we conducted several mini-surveys using three large-field telescopes, namely the ING's INT, the Korean KASI and the Brazilian T80S telescopes. Most images were processed using our Umbrella Image Processing Pipeline (IPP) module. The ST search was conducted using our Synthetic Tracking via Umbrella (STU) module and the commercial Tycho Tracker software, which allowed to compare and complement the findings. The source validation was supported by reducers using our new Webrella platform. Most of the nights were reduced in near-real-time, demonstrating the ability to process, sort, and report large volumes of data. We discovered 5 credited and 4 one-night NEAs, co-discovered other 3 NEAs and recovered 3 poorly known NEAs. We flagged 59 NEA candidates for recovery and orbital classification, discovering, co-discovering and recovering other 18 orbitally related NEAs, additionally improving the orbits of 23,428 known asteroids and reporting 1,374 unknown objects. A comparison between ST and traditional blinking detection using the new EURONEAR tool MagLim shows improvements of two magnitudes and a two-fold increase in the number of detections. A preliminary comparison between STU and Tycho shows that STU detects about 70% of Tycho findings, however STU detects rapid objects much faster than Tycho, 7 NEAs with speeds between 2-10--/min being found exclusively by STU. Based on our surveys, we assessed the current NEA discovery rate using 1-2-m class telescopes and ST methods, finding that one NEA candidate can be discovered in every 9-12 square degrees up to magnitude ∼23.

Tier 1otherclassical-ML

STag. II. Classification of Serendipitous Supernovae Observed by Galaxy Redshift Surveys

Davison et al. (2025)

KASI: W. Davison (#1), D. Parkinson (#2)

With the number of supernovae observed expected to drastically increase thanks to large-scale surveys like the Dark Energy Spectroscopic Instrument (DESI), it is necessary that the tools we use to classify these objects keep up with this increase. We previously created Supernova Tagging and Classification (STag) to address this problem by employing machine learning techniques alongside logistic regression in order to assign `tags' to spectra based on spectral features. STag II is a continuation of this work, which now makes use of model supernova spectra combined with real DESI spectra in order to train STag to better deal with realistic data. We also make use of the rlap score as a trustworthiness cut, making for a more robust and accurate supernova classifier than before.

Tier 1transients/time-domainclassical-ML

Capturing Star Formation Activity from Compressed Photometric Images of Galaxies

Oh & Turp (2025)

KASI: Kyuseok Oh (#1)

We present a novel approach for classifying star-forming galaxies using photometric images. By utilizing approximately 124,000 optical color composite images and spectroscopic data of nearby galaxies at 0.01 < z < 0.06 from the Sloan Digital Sky Survey, along with follow-up spectroscopic line measurements from the OSSY catalog, and leveraging the vision transformer machine learning technique, we demonstrate that galaxy images in JPEG format alone can be directly used to determine whether star-forming activity dominates the galaxy, bypassing traditional spectroscopic analyses such as emission-line diagnostic diagrams. We anticipate that this method holds significant potential for application in current and future large-scale surveys, such as Euclid, the Dark Energy Survey, and the Legacy Survey of Space and Time.

Tier 2galaxiestransformer

Active galactic nuclei with massive black holes have closer galactic neighbors

Mhatre et al. (2025)

KASI: Kyuseok Oh (#8)

Context. The large-scale environments of active galactic nuclei (AGNs) reveal important information on the growth and evolution of supermassive black holes (SMBHs). Previous AGN clustering measurements using two-point correlation functions have hinted that AGNs with massive black holes preferentially reside in denser cosmic regions than AGNs with less massive SMBHs. At the same time, little to no dependence on the accretion rate has been found; however, the significance of such trends has been limited. Aims. Here, we present kth-nearest-neighbor (kNN) statistics of 2MASS galaxies around AGNs from the Swift/BAT AGN Spectroscopic survey. These statistics have been shown to contribute additional higher order clustering information on the cosmic density field. Methods. By calculating the distances to the nearest seven galaxy neighbors in angular separation to each AGN within two redshift ranges (0.01 < z < 0.03 and 0.03 < z < 0.06), we compared their cumulative distribution functions to that of a randomly distributed sample to demonstrate the sensitivity of this method to the clustering of AGNs. We also split the AGNs into bins of bolometric luminosity, black hole mass, and Eddington ratio (while controlling for redshift) to search for trends between kNN statistics and fundamental AGN properties. Results. We find that AGNs with massive SMBHs have significantly closer neighbors than AGNs with less massive SMBHs (at the 99.98% confidence level), especially in our lower redshift range. We find less significant trends with luminosity or Eddington ratio. By comparing our results to empirical SMBH-galaxy-halo models implemented in N-body simulations, we show that small-scale kNN trends with black hole mass may go beyond stellar mass dependencies. Conclusions. This suggests that massive SMBHs in the local universe reside in more massive dark matter halos and denser regions of the cosmic web, which may indicate that environment is important for the growth of SMBHs, bolstering prior conclusions based on correlation functions.

Tier 1galaxiesclassical-ML

Comparison of Empirical and Deep Learning Models for Solar Wind Speed Prediction

Ahn et al. (2025)

KASI: Jihyeon Son (#2)

In this study, we compare representative empirical models with a deep learning model for predicting solar wind speed at 1 au. The empirical models are the Wang-Sheeley-Arge-ENLIL model, which combines empirical methods with a magnetohydrodynamic model, and the empirical solar wind forecast model, which uses the relationship between the fractional coronal hole area and solar wind speed. Our deep learning model predicts solar wind speed over 3 days ahead using extreme-ultraviolet images and up to 5 days of solar wind speed before the prediction date. We evaluate the models over the test period (October-December in each year from 2012 to 2020) in view of solar activity phases and the entire period. To validate the model’s performance, we use two evaluation methods: a statistical approach and an event-based approach. For statistical verification during the entire period, our model outperforms the other empirical models, with a much lower mean absolute error of 51.4 km s-1 and rms error of 68.6 km s-1, along with a much higher correlation coefficient of 0.69. For the event-based verification for high-speed solar wind streams, our model has superior performance in most of the six metrics evaluated within a ±1 day time window. In particular, it achieves a high success ratio of 0.82, emphasizing the model’s stable performance and ability to minimize false alarms. These results show that our deep learning model has strong potential for practical application as a reliable tool for fast solar wind forecasting with its high accuracy and stability.

Tier 2solar/heliophysicsCNN

ODIN: Star Formation Histories Reveal Formative Starbursts Experienced by Lyα- emitting Galaxies at Cosmic Noon

Firestone et al. (2025)

KASI: Yujin Yang (#5), Sungryong Hong (#11), Jeong Woong-Seob (#14) +2

In this work, we test the frequent assumption that Lyα-emitting galaxies (LAEs) are experiencing their first major burst of star formation at the time of observation. To this end, we identify 74 LAEs from the ODIN Survey with rest-UV-through-NIR photometry from UVCANDELS. For each LAE, we perform nonparametric star formation history (SFH) reconstruction using the Dense Basis Gaussian-process-based method of spectral energy distribution fitting. We find that a strong majority (67%) of our LAE SFHs align with the frequently assumed archetype of a first major star formation burst, with at most modest star formation rates (SFRs) in the past. However, the rest of our LAE SFHs have significant amounts of star formation in the past, with 28% exhibiting earlier bursts of star formation, with the ongoing burst having the highest SFR (dominant bursts) and the final 5% having experienced their highest SFR in the past (nondominant bursts). Combining the SFHs indicating first and dominant bursts, ∼95% of LAEs are experiencing their largest burst yet: a formative burst. We also find that the fraction of total stellar mass created in the last 200 Myr is ∼1.3 times higher in LAEs than in mass-matched Lyman break galaxy (LBG) samples, and that a majority of LBGs are experiencing dominant bursts, reaffirming that LAEs differ from other star-forming galaxies. Overall, our results suggest that multiple evolutionary paths can produce galaxies with strong observed Lyα emission.

Tier 1galaxiesgaussian-process

Recovering coherent flow structures in active regions using machine learning

Lennard et al. (2025)

KASI: Sung-Hong park (#10)

Analysing high-resolution solar atmospheric observations requires robust techniques to recover plasma flow features across different scales, especially in active regions. Current methodologies often fall short in capturing subgranular-scale flows, and there is limited research on the errors introduced by velocity estimation techniques and analysing the properties of recovered flows in the presence of kG magnetic flux density. This study concentrates on validating the effectiveness of the DeepVel neural network in recovering subgranular to mesogranular-scale topological plasma flow features throughout the total evolution of a simulated active region by tracking tracers, and reproducing coherent patterns. The neural network was trained on the r2d2 radiative MHD simulation depicting the emergence and decay of a magnetic flux tube. DeepVel achieved strong correlations (exceeding 0.7) with flows from an unseen muram simulation, despite being trained on a model with a simpler radiative transfer and lacking thermal resistivity. DeepVel was able to capture the detailed topology well, e.g. the structure of vortical and diverging structures across all scales present in the flows. DeepVel performed slightly less well in the umbra, this is likely explained by magnetic field suppression and reduced contrast. Differences in velocities introduced by DeepVel did not affect Lagrangian analysis; consequently, we demonstrate for the first time that the DeepVel-recovered velocities accurately reflected the flow’s transport barriers. These findings highlight the precision and reliability of the DeepVel and its ability to emulate plasma flows surrounding and within active regions.

Tier 2solar/heliophysicsCNN

Model-independent cosmology with joint observations of gravitational waves and γ-ray bursts

Cozzumbo et al. (2025)

KASI: Rodrigo Calderon Bruni (#3)

Multi-messenger (MM) observations of binary neutron star (BNS) mergers provide a promising approach to trace the distance-redshift relation, crucial for understanding the expansion history of the Universe and, consequently, testing the nature of Dark Energy (DE). While the gravitational wave (GW) signal offers a direct measure of the distance to the source, high-energy observatories can detect the electromagnetic counterpart and drive the optical follow-up providing the redshift of the host galaxy. In this work, we exploit up-to-date catalogs of γ-ray bursts (GRBs) supposedly coming from BNS mergers observed by the Fermi γ-ray Space Telescope and the Neil Gehrels Swift Observatory, to construct a large set of mock MM data. We explore how combinations of current and future generations of GW observatories operating under various underlying cosmological models would be able to detect GW signals from these GRBs. We achieve the reconstruction of the GW parameters by means of a novel prior-informed Fisher matrix approach. We then use these mock data to perform an agnostic reconstruction of the DE phenomenology, thanks to a machine learning method based on forward modeling and Gaussian Processes (GP). Our study highlights the paramount importance of observatories capable of detecting GRBs and identifying their redshift. In the best-case scenario, the GP constraints are 1.5 times more precise than those produced by classical parametrizations of the DE evolution. We show that, in combination with forthcoming cosmological surveys, fewer than 40 GW-GRB detections will enable unprecedented precision on H0 and Ωm, and accurately reconstruct the DE density evolution.

Tier 1cosmologygaussian-processclassical-ML

Six-hour Prediction of Interplanetary Magnetic Field Bz Profiles for Strong Southward Cases by Deep Learning

Son et al. (2025)

KASI: 손지현 (#1), 곽영실 (#3)

In this study, we develop deep learning models to forecast the 6 hr interplanetary magnetic field (IMF) Bz component for southward cases. The models are based on a bidirectional long short-term memory method, and input parameters are solar wind data (V, N, T) and IMF components (Bt, Bx, By, Bz). The data are obtained from OMNI, whose period is from 2000 to 2022. We use the preceding 12 hr of data as input and the subsequent 6 hr of Bz data as target. To focus on strong geomagnetic conditions, we consider periods where Bz values drop below the negative standard deviation (approximately -3 nT) for at least 6 hr. The models are trained and validated using a 12-fold cross-validation process, with each model trained over 8 months of data and tested over 4 months. The ensemble model, which averages 12-fold model results, achieves an RMSE ranging from 1.75 (30 minutes prediction) to 2.55 nT (6 hr prediction), significantly outperforming two baseline methods: multilayer perception and multiple linear regression. Our model can capture both decreasing and increasing phases of Bz, showing reliable performance across varying geomagnetic conditions. Our results suggest a sufficient possibility for predicting Bz under noticeable southward conditions. We expect that our model can be used for subsequent space weather predictions, such as global magnetohydrodynamic simulations in the magnetosphere.

Tier 2solar/heliophysicsRNN/LSTM

Redshift Evolution of the X-Ray and Ultraviolet Luminosity Relation of Quasars: Calibrated Results from SNe Ia

Li et al. (2025)

KASI: Arman Shafieloo (#3)

Quasars could serve as standard candles if the relation between their ultraviolet (UV) and X-ray luminosities can be accurately calibrated. Previously, we developed a model-independent method to calibrate quasar standard candles using the distance-redshift relation reconstructed from TypeIa supernovae (SNeIa) at z<2 using Gaussian process regression. Interestingly, we found that the calibrated quasar standard candle data set preferred a deviation from ΛCDM at redshifts above z > 2. One possible interpretation of these findings is that the calibration parameters of the quasar UV and X-ray luminosity relationship evolves with redshift. In order to test the redshift dependence of the quasar calibration in a model-independent manner, we divided the quasar sample whose redshift overlaps with the redshift coverage of Pantheon+ SNe Ia compilation into two subsamples: a low-redshift quasar subsample and a high-redshift quasar subsample. Assuming all the quasar samples are reliable, our results show that there is about a 4σ inconsistency between the quasar parameters inferred from the subsamples without considering evolution. This inconsistency suggests the possibility of considering redshift evolution for the relationship between the quasars’ UV and X-ray luminosities. We then test an explicit parameterization of the redshift evolution of the quasar calibration parameters via γ(z)=γ0+γ1(1+z) and β(z)=β0+β1(1+z). Combining this redshift- dependent calibration relationship with the distance-redshift relationship reconstructed from the Pantheon+ supernova compilation, we find the high-redshift subsample and low-redshift subsample become consistent at the 2σ level, which means that the parameterized form of γ(z) and β(z) works well at describing the evolution of the quasar calibration parameters.

Tier 1cosmologygaussian-processclassical-ML

Prediction of the Next Solar Rotation Synoptic Maps Using an Artificial Intelligence-based Surface Flux Transport Model

Jeong et al. (2025)

KASI: Ji-Hye Baek (#5), Seonghwan Choi (#7)

In this study, we develop an artificial intelligence (AI)-based solar surface flux transport (SFT) model. We predict synoptic maps for the next solar rotation (27.2753 days) using deep learning. Our model takes the latest synoptic maps and their sine-latitude grid data as inputs. Synoptic maps, which represent global magnetic field distributions on the solar surface, have been widely used as initial boundary conditions in the Sun and space-weather prediction models. Here we train and evaluate our deep-learning model, based on the Pix2PixCC architecture, using data sets of Solar Dynamics Observatory/Helioseismic and Magnetic Imager, Solar and Heliospheric Observatory/ Michelson Doppler Imager, and National Solar Observatory/Global Oscillation Network Group synoptic maps with a resolution of 360 by 180 (longitude and sine latitude) from 1996 to 2023. We present results of our model and compare them with those from the persistent model and the conventional SFT model, including the effects of differential rotation, meridional flow, and diffusion on the solar surface. The average pixel-to-pixel correlation coefficient between the target and our AI-generated data, after 10 by 10 binning with a 10° resolution in longitude, is 0.71. This result is qualitatively similar to the results of the conventional SFT model (0.65-0.68) and better than the results of the persistent model (0.56). Our model successfully generates magnetic features, such as the diffusion of solar active regions and the motions of supergranules. Using synthetic input data with bipolar structures, we confirm that our model successfully reproduces differential rotation and meridional flow. Finally, we discuss the advantages and limitations of our model in view of magnetic field evolution and its potential applications.

Tier 3solar/heliophysicsCNN

Machine Learning-based Photometric Redshifts for Galaxies in the North Ecliptic Pole Wide Field: Catalogs of Spectroscopic and Photometric Redshifts

Kim et al. (2025)

KASI: Jeong Woong-Seob (#8), Narae Hwang (#15), PARK, BYEONG GON (#16)

We perform an MMT/Hectospec redshift survey of the North Ecliptic Pole Wide (NEPW) field covering 5.4 deg2 and use it to estimate the photometric redshifts for the sources without spectroscopic redshifts. By combining 2572 newly measured redshifts from our survey with existing data from the literature, we create a large sample of 4421 galaxies with spectroscopic redshifts in the NEPW field. Using this sample, we estimate photometric redshifts of 77,755 sources in the band-merged catalog of the NEPW field with a random forest model. The estimated photometric redshifts are generally consistent with the spectroscopic redshifts, with a dispersion of 0.028, an outlier fraction of 7.3%, and a bias of -0.01. We find that the standard deviation of the prediction from each decision tree in the random forest model can be used to infer the fraction of catastrophic outliers and the measurement uncertainties. We test various combinations of input observables, including colors and magnitude uncertainties, and find that the details of these various combinations do not change the prediction accuracy much. As a result, we provide a catalog of 77,755 sources in the NEPW field, which includes both spectroscopic and photometric redshifts up to z ∼ 2. This data set has significant legacy value for studies in the NEPW region, especially with upcoming space missions such as JWST, Euclid, and SPHEREx.

Tier 1galaxiesrandom-forest/GBDT

Can we properly determine differential emission measures from Solar Orbiter/EUI/FSI with deep learning?

Youn et al. (2025)

KASI: Eunsu Park (#5)

In this study, we address the question of whether we can properly determine differential emission measures (DEMs) using Solar Orbiter/Extreme Ultraviolet Imager (EUI)/Full Sun Imager (FSI) and AI-generated extreme UV (EUV) data. The FSI observes only two full-disk EUV channels (174 and 304 A), which is insufficient for accurately determining DEMs and can lead to significant uncertainties. To solve this problem, we trained and tested deep learning models based on Pix2PixCC using the Solar Dynamics Observatory (SDO)/Atmospheric Imaging Assembly (AIA) dataset. The models successfully generated five-channel (94, 131, 193, 211, and 335 A) EUV data from 171 and 304 A EUV observations with high correlation coefficients. Then we applied the trained models to the Solar Orbiter/EUI/FSI dataset and generated the five-channel data that the FSI cannot observe. We used the regularized inversion method to compare the DEMs from the SDO/AIA dataset with those from the Solar Orbiter/EUI/FSI dataset, which includes AI-generated data. We demonstrate that, when SDO and Solar Orbiter are at the inferior conjunction, the main peaks and widths of both DEMs are consistent with each other at the same coronal structures. Our study suggests that deep learning can make it possible to properly determine DEMs using Solar Orbiter/EUI/FSI and AI-generated EUV data.

Tier 2solar/heliophysicsCNN

LensNet: Enhancing Real-time Microlensing Event Discovery with Recurrent Neural Networks in the Korea Microlensing Telescope Network

Viaña et al. (2025)

KASI: Kyu-Ha Hwang (#2), Chung, Sun-Ju (#7), Jung Youn Kil (#10) +8

Traditional microlensing event vetting methods require highly trained human experts, and the process is both complex and time consuming. This reliance on manual inspection often leads to inefficiencies and constrains the ability to scale for widespread exoplanet detection, ultimately hindering discovery rates. To address the limits of traditional microlensing event vetting, we have developed LensNet, a machine learning pipeline specifically designed to distinguish legitimate microlensing events from false positives caused by instrumental artifacts, such as pixel bleed trails and diffraction spikes. Our system operates in conjunction with a preliminary algorithm that detects increasing trends in flux. These agged instances are then passed to LensNet for further classi cation, allowing for timely alerts and follow-up observations. Tailored for the multiobservatory setup of the Korea Microlensing Telescope Network and trained on a rich data set of manually classified events, LensNet is optimized for early detection and warning of microlensing occurrences, enabling astronomers to organize follow-up observations promptly. The internal model of the pipeline employs a multibranch Recurrent Neural Network architecture that evaluates time-series flux data with contextual information, including sky background, the full width at half-maximum of the target star, flux errors, point-spread function quality flags, and air mass for each observation. We demonstrate a classification accuracy above 87.5% and anticipate further improvements as we expand our training set and continue to refine the algorithm.

Tier 2exoplanetsRNN/LSTM

Weak-lensing Mass Reconstruction of Galaxy Clusters with a Convolutional Neural Network. II. Application to Next-generation Wide-field Surveys

Cha et al. (2025)

KASI: Sungwook E. Hong (#3)

Traditional weak-lensing mass reconstruction techniques suffer from various artifacts, including noise amplification and the mass-sheet degeneracy. In S. E. Hong et al., we demonstrated that many of these pitfalls of traditional mass reconstruction can be mitigated using a deep learning approach based on a convolutional neural network (CNN). In this paper, we present our improvements and report on the detailed performance of our CNN algorithm applied to next-generation wide-field (WF) observations. Assuming the field of view (3.5 deg x 3.5 deg) and depth (27 mag at 5σ) of the Vera C. Rubin Observatory, we generated training data sets of mock shear catalogs with a source density of 33 arcmin^-2 from cosmological simulation ray-tracing data. We find that the current CNN method provides high-fidelity reconstructions consistent with the true convergence field, restoring both small- and large-scale structures. In addition, the cluster detection utilizing our CNN reconstruction achieves ∼75% completeness down to ∼10^14 M_⊙. We anticipate that this CNN-based mass reconstruction will be a powerful tool in the Rubin era, enabling fast and robust WF mass reconstructions on a routine basis.

Tier 2galaxiesCNN

GECKO Follow-up Observations of the Binary Neutron Star-Black Hole Merger Candidate S230518h

Paek et al. (2025)

KASI: Lee, Chung-Uk (#14), Kim Seung-Lee (#15)

The gravitational-wave (GW) event S230518h is a potential binary neutron star-black hole merger (NSBH) event that was detected during engineering run 15, which served as the commissioning period before the LIGO-Virgo-KAGRA O4a observing run. Despite its low probability of producing detectable electromagnetic emissions, we performed extensive follow-up observations of this event using the Gravitational-wave Electromagnetic Counterpart Korean Observatories (GECKO) telescopes in the Southern Hemisphere. Our observations covered 61.7% of the 90% credible region, a 284 deg2 area accessible from the Southern Hemisphere, reaching a median limiting magnitude of R = 21.6 mag. In these images, we conducted a systematic search for an optical counterpart of this event by combining a convolutional-neural-network-based classifier and human verification. We identified 128 transient candidates, but no significant optical counterpart was found that could have caused the GW signal. Furthermore, we provide feasible kilonova properties that are consistent with the upper limits of the observations. Although no optical counterpart was found, our result demonstrates both GECKO's efficient wide-field follow-up capabilities and usefulness for constraining properties of kilonovae from NSBH mergers at distances of ∼200 Mpc.

Tier 2transients/time-domainCNN

Tracing magnetic switchbacks to their source: An assessment of solar coronal jets as switchback precursors

Bizien et al. (2025)

KASI: Maria S. Madjarska (#3)

Context. The origin of large-amplitude magnetic field deflections in the solar wind, known as magnetic switchbacks, is still under debate. These structures, which are ubiquitous in the in situ observations made by Parker Solar Probe (PSP), likely have their seed in the lower solar corona, where small-scale energetic events driven by magnetic reconnection could provide conditions ripe for either direct or indirect generation. Aims. We investigated potential links between in situ measurements of switchbacks and eruptions originating from the clusters of small-scale solar coronal loops known as coronal bright points to establish whether these eruptions act as precursors to switchbacks. Methods. We traced solar wind switchbacks from PSP back to their source regions using the ballistic back-mapping and potential field source surface methods, and analyzed the influence of the source surface height and solar wind propagation velocity on magnetic connectivity. Using extreme ultraviolet images, we combined automated and visual approaches to identify small-scale eruptions (e.g., jets) in the source regions. The jet occurrence rate was then compared with the rate of switchbacks captured by PSP. Results. We find that the source region connected to the spacecraft varies significantly depending on the source surface height, which exceeds the expected dependence on the solar cycle and cannot be detected via polarity checks. For two corotation periods that are straightforwardly connected, we find a matching level of activity (jets and switchbacks), which is characterized by the hourly rate of events and depends on the size of the region connected to PSP. However, no correlation is found between the two time series of hourly event rates. Modeling constraints and the event selection may be the main limitations in the investigation of a possible correlation. Evolutionary phenomena occurring during the solar wind propagation may also influence our results. These results do not allow us to conclude that the jets are the main switchback precursors, nor do they rule out this hypothesis. They may also indicate that a wider range of dynamical phenomena are the precursors of switchbacks.

Tier 1solar/heliophysicsclassical-ML

2024 (17 papers)

Selecting variable sources with median colours using a self-organising map

Venville et al. (2024)

KASI: Bomee Lee (#4)

A key objective for upcoming surveys, and when re-analysing archival data, is the identification of variable stellar sources. However, the selection of these sources is often complicated by the unavailability of light curve data. Utilising a self-organising map (SOM), we demonstrate the selection of diverse variable source types from a catalogue of variable and non-variable SDSS Stripe 82 sources whilst employing only the median -----, -----, ----- , and ----- photometric colours for each source as input, without using source magnitudes. This includes the separation of main sequence variable stars that are otherwise degenerate with non-variable sources ( ----- , -----) and ( -----, ---- ) colour-spaces. We separate variable sources on the main sequence from all other variable and non-variable sources with a purity of 80.0% and completeness of 25.1% , figures which can be modified depending on the application. We also explore the varying ability of the same method to simultaneously select other types of variable sources from the heterogeneous sample, including variable quasars and RR-Lyrae stars. The demonstrated ability of this method to select variable main sequence stars in colour-space holds promise for application in future survey reduction pipelines and for the analysis of archival data, where light curves may not be available or may be prohibitively expensive to obtain.

Tier 1starsclassical-ML

High precision accelerator for our hybrid model of the redshift space power spectrum

Icaza-Lizaola et al. (2024)

KASI: Miguel Angel C. de Icaza Lizaola (#1), Yong-Seon Song (#2)

Upcoming Large Scale Structure surveys aim to achieve an unprecedented level of pre- cision in measuring galaxy clustering. However, accurately modelling these statistics may require theoretical templates that go beyond two-loop order perturbation theory, especially for achieving precision at smaller scales. In our previous work, we introduced a hybrid model for the redshift space power spectrum of galaxies. This model combines two-loop order templates with N-body simulations to capture the influence of scale-independent parameters on the galaxy power spectrum. However, the impact of scale-dependent parameters was addressed by precomputing a set of input statistics derived from computationally expensive N-body simulations. As a result, exploring the scale-dependent parameter space was not feasible in this approach. To address this challenge, we present an accelerated methodology that utilizes Gaussian processes, a machine learning technique, to emulate these input statistics. Our emu- lators exhibit remarkable accuracy, achieving reliable results with just 13 N-body simulations for training. Our emulators can reproduce the set of statistics we are interested in with less than 0.1% error in the parameter space within 5-- of the Planck ΛCDM predictions, specifically for scales around -- > 0.1 h Mpc-1. Following the training of our emulators, we can predict all inputs for our hybrid model in approximately 0.2 seconds at a specified redshift. Given that performing 13 N-body simulations is a manageable task, our present methodology enables us to construct efficient and highly accurate models of the galaxy power spectra within a manageable time frame.

Tier 1cosmologygaussian-processclassical-ML

The MOST Hosts Survey: Spectroscopic Observation of the Host Galaxies of ∼40,000 Transients Using DESI

Soumagnac et al. (2024)

KASI: David Parkinson (#45)

We present the Multi-Object Spectroscopy of Transient (MOST) Hosts survey. The survey is planned to run throughout the 5 yr of operation of the Dark Energy Spectroscopic Instrument (DESI) and will generate a spectroscopic catalog of the hosts of most transients observed to date, in particular all the supernovae observed by most public, untargeted, wide-field, optical surveys (Palomar Transient Factory, PTF/intermediate PTF, Sloan Digital Sky Survey II, Zwicky Transient Facility, DECAT, DESIRT). Science cases for the MOST Hosts survey include Type Ia supernova cosmology, fundamental plane and peculiar velocity measurements, and the understanding of the correlations between transients and their host-galaxy properties. Here we present the first release of the MOST Hosts survey: 21,931 hosts of 20,235 transients. These numbers represent 36% of the final MOST Hosts sample, consisting of 60,212 potential host galaxies of 38,603 transients (a transient can be assigned multiple potential hosts). Of all the transients in the MOST Hosts list, only 26.7% have existing classifications, and so the survey will provide redshifts (and luminosities) for nearly 30,000 transients. A preliminary Hubble diagram and a transient luminosity-duration diagram are shown as examples of future potential uses of the MOST Hosts survey. The survey will also provide a training sample of spectroscopically observed transients for classifiers relying only on photometry, as we enter an era when most newly observed transients will lack spectroscopic classification. The MOST Hosts DESI survey data will be released on a rolling cadence and updated to match the DESI releases. Dates of future releases and updates are available through the https://mosthosts.desi.lbl.gov website.

Tier 1transients/time-domainclassical-ML

Inferring Cosmological Parameters on SDSS via Domain-generalized Neural Networks and Light-cone Simulations

Lee et al. (2024)

KASI: Jaehyun Lee (#7)

We present a proof-of-concept simulation-based inference on Ωm and σ8 from the Sloan Digital Sky Survey (SDSS) Baryon Oscillation Spectroscopic Survey (BOSS) LOWZ Northern Galactic Cap (NGC) catalog using neural networks and domain generalization techniques without the need of summary statistics. Using rapid light- cone simulations L-PICOLA, mock galaxy catalogs are produced that fully incorporate the observational effects. The collection of galaxies is fed as input to a point cloud-based network, Minkowski-PointNet. We also add relatively more accurate GADGET mocks to obtain robust and generalizable neural networks. By explicitly learning the representations that reduce the discrepancies between the two different data sets via the semantic alignment loss term, we show that the latent space configuration aligns into a single plane in which the two cosmological parameters form clear axes. Consequently, during inference, the SDSS BOSS LOWZ NGC catalog maps onto the plane, demonstrating effective generalization and improving prediction accuracy compared to non-generalized models. Results from the ensemble of 25 independently trained machines find Ωm=0.339±0.056 and σ8 = 0.801 ± 0.061, inferred only from the distribution of galaxies in the light-cone slices without relying on any indirect summary statistics. A single machine that best adapts to the GADGET mocks yields a tighter prediction of Ωm = 0.282 ± 0.014 and σ8 = 0.786 ± 0.036. We emphasize that adaptation across multiple domains can enhance the robustness of the neural networks in observational data.

Tier 3cosmologyCNNclassical-ML

Detecting unresolved lensed SNe Ia in LSST using blended light curves

Bag et al. (2024)

KASI: Kushal Lodha (#9), 샤피엘루알만 (#13)

Strongly gravitationally lensed supernovae (LSNe) are promising probes for providing absolute distance measurements using gravitational-lens time delays. Spatially unresolved LSNe offer an opportunity to enhance the sample size for precision cosmology. We predict that there will be approximately three times as many unresolved as resolved LSNe Ia in the Legacy Survey of Space and Time (LSST) by the Rubin Observatory. In this article, we explore the feasibility of detecting unresolved LSNe Ia from a pool of preclassified SNe Ia light curves using the shape of the blended light curves with deep-learning techniques. We find that ∼30% unresolved LSNe Ia can be detected with a simple 1D convolutional neural network (CNN) using well-sampled rizy-band light curves (with a false-positive rate of ∼3%). Even when the light curve is well observed in only a single band among r, i, and z, detection is still possible with false-positive rates ranging from ∼4 to 7% depending on the band. Furthermore, we demonstrate that these unresolved cases can be detected at an early stage using light curves up to ∼20 days from the first observation with well-controlled false-positive rates, providing ample opportunity to trigger follow-up observations. Additionally, we demonstrate the feasibility of time-delay estimations using solely LSST-like data of unresolved light curves, particularly for doubles, when excluding systems with low time delays and magnification ratios. However, the abundance of such systems among those unresolved in LSST poses a significant challenge. This approach holds potential utility for upcoming wide-field surveys, and overall results could significantly improve with enhanced cadence and depth in the future surveys.

Tier 2cosmologyCNN

Construction of global IGS-3D electron density (Ne) model by deep learning

Ji et al. (2024)

KASI: Young-Sil Kwak (#3), Jeong-Heon Kim (#5)

In this study, we construct a global IGS-3D Ne model that generates global 3-D electron density (Ne) from International Global Navigation Satellite Systems (GNSS) Service (IGS) total electron content (TEC) data through deep learning. As a first step towards this, we make a model to generate a vertical electron density profile from a TEC value using Multi-Layer Perceptron (MLP). In this process, we use the vertical electron density profiles and the corresponding TEC values of the IRI-2016 model from 2001 to 2008 for training, 2009 and 2014 for validation, and 2010 to 2013 for a test. The next step is to generate global IGS electron density profiles using the global IGS TECs as input data for the model, which is called the global IGS-3D Ne model. We evaluate the IGS-3D Ne model by comparing the electron density profiles from the incoherent scatter radars (ISRs) at three stations with the IGS-3D Ne model from 2010 to 2013. The evaluation shows that the electron density profiles from the IGS-3D Ne model are closer to the ISR data than those of the IRI model, especially at high latitudes. The IGS-3D Ne model shows that the averaged root mean square error (RMSE) values between IGS and ISR electron density profiles are 0.37 log(m- 3), 0.22 log(m- 3), and 0.34 log(m- 3) for all test datasets at Jicamarca, Millstone Hill, and EISCAT stations, respectively. These results suggest that our method has sufficient potential to enhance the ability to predict global electron density profiles.

Tier 1solar/heliophysicsclassical-ML

Systematic KMTNet Planetary Anomaly Search. XI. Complete Sample of 2016 Subprime Field Planets

Shin et al. (2024)

KASI: Lee, Chung-Uk (#7), Chung, Sun-Ju (#11), Kyu-Ha Hwang (#12) +8

Following Shin et al. (2023b), which is a part of the "Systematic KMTNet Planetary Anomaly Search" series (i.e., a search for planets in the 2016 KMTNet prime fields), we conduct a systematic search of the 2016 KMTNet subprime fields using a semi-machine-based algorithm to identify hidden anomalous events missed by the conventional by-eye search. We find four new planets and seven planet candidates that were buried in the KMTNet archive. The new planets are OGLE-2016-BLG-1598Lb, OGLE-2016-BLG-1800Lb, MOA-2016-BLG-526Lb, and KMT-2016-BLG-2321Lb, which show typical properties of microlensing planets, i.e., giant planets orbit M-dwarf host stars beyond their snow lines. For the planet candidates, we find planet/binary or 2L1S/1L2S degeneracies, which are an obstacle to firmly claiming planet detections. By combining the results of Shin et al. (2023b) and this work, we find a total of nine hidden planets, which is about half the number of planets discovered by eye in 2016. With this work, we have met the goal of the systematic search series for 2016, which is to build a complete microlensing planet sample. We also show that our systematic searches significantly contribute to completing the planet sample, especially for planet/host mass ratios smaller than 10-3, which were incomplete in previous by-eye searches of the KMTNet archive.

Tier 1exoplanetsclassical-ML

Local primordial non-Gaussianity from the large-scale clustering of photometric DESI luminous red galaxies

Rezaie et al. (2024)

KASI: Benedict Bahr-Kalus (#15)

We use angular clustering of luminous red galaxies from the Dark Energy Spectroscopic Instrument (DESI) imaging surveys to constrain the local primordial non-Gaussianity parameter fNL. Our sample comprises over 12 million targets, covering 14 000 deg2 of the sky, with redshifts in the range 0.2 < z < 1.35. We identify Galactic extinction, survey depth, and astronomical seeing as the primary sources of systematic error, and employ linear regression and artificial neural networks to alleviate non-cosmological excess clustering on large scales. Our methods are tested against simulations with and without fNL and systematics, showing superior performance of the neural network treatment. The neural network with a set of nine imaging property maps passes our systematic null test criteria, and is chosen as the fiducial treatment. Assuming the universality relation, we find fNL = 34+24(+50)-44(-73) at 68 per cent (95 per cent) confidence. We apply a series of robustness tests (e.g. cuts on imaging, declination, or scales used) that show consistency in the obtained constraints. We study how the regression method biases the measured angular power spectrum and degrades the fNL constraining power. The use of the nine maps more than doubles the uncertainty compared to using only the three primary maps in the regression. Our results thus motivate the development of more efficient methods that avoid overcorrection, protect large-scale clustering information, and preserve constraining power. Additionally, our results encourage further studies of fNL with DESI spectroscopic samples, where the inclusion of 3D clustering modes should help separate imaging systematics and lessen the degradation in the fNL uncertainty.

Tier 2cosmologyclassical-ML

Merger tree-bsasd galaxy matching: A comparative study across different resolution

Jung et al. (2024)

KASI: Sungwook E. Hong (#4), Jaehyun Lee (#5)

We introduce a novel halo/galaxy matching technique between two cosmological simulations with different resolutions, which utilizes the positions and masses of halos along their subhalo merger tree. With this tool, we conduct a study of resolution biases through the galaxy-by-galaxy inspection of a pair of simulations that have the same simulation configuration but different mass resolutions, utilizing a suite of IllustrisTNG simulations to assess the impact on galaxy properties. We find that, with the subgrid physics model calibrated for TNG100-1, subhalos in TNG100-1 (high resolution) have -0.5 dex higher stellar masses than their counterparts in the TNG100-2 (low resolution). It is also discovered that the subhalos with M_gas ∼ 10^8.5 M_sun in TNG100-1 have ∼0.5 dex higher gas mass than those in TNG100-2. The mass profiles of the subhalos reveal that the dark matter masses of subhalos in TNG100-2 converge well with those from TNG100-1, except within 4 kpc of the resolution limit. The differences in stellar mass and hot gas mass are most pronounced in the central region. We exploit machine learning to build a correction mapping for the physical quantities of subhalos from low- to high-resolution simulations (TNG300-1 and TNG100-1), which enables us to find an efficient way to compile a high-resolution galaxy catalog even from a low-resolution simulation. Our tools can easily be applied to other large cosmological simulations, testing and mitigating the resolution biases of their numerical codes and subgrid physics models.

Tier 1galaxiesclassical-ML

Deblurring the early Universe: reconstruction of primordial power spectrum from Planck CMB using image analysis techniques

Sohn et al. (2024)

KASI: Wuhyun Sohn (#1), Arman Shafieloo (#2)

While the simplest inflationary models predict the primordial perturbations to be near scale-invariant, the primordial power spectrum (PPS) can exhibit oscillatory features in many physically well-motivated models. We search for hints of such features via free-form reconstructions of the PPS based on Planck 2018 CMB temperature and polarization anisotropies. In order to robustly invert the oscillatory integrals and handle noisy unbinned data, we draw inspiration from image analysis techniques. In previous works, the Richardson-Lucy deconvolution algorithm for deblurring images has been modified for reconstructing PPS from the CMB temperature angular power spectrum. We extensively develop the methodology by including CMB polarization and introducing two new regularization techniques, also inspired by image analysis and adapted for our cosmological context. Regularization is essential for improving the fit to the temperature and polarization channels (TT, TE and EE) simultaneously without sacrificing one for another. The reconstructions we obtain are consistent with previous findings from temperature-only analyses. We evaluate the statistical significance of the oscillatory features in our reconstructions using mock data and find the observations to be consistent with having a featureless PPS. The machinery developed here will be a complimentary tool in the search for features with upcoming CMB surveys. Our methodology also shows competitive performance in image deconvolution tasks, which have various applications from microscopy to medical imaging.

Tier 1cosmologyclassical-ML

Searching for local features in primordial power spectrum using genetic algorithms

Lodha et al. (2024)

KASI: Kushal Lodha (#1), Arman Shafieloo (#4), Wuhyun Sohn (#5)

We present a novel methodology for exploring local features directly in the primordial power spectrum using a genetic algorithm pipeline coupled with a Boltzmann solver and Cosmic Microwave Background data (CMB). After testing the robustness of our pipeline using mock data, we apply it to the latest CMB data, including Planck 2018 and CamSpec PR4. Our model-independent approach provides an analytical reconstruction of the power spectra that best fits the data, with the unsupervised machine learning algorithm exploring a functional space built off simple ‘grammar’ functions. We find significant improvements upon the simple power-law behaviour, by Δχ2>=21, consistently with more traditional model-based approaches. These best-fits always address both the low-ℓ anomaly in the TT spectrum and the residual high-ℓ oscillations in the TT, TE, and EE spectra. The proposed pipeline provides an adaptable tool for exploring features in the primordial power spectrum in a model-independent way, providing valuable hints to theorists for constructing viable inflationary models that are consistent with the current and upcoming CMB surveys.

Tier 1cosmologyclassical-ML

Deep learning-based solar image captioning

Baek et al. (2024)

KASI: Ji-Hye Baek (#1), Kim Sujin (#2), Seonghwan Choi (#3) +1

Solar images are essential for identifying and predicting solar phenomena, and have been used as key information for analyzing space weather. In this paper, we propose a solar image captioning method that applies a transformer-based deep learning (DL) natural language processing method. In addition, we provide a new DeepSDO description dataset for training solar image captioning models. First, we develop the DeepSDO description dataset using solar image data from Korean Data Center for solar dynamics observatory (SDO) and scripts from the National Aeronautics and Space Administration (NASA) SDO gallery website. The DeepSDO description dataset includes nine solar events: sunspots, flares, prominences, prominent eruptions, coronal holes, coronal loops, filaments, active regions, and eclipses. Second, we train the DL-based image captioning model, the meshed-memory transformer, using the DeepSDO description dataset. The experimental results show that the proposed method outperforms other benchmark methods in terms of four evaluation metrics. This study demonstrates that DL-based image captioning can successfully generate solar image captions for multiple solar features, and could potentially be used in other themes of solar physics and space weather.

Tier 2solar/heliophysicstransformerclassical-ML

Systematic reanalysis of KMTNet microlensing events, paper I: Updates of the photometry pipeline and a new planet candidate

Yang et al. (2024)

KASI: Kyu-Ha Hwang (#3), Chung, Sun-Ju (#12), Kim Seung-Lee (#13) +8

In this work, we update and develop algorithms for KMTNet tender-love care (TLC) photometry in order to create a new, mostly automated, TLC pipeline. We then start a project to systematically apply the new TLC pipeline to the historic KMTNet microlensing events, and search for buried planetary signals. We report the discovery of such a planet candidate in the microlensing event MOA-2019-BLG-421/KMT-2019-BLG-2991. The anomalous signal can be explained by either a planet around the lens star or the orbital motion of the source star. For the planetary interpretation, despite many degenerate solutions, the planet is most likely to be a Jovian planet orbiting an M or K dwarf, which is a typical microlensing planet. The discovery proves that the project can indeed increase the sensitivity of historic events and find previously undiscovered signals.

Tier 1exoplanetsclassical-ML

Extended ionized Fe objects in the UWIFE survey

Kim et al. (2024)

KASI: Yesol Kim (#1), Jeong Woong-Seob (#5), Lee, Jae-Joon (#6) +1

We explore systematically the shocked gas in the first Galactic quadrant of the Milky Way using the United Kingdom Infrared Telescope (UKIRT) Wide-field Infrared Survey for Fe+ (UWIFE). The UWIFE survey is the first imaging survey of the Milky Way in the [Fe II] 1.644 μm emission line and covers the Galactic plane in the first Galactic quadrant (7° < l < 62°; |b| - 1 -.5). We identify 204 extended ionized Fe objects (IFOs) using a combination of a manual and automatic search. Most of the IFOs are detected for the first time in the [Fe II] 1.644 μm line. We present a catalogue of the measured sizes and fluxes of the IFOs and searched for their counterparts by performing positional cross-matching with known sources. We found that IFOs are associated with supernova remnants (25), young stellar objects (100), H II regions (33), planetary nebulae (17), and luminous blue variables (4). The statistical and morphological properties are discussed for each of these.

Tier 1ISMclassical-ML

Gravitational-wave Electromagnetic Counterpart Korean Observatory (GECKO): GECKO Follow-up Observation of GW190425

Paek et al. (2024)

KASI: Changsu Choi (#6), Lee, Chung-Uk (#14), Kim Seung-Lee (#15) +1

One of the keys to the success of multimessenger astronomy is the rapid identification of the electromagnetic wave counterpart, kilonova (KN), of the gravitational-wave (GW) event. Despite its importance, it is hard to find a KN associated with a GW event, due to a poorly constrained GW localization map and numerous signals that could be confused as a KN. Here, we present the Gravitational-wave Electromagnetic wave Counterpart Korean Observatory (GECKO) project, the GECKO observation of GW190425, and prospects of GECKO in the fourth observing run (O4) of the GW detectors. We outline our follow-up observation strategies during O3. In particular, we describe our galaxy-targeted observation criteria that prioritize based on galaxy properties. Armed with this strategy, we performed an optical and/or near-infrared follow-up observation of GW190425, the first binary neutron star merger event during the O3 run. Despite a vast localization area of 7460 deg2, we observed 621 host galaxy candidates, corresponding to 29.5% of the scores we assigned, with most of them observed within the first 3 days of the GW event. Ten transients were discovered during this search, including a new transient with a host galaxy. No plausible KN was found, but we were still able to constrain the properties of potential KNe using upper limits. The GECKO observation demonstrates that GECKO can possibly uncover a GW170817-like KN at a distance <200 Mpc if the localization area is of the order of hundreds of square degrees, providing a bright prospect for the identification of GW electromagnetic wave counterparts during the O4 run.

Tier 1transients/time-domainclassical-ML

Deep Learning-Based Regional Ionospheric Total Electron Content Prediction-Long Short-Term Memory (LSTM) and Convolutional LSTM Approach

Jeong et al. (2024)

KASI: Se-Heon Jeong (#1), Woo Kyoung Lee (#2), Jeong-Heon Kim (#5) +1

This study evaluates the performance of deep learning approach in the prediction of the ionospheric total electron content (TEC) during magnetically quiet periods. Two deep learning techniques, long short-term memory (LSTM) and convolutional LSTM (ConvLSTM), are employed to predict TEC values 24 hr ahead in the vicinity of the Korean Peninsula (26.5°-40°N, 121°-134.5°E). The LSTM method predicts TEC at a single point based on time series of data at that point, whereas the ConvLSTM method simultaneously predicts TEC values at multiple points using spatiotemporal distribution of TEC. Both the LSTM and ConvLSTM models are trained using the complete regional TEC maps reconstructed by applying the Deep Convolutional Generative Adversarial Network-Poisson Blending (DCGAN-PB) method to observed TEC data. The training period spans from 2002 to 2018, and the model performance is evaluated using 2019 data. Our results show that the ConvLSTM method outperforms the LSTM method, generating more reliable TEC maps with smaller root mean square errors when compared to the ground truth (DCGAN-PB TEC maps). This outcome indicates that deep learning models can improve the prediction accuracy of TEC at a specific point by taking into account spatial information of TEC. We conclude that ConvLSTM is a reliable and efficient approach for the prompt ionospheric prediction.

Tier 2solar/heliophysicsRNN/LSTMCNN

A Model-independent Method to Determine H0 Using Time-delay Lensing, Quasars, and Type Ia Supernovae

Li et al. (2024)

KASI: 샤피엘루알만 (#3)

Absolute distances from strong lensing can anchor Type Ia Supernovae (SNe Ia) at cosmological distances giving a model-independent inference of the Hubble constant (H0). Future observations could provide strong lensing time- delay distances with source redshifts up to z ; 4, which are much higher than the maximum redshift of SNe Ia observed so far. In order to make full use of time-delay distances measured at higher redshifts, we use quasars as a complementary cosmic probe to measure cosmological distances at redshifts beyond those of SNe Ia and provide a model-independent method to determine H0. In this work, we demonstrate a model-independent, joint constraint of SNe Ia, quasars, and time-delay distances from strong lensed quasars. We first generate mock data sets of SNe Ia, quasar, and time-delay distances based on a fiducial cosmological model. Then, we calibrate the quasar parameters model independently using Gaussian process (GP) regression with mock SNe Ia data. Finally, we determine the value of H0 model-independently using GP regression from mock quasars and time-delay distances from strong lensing systems. As a comparison, we also show the H0 results obtained from mock SNe Ia in combination with time-delay lensing systems whose redshifts overlap with SNe Ia. Our results show that quasars at higher redshifts show great potential to extend the redshift coverage of SNe Ia and thus enable the full use of strong lens time-delay distance measurements from ongoing cosmic surveys and improve the accuracy of the estimation of H0 from 2.1% to 1.3% when the uncertainties of the time-delay distances are 5% of the distance values.

Tier 1cosmologygaussian-process

2023 (26 papers)

The Principal Component Analysis Filtering Method for an Unbiased Spectral Survey of Complex Organic Molecules

Yun & Lee (2023)

KASI: Hyeong-Sik Yun (#1)

A variety of interstellar complex organic molecules (COMs) have been detected in various physical conditions. However, in the protostellar and protoplanetary environments, their complex kinematics make line profiles blend together and the line strength of weak lines weaker. In this paper, we utilize the principal component analysis technique to develop a filtering method that can extract COM spectra from the main kinematic component associated with COM emission and increase the signal-to-noise ratio (S/N) of spectra. This filtering method corrects non-Gaussian line profiles caused by the kinematics. For this development, we adopt the ALMA BAND 6 spectral survey data of V883 Ori, an eruptive young star with a Keplerian disk. A filter was, first, created using 34 strong and well-isolated COM lines and then applied to the entire spectral range of the data set. The first principal component (PC1) describes the most common emission structure of the selected lines, which is confined within the water sublimation radius (~0.″3) in the Keplerian disk of V883 Ori. Using this PC1 filter, we extracted high-S/N kinematics-corrected spectra of V883 Ori over the entire spectral coverage of ~50 GHz. The PC1-filtering method reduces the noise by a factor of ~2 compared to the average spectra over the COM emission region. One important advantage of this PC1-filtering method over the previously developed matched-filtering method is the ability to preserve the original integrated intensities of COM lines.

Tier 1starsclassical-ML

Probabilistic classification of infrared-selected targets for SPHEREx mission: in search of young stellar objects

Lakshmipathaiah et al. (2023)

KASI: Miju Kang (#5)

We apply machine learning algorithms to classify Infrared (IR)-selected targets for NASA's upcoming SPHEREx mission. In particular, we are interested in classifying Young Stellar Objects (YSOs), which are essential for understanding the star formation process. Our approach differs from previous work, which has relied heavily on broadband color criteria to classify IR-bright objects, and are typically implemented in color-color and color-magnitude diagrams. However, these methods do not state the confidence associated with the classification and the results from these methods are quite ambiguous due to the overlap of different source types in these diagrams. Here, we utilize photometric colors and magnitudes from seven near and mid-infrared bands simultaneously and employ machine and deep learning algorithms to carry out probabilistic classification of YSOs, Asymptotic Giant Branch (AGB) stars, Active Galactic Nuclei (AGN) and main-sequence (MS) stars. Our approach also sub-classifies YSOs into Class I, II, III and flat spectrum YSOs, and AGB stars into carbon-rich and oxygen-rich AGB stars. We apply our methods to infrared-selected targets compiled in preparation for SPHEREx which are likely to include YSOs and other classes of objects. Our classification indicates that out of 8,308,384 sources, 1,966,340 have class prediction with probability exceeding 90%, amongst which ∼1.7% are YSOs, ∼58.2% are AGB stars, ∼40% are (reddened) MS stars, and ∼0.1% are AGN whose red broadband colors mimic YSOs. We validate our classification using the spatial distributions of predicted YSOs towards the Cygnus-X star-forming complex, as well as AGB stars across the Galactic plane.

Tier 2starsclassical-ML

The universe is worth 64^3 pixels: convolution neural network and vision transformers for cosmology

Hwang et al. (2023)

KASI: Sungwook E. Hong (#4)

We present a novel approach for estimating cosmological parameters, Ωm, σ8, w0, and one derived parameter, S8, from 3D lightcone data of dark matter halos in redshift space covering a sky area of 40- × 40- and redshift range of 0.3 < z < 0.8, binned to 643 voxels. Using two deep learning algorithms - Convolutional Neural Network (CNN) and Vision Transformer (ViT) - we compare their performance with the standard two-point correlation (2pcf) function. Our results indicate that CNN yields the best performance, while ViT also demonstrates significant potential in predicting cosmological parameters. By combining the outcomes of Vision Transformer, Convolution Neural Network, and 2pcf, we achieved a substantial reduction in error compared to the 2pcf alone. To better understand the inner workings of the machine learning algorithms, we employed the Grad-CAM method to investigate the sources of essential information in heatmaps of the CNN and ViT. Our findings suggest that the algorithms focus on different parts of the density field and redshift depending on which parameter they are predicting. This proof-of-concept work paves the way for incorporating deep learning methods to estimate cosmological parameters from large-scale structures, potentially leading to tighter constraints and improved understanding of the Universe.

Tier 2cosmologyCNNtransformer

Near-IR Weak-lensing (NIRWL) Measurements in the CANDELS Fields. I. Point-spread Function Modeling and Systematics

Finner et al. (2023)

KASI: Bomee Lee (#2)

We have undertaken a near-IR weak-lensing (NIRWL) analysis of the CANDELS HST/WFC3-IR F160W observations. With the Gaia proper motion-corrected catalog as an astrometric reference, we updated the astrometry of the five CANDELS mosaics and achieved an absolute alignment within 0farcs02 ± 0farcs02, on average, which is a factor of several superior to existing mosaics. These mosaics are available to download (https://drive.google.com/drive/folders/1k9WEV3tBOuRKBlcaTJ0-wTZnUCisS__r). We investigated the systematic effects that need to be corrected for weak-lensing measurements. We find that the largest contributing systematic effect is caused by undersampling. We find a subpixel centroid dependence on the PSF shape that causes the PSF ellipticity and size to vary by up to 0.02% and 3%, respectively. Using the UDS as an example field, we show that undersampling induces a multiplicative shear bias of -0.025. We find that the brighter-fatter effect causes a 2% increase in the size of the PSF and discover a brighter-rounder effect that changes the ellipticity by 0.006. Based on the small range of slopes in a galaxy's spectral energy distribution (SED) within the WFC3-IR bandpasses, we suggest that the impact of the galaxy SED on the PSF is minor. Finally, we model the PSF of WFC3-IR F160W for weak lensing using a principal component analysis. The PSF models account for temporal and spatial variations of the PSF. The PSF corrections result in residual ellipticities and sizes, -de1- < 0.0005 ± 0.0003, -de2- < 0.0005 ± 0.0003, and -dR- < 0.0005 ± 0.0001, that are sufficient for the upcoming NIRWL search for massive overdensities in the five CANDELS fields.

Tier 1galaxiesclassical-ML

A data compression and optimal galaxy weights scheme for Dark Energy Spectroscopic Instrument and weak lensing data sets

Ruggeri et al. (2023)

KASI: Christoph Saulder (#14)

Combining different observational probes, such as galaxy clustering and weak lensing, is a promising technique for unveiling the physics of the Universe with upcoming dark energy experiments. The galaxy redshift sample from the Dark Energy Spectroscopic Instrument (DESI) will have a significant overlap with major ongoing imaging surveys specifically designed for weak lensing measurements: the Kilo-Degree Survey (KiDS), the Dark Energy Survey (DES), and the Hyper Suprime-Cam (HSC) survey. In this work, we analyse simulated redshift and lensing catalogues to establish a new strategy for combining high-quality cosmological imaging and spectroscopic data, in view of the first-year data assembly analysis of DESI. In a test case fitting for a reduced parameter set, we employ an optimal data compression scheme able to identify those aspects of the data that are most sensitive to cosmological information and amplify them with respect to other aspects of the data. We find this optimal compression approach is able to preserve all the information related to the growth of structures.

Tier 1cosmologyclassical-ML

Automatic Detection of Type II Solar Radio Burst by Using 1-D Convolution Neutral Network

Cho et al. (2023)

KASI: CHO KYUNG SUK (#1), Rok-Soon Kim (#3), Eunsu Park (#4)

Type II solar radio bursts show frequency drifts from high to low over time. They have been known as a signature of coronal shock associated with Coronal Mass Ejections (CMEs) and/or flares, which cause an abrupt change in the space environment near the Earth (space weather). Therefore, early detection of type II bursts is important for forecasting of space weather. In this study, we develop a deep-learning (DL) model for the automatic detection of type II bursts. For this purpose, we adopted a 1-D Convolution Neutral Network (CNN) as it is well-suited for processing spatiotemporal information within the applied data set. We utilized a total of 286 radio burst spectrum images obtained by Hiraiso Radio Spectrograph (HiRAS) from 1991 and 2012, along with 231 spectrum images without the bursts from 2009 to 2015, to recognizes type II bursts. The burst types were labeled manually according to their spectra features in an answer table. Subsequently, we applied the 1-D CNN technique to the spectrum images using two filter windows with different size along time axis. To develop the DL model, we randomly selected 412 spectrum images (80%) for training and validation. The train history shows that both train and validation losses drop rapidly, while train and validation accuracies increased within approximately 100 epoches. For evaluation of the model’s performance, we used 105 test images (20%) and employed a contingence table. It is found that false alarm ratio (FAR) and critical success index (CSI) were 0.14 and 0.83, respectively. Furthermore, we confirmed above result by adopting five-fold cross-validation method, in which we re-sampled five groups randomly. The estimated mean FAR and CSI of the five groups were 0.05 and 0.87, respectively. For experimental purposes, we applied our proposed model to 85 HiRAS type II radio bursts listed in the NGDC catalogue from 2009 to 2016 and 184 quiet (no bursts) spectrum images before and after the type II bursts. As a result, our model successfully detected 79 events (93%) of type II events. This results demonstrates, for the first time, that the 1-D CNN algorithm is useful for detecting type II bursts.

Tier 2solar/heliophysicsCNN

The SN 2023ixf Progenitor in M101. I. Infrared Variability

Soraisam et al. (2023)

KASI: Sang-Hyun Chun (#6)

Observational evidence points to a red supergiant (RSG) progenitor for SN 2023ixf. The progenitor candidate has been detected in archival images at wavelengths (-0.6 μm) where RSGs typically emit profusely. This object is distinctly variable in the infrared (IR). We characterize the variability using pre-explosion mid-IR (3.6 and 4.5 μm) Spitzer and ground-based near-IR (JHKs) archival data jointly covering 19yr. The IR light curves exhibit significant variability with rms amplitudes in the range 0.2-0.4 mag, increasing with decreasing wavelength. From a robust period analysis of the more densely sampled Spitzer data, we measure a period of 1091 ± 71 days. We demonstrate using Gaussian process modeling that this periodicity is also present in the near-IR light curves, thus indicating a common physical origin, which is likely pulsational instability. We use a period-luminosity relation for RSGs to derive a value of MK = -11.58 ± 0.31 mag. Assuming a late M spectral type, this corresponds to log(L L-) = 5.27 - 0.12 at Teff =3200 K and to log(L L-) = 5.37 - 0.12 at Teff =3500 K. This gives an independent estimate of the progenitor’s luminosity, unaffected by uncertainties in extinction and distance. Assuming the progenitor candidate underwent enhanced dust-driven mass loss during the time of these archival observations, and using an empirical period-luminosity-based mass-loss prescription, we obtain a mass-loss rate of around (2-4) × 10-4 Me yr-1. Comparing the above luminosity with stellar evolution models, we infer an initial mass for the progenitor candidate of 20 ± 4 Me, making this one of the most massive progenitors for a Type II SN detected to date.

Tier 1starsgaussian-process

Flow of gas detected from beyond the filaments to protostellar scales in Barnard 5

Valdivia-Mena et al. (2023)

KASI: Spandan Choudhury (#6)

Context. The infall of gas from outside natal cores has proven to feed protostars after the main accretion phase (Class 0). This changes our view of star formation to a picture that includes asymmetric accretion (streamers), and a larger role of the environment. However, the connection between streamers and the filaments that prevail in star-forming regions is unknown. Aims: We investigate the flow of material toward the filaments within Barnard 5 (B5) and the infall from the envelope to the protostellar disk of the embedded protostar B5-IRS1. Our goal is to follow the flow of material from the larger, dense core scale, to the protostellar disk scale. Methods: We present new HC3N line data from the NOEMA and 30 m telescopes covering the coherence zone of B5, together with ALMA H2CO and C18O maps toward the protostellar envelope. We fit multiple Gaussian components to the lines so as to decompose their individual physical components. We investigated the HC3N velocity gradients to determine the direction of chemically fresh gas flow. At envelope scales, we used a clustering algorithm to disentangle the different kinematic components within H2CO emission. Results: At dense core scales, HC3N traces the infall from the B5 region toward the filaments. HC3N velocity gradients are consistent with accretion toward the filament spines plus flow along them. We found a ~2800 au streamer in H2CO emission, which is blueshifted with respect to the protostar and deposits gas at outer disk scales. The strongest velocity gradients at large scales curve toward the position of the streamer at small scales, suggesting a connection between both flows. Conclusions: Our analysis suggests that the gas can flow from the dense core to the protostar. This implies that the mass available for a protostar is not limited to its envelope, and it can receive chemically unprocessed gas after the main accretion phase.

Tier 1starsclassical-ML

On the consistency of LCDM with CMB measurements in light of the latest Planck, ACT and SPT data

Calderon et al. (2023)

KASI: Rodrigo Calderon (#1), Arman Shafieloo (#2), Wuhyun Sohn (#4)

Using Gaussian Processes we perform a thorough, non-parametric consistency test of the $\Lambda$CDM model when confronted with state-of-the-art TT, TE, and EE measurements of the anisotropies in the Cosmic Microwave Background by the Planck, ACT, and SPT collaborations. Using $\Lambda$CDM's best-fit predictions to the TTTEEE data from Planck, we find no statistically significant deviations when looking for signatures in the residuals across the different datasets. The results of SPT are in good agreement with the $\Lambda$CDM best-fit predictions to the Planck data, while the results of ACT are only marginally consistent. However, when using the best-fit predictions to CamSpec -- a recent reanalysis of the Planck data -- as the mean function, we find larger discrepancies between the datasets. Our analysis also reveals an interesting feature in the polarisation (EE) measurements from the CamSpec analysis, which could be explained by a slight underestimation of the covariance matrix. Interestingly, the disagreement between CamSpec and Planck/ACT is mainly visible in the residuals of the TT spectrum, the latter favoring a scale-invariant tilt $n_s\simeq1$, which is consistent with previous findings from parametric analyses. We also report some features in the EE measurements captured both by ACT and SPT which are independent of the chosen mean function and could be hinting towards a common physical origin. For completeness, we repeat our analysis using the best-fit spectra to ACT+WMAP as the mean function. Finally, we test the internal consistency of the Planck data alone by studying the high and low-$\ell$ ranges separately, finding no discrepancy between small and large angular scales.

Tier 1cosmologygaussian-process

The Kinematics of the Young Stellar Population in the W5 Region of the Cassiopeia OB6 Association: Implication for the Formation Process of Stellar Associations

Lim et al. (2023)

KASI: 임범두 (#1), 홍종석 (#2), 이진희 (#3) +3

The star-forming region W5 is a major part of the Cassiopeia OB6 association. Its internal structure and kinematics may provide hints of the star formation process in this region. Here, we present a kinematic study of young stars in W5 using the Gaia data and our radial velocity data. A total 490 out of 2000 young stars are confirmed as members. Their spatial distribution shows that W5 is highly substructured. We identify a total of eight groups using the k-means clustering algorithm. There are three dense groups in the cavities of H ii bubbles, and the other five sparse groups are distributed at the edges of the bubbles. The three dense groups have almost the same age (5 Myr) and show a pattern of expansion. The scale of their expansion is not large enough to account for the overall structure of W5. The three northern groups are, in fact, 3 Myr younger than the dense groups, which indicates independent star formation events. Only one of these groups shows the signature of feedback-driven star formation as its members move away from the eastern dense group. The other two groups might have formed in a spontaneous way. On the other hand, the properties of two southern groups are not understood as those of a coeval population. Their origins can be explained by dynamical ejection of stars and multiple star formation. Our results suggest that the substructures in W5 formed through multiple star-forming events in a giant molecular cloud.

Tier 1starsclassical-ML

Joint reconstructions of growth and expansion histories from stage-IV surveys with minimal assumptions. II. Modified gravity and massive neutrinos

Calderón et al. (2023)

KASI: Rodrigo Calderon (#1), Arman Shafieloo (#4)

Based on a formalism introduced in our previous work, we reconstruct the phenomenological function Geff(z) describing deviations from general relativity (GR) in a model-independent manner. In this alternative approach, we model & mu; & EQUIV; Geff/G as a Gaussian process and use forecasted growth-rate measurements from a stage-IV survey to reconstruct its shape for two different toy-models. We follow a two-step procedure: (i) we first reconstruct the background expansion history from supernovae (SNe) and baryon acoustic oscillation (BAO) measurements; (ii) we then use it to obtain the growth history f & sigma;8, that we fit to redshift-space distortions (RSD) measurements to reconstruct Geff. We find that upcoming surveys such as the dark energy spectroscopic instrument (DESI) might be capable of detecting deviations from GR, provided the dark energy behavior is accurately determined. We might even be able to constrain the transition redshift from G & RARR; Geff for some particular models. We further assess the impact of massive neutrinos on the reconstructions of Geff (or & mu;) assuming the expansion history is given, and only the neutrino mass is free to vary. Given the tight constraints on the neutrino mass, and for the profiles we considered in this work, we recover numerically that the effect of such massive neutrinos do not alter our conclusions. Finally, we stress that incorrectly assuming a ACDM expansion history leads to a degraded reconstruction of & mu;, and/or a non-negligible bias in the (& omega;m;0; & sigma;8;0)-plane.

Tier 1cosmologygaussian-process

Reconstructing the cosmological density and velocity fields from redshifted galaxy distributions using V-net

Qin et al. (2023)

KASI: Fei Qin (#1), David Parkinson (#2), Sungwook E. Hong (#3)

The distribution of matter that is measured through galaxy redshift and peculiar velocity surveys can be harnessed to learn about the physics of dark matter, dark energy, and the nature of gravity. To improve our understanding of the matter of the Universe, we can reconstruct the full density and velocity fields from the galaxies that act as tracer particles. In this paper, we use the simulated halos as proxies for the galaxies. We use a convolutional neural network, a V-net, trained on numerical simulations of structure formation to reconstruct the density and velocity fields. We find that, with detailed tuning of the loss function, the V-net could produce better fits to the density field in the high-density and low-density regions, and improved predictions for the probability distribution of the amplitudes of the velocities. However, the weights will reduce the precision of the estimated β parameter. We also find that the redshift-space distortions of the halo catalogue do not significantly contaminate the reconstructed real-space density and velocity field. We estimate the velocity field β parameter by comparing the peculiar velocities of halo catalogues to the reconstructed velocity fields, and find the estimated β values agree with the fiducial value at the 68% confidence level.

Tier 2cosmologyCNN

Systematic KMTNet Planetary Anomaly Search. VIII. Complete Sample of 2019 Subprime Field Planets

Jung et al. (2023)

KASI: Jung Youn Kil (#1), Chung, Sun-Ju (#8), Kyu-Ha Hwang (#9) +8

We complete the publication of all microlensing planets (and ``possible planets'') identified by the uniform approach of the KMT AnomalyFinder system in the 21 KMT subprime fields during the 2019 observing season, namely, KMT-2019-BLG-0298, KMT-2019-BLG-1216, KMT-2019-BLG-2783, OGLE-2019-BLG-0249, and OGLE-2019-BLG-0679 (planets), as well as OGLE-2019-BLG-0344, and KMT-2019-BLG-0304 (possible planets). The five planets have mean log mass-ratio measurements of $(-2.6,-3.6,-2.5,-2.2,-2.3)$, median mass estimates of $(1.81,0.094,1.16,7.12,3.34)\, M_{\rm Jup}$, and median distance estimates of $(6.7,2.7,5.9,6.4,5.6)\,\kpc$, respectively. The main scientific interest of these planets is that they complete the AnomalyFinder sample for 2019, which has a total of 25 planets that are likely to enter the statistical sample. We find statistical consistency with the previously published 33 planets from the 2018 AnomalyFinder analysis according to an ensemble of five tests. Of the 58 planets from 2018-2019, 23 were newly discovered by AnomalyFinder. Within statistical precision, half of all the planets have caustic crossings while half do not; an equal number of detected planets result from major-image and minor-image light-curve perturbations, and an equal number come from KMT prime fields versus subprime fields.

Tier 1exoplanetsclassical-ML

Variation of ionosonde foE and its comparison with IRI-2016 & a local model over Pakistan longitude sector during solar cycle 22

Ameen et al. (2023)

KASI: Madeeha Talha (#5)

The study presents the variation of ionospheric critical frequency of E-layer (f oE), its comparison with International Reference Ionosphere (IRI-2016) over Pakistan, and suggests an artificial neural network (ANN) based algorithm to predict noontime daily hourly values. The ionosonde measurements of Karachi (Geog. Coord: 24.95N, 67.13E, dip latitude = 17.0N), Multan (30.18N, 71.48E, 21.8N), and Islamabad (33.75N, 73.13E, 25.2N) have been used for the purpose. The observed f oE values have been analyzed during the high, moderate and low solar activity years 1989, 1992, and 1996, respectively. Results show a strong correlation between f oE and solar activity which was expected because the E-layer is essentially solar controlled. It is found that f oE peaks around noon with higher values during summer. The annual relative percentage deviation between data and IRI-2016 modeled values remains less than 6% for all locations and years under study. The ANN-based algorithm predicts daily noontime f oE values with a goodness of fit of 78%, compared to the IRI-2016 prediction, which has a goodness of fit of 71%. It is noted that the performance of IRI-2016 in predicting monthly hourly median values of f oE improves with decreasing solar activity. The study confirms the data authenticity of f oE over Pakistan and suggests an ANN-based algorithm in predicting noontime values with no immediate updates to IRI-2016 for this Asian longitude sector.

Tier 1solar/heliophysicsclassical-ML

Identifying anomalous radio sources in the Evolutionary Map of the Universe Pilot Survey using a complexity-based approach

Segal et al. (2023)

KASI: David Parkinson (#2)

The Evolutionary Map of the Universe (EMU) large-area radio continuum survey will detect tens of millions of radio galaxies, giving an opportunity for the detection of previously unknown classes of objects. To maximize the scientific value and make new discoveries, the analysis of these data will need to go beyond simple visual inspection. We propose the coarse-grained complexity, a simple scalar quantity relating to the minimum description length of an image that can be used to identify unusual structures. The complexity can be computed without reference to the broader sample or existing catalogue data, making the computation efficient on new surveys at very large scales (such as the full EMU survey). We apply our coarse-grained complexity measure to data from the EMU Pilot Survey to detect and confirm anomalous objects in this data set and produce an anomaly catalogue. Rather than work with existing catalogue data using a specific source detection algorithm, we perform a blind scan of the area, computing the complexity using a sliding square aperture. The effectiveness of the complexity measure for identifying anomalous objects is evaluated using crowd-sourced labels generated via the Zooniverse.org platform. We find that the complexity scan identifies unusual sources, such as odd radio circles, by partitioning on complexity. We achieve partitions where 5 per cent of the data is estimated to be 86 per cent complete, and 0.5 per cent is estimated to be 94 per cent pure, with respect to anomalies and use this to produce an anomaly catalogue.

Tier 1galaxiesclassical-ML

Identification of tidal features in deep optical galaxy images with Convolutional Neural Networks

Domínguez Sánchez et al. (2023)

KASI: Garreth William Martin (#2)

Interactions between galaxies leave distinguishable imprints in the form of tidal features, which hold important clues about their mass assembly. Unfortunately, these structures are difficult to detect because they are low surface brightness features, so deep observations are needed. Upcoming surveys promise several orders of magnitude increase in depth and sky coverage, for which automated methods for tidal feature detection will become mandatory. We test the ability of a convolutional neural network to reproduce human visual classifications for tidal detections. We use as training ∼6000 simulated images classified by professional astronomers. The mock Hyper Suprime Cam Subaru (HSC) images include variations with redshift, projection angle, and surface brightness (μlim = 26-35-mag-arcsec-2). We obtain satisfactory results with accuracy, precision, and recall values of Acc = 0.84, P = 0.72, and R = 0.85 for the test sample. While the accuracy and precision values are roughly constant for all surface brightness, the recall (completeness) is significantly affected by image depth. The recovery rate shows strong dependence on the type of tidal features: we recover all the images showing shell features and 87 per-cent of the tidal streams; these fractions are below 75 per-cent for mergers, tidal tails, and bridges. When applied to real HSC images, the performance of the model worsens significantly. We speculate that this is due to the lack of realism of the simulations, and take it as a warning on applying deep learning models to different data domains without prior testing on the actual data.

Tier 2galaxiesCNN

Fast Reconstruction of 3D Density Distribution around the Sun Based on the MAS by Deep Learning

Rahman et al. (2023)

KASI: Eunsu Park (#6)

This study is the first attempt to generate a three-dimensional (3D) coronal electron density distribution based on the pix2pixHD model, whose computing time is much shorter than that of the magnetohydrodynamic (MHD) simulation. For this, we consider photospheric solar magnetic fields as input, and electron density distribution simulated with the MHD Algorithm outside a Sphere (MAS) at a given solar radius is taken as output. We consider 155 pairs of Carrington rotations as inputs and outputs from 2010 June to 2022 April for training and testing. We train 152 deep-learning models for 152 solar radii, which are taken up to 30 solar radii. The artificial intelligence (AI) generated 3D electron densities from this study are quite consistent with the simulated ones from lower radii to higher radii, with an average correlation coefficient 0.97. The computing time of testing data sets up to 30 solar radii of 152 deep-learning models is about 45.2 s using the NVIDIA TITAN XP graphics-processing unit, which is much less than the typical simulation time of MAS. We find that the synthetic coronagraphic images estimated from the deep-learning models are similar to the Solar Heliospheric Observatory (SOHO)/Large Angle and Spectroscopic Coronagraph C3 coronagraph data, especially during the solar minimum period. The AI-generated coronal density distribution from this study can be used for space weather models on a near-real-time basis.

Tier 3solar/heliophysicsCNN

Swarm-intelligence-based extraction and manifold crawling along the Large-Scale Structure

Awad et al. (2023)

KASI: Jihye Shin (#7)

The distribution of galaxies and clusters of galaxies on the mega-parsec scale of the Universe follows an intricate pattern now famously known as the Large-Scale Structure or the Cosmic Web. To study the environments of this network, several techniques have been developed that are able to describe its properties and the properties of groups of galaxies as a function of their environment. In this work, we analyse the previously introduced framework: 1-Dimensional Recovery, Extraction, and Analysis of Manifolds (1-DREAM) on N-body cosmological simulation data of the Cosmic Web. The 1-DREAM toolbox consists of five Machine Learning methods, whose aim is the extraction and modelling of one-dimensional structures in astronomical big data settings. We show that 1-DREAM can be used to extract structures of different density ranges within the Cosmic Web and to create probabilistic models of them. For demonstration, we construct a probabilistic model of an extracted filament and move through the structure to measure properties such as local density and velocity. We also compare our toolbox with a collection of methodologies which trace the Cosmic Web. We show that 1-DREAM is able to split the network into its various environments with results comparable to the state-of-the-art methodologies. A detailed comparison is then made with the public code DISPERSE, in which we find that 1-DREAM is robust against changes in sample size making it suitable for analysing sparse observational data, and finding faint and diffuse manifolds in low-density regions.

Tier 1cosmologyclassical-ML

Taxonomic Classification of Asteroids Using the KMTNet Multiband Photometry Data Set

Choi et al. (2023)

KASI: MOON, HONG KYU (#2), Dong-Goo Roh (#3), Min-Su Shin (#4) +1

We report the multiband photometry of asteroids observed over 14 nights from 2015 December to 2017 December using the Korea Microlensing Telescope Network telescopes with the taxonomic classification of those objects. The data set contains the photometry of 6793 asteroids in the Sloan Digital Sky Survey griz bands. Following the method of DeMeo & Carry, we define classification criteria on the 2D color plane to assign nine taxonomic types (A, B, C, K, L&D, O, S, V, and X) for the observed objects. We also determine asteroid taxonomy in the newly defined 3D color space as suggested by Roh et al. with seven distinct types based on their novel semisupervised machine-learning model. Both methods distinguish between the S type and others but have difficulty separating the X and C types due to their weak and indistinguishable features and broad distribution in the color spaces. The heliocentric distribution of the observed asteroids with their taxonomic assignments confirms similar trends in the previous works; the number of S types decreases, while the fraction of C types increases with the heliocentric distance in the main belt. On the other hand, the D type dominates in the Jupiter Trojans.

Tier 1otherclassical-ML

JCMT BISTRO Observations: Magnetic Field Morphology of Bubbles Associated with NGC 6334

Tahani et al. (2023)

KASI: Thiem Hoang (#25), 황지혜 (#26), 강지현 (#27) +17

We study the Hii regions associated with the NGC 6334 molecular cloud observed in the submillimeter and taken as part of the B-fields In STar-forming Region Observations Survey. In particular, we investigate the polarization patterns and magnetic field morphologies associated with these Hii regions. Through polarization pattern and pressure calculation analyses, several of these bubbles indicate that the gas and magnetic field lines have been pushed away from the bubble, toward an almost tangential (to the bubble) magnetic field morphology. In the densest part of NGC 6334, where the magnetic field morphology is similar to an hourglass, the polarization observations do not exhibit observable impact from Hii regions. We detect two nested radial polarization patterns in a bubble to the south of NGC 6334 that correspond to the previously observed bipolar structure in this bubble. Finally, using the results of this study, we present steps (incorporating computer vision; circular Hough transform) that can be used in future studies to identify bubbles that have physically impacted magnetic field lines.

Tier 1ISMclassical-ML

Systematic KMTNet Planetary Anomaly Search. VII. Complete Sample of q < 10-4 Planets from the First 4 yr Survey

Zang et al. (2023)

KASI: Jung Youn Kil (#2), Chung, Sun-Ju (#10), Kyu-Ha Hwang (#12) +8

We present the analysis of seven microlensing planetary events with planet/host mass ratios q < 10-4: KMT-2017-BLG-1194, KMT-2017-BLG-0428, KMT-2019-BLG-1806, KMT-2017-BLG-1003, KMT-2019-BLG-1367, OGLE-2017-BLG-1806, and KMT-2016-BLG-1105. They were identified by applying the Korea Microlensing Telescope Network (KMTNet) AnomalyFinder algorithm to 2016-2019 KMTNet events. A Bayesian analysis indicates that all the lens systems consist of a cold super-Earth orbiting an M or K dwarf. Together with 17 previously published and three that will be published elsewhere, AnomalyFinder has found a total of 27 planets that have solutions with q < 10-4 from 2016-2019 KMTNet events, which lays the foundation for the first statistical analysis of the planetary mass-ratio function based on KMTNet data. By reviewing the 27 planets, we find that the missing planetary caustics problem in the KMTNet planetary sample has been solved by AnomalyFinder. We also find a desert of high-magnification planetary signals (A >~65), and a follow-up project for KMTNet high-magnification events could detect at least two more q < 10 -4 planets per year and form an independent statistical sample.

Tier 1exoplanetsclassical-ML

How to use GP: effects of the mean function and hyperparameter selection on Gaussian process regression

Hwang et al. (2023)

KASI: Ryan Keeley (#3), Arman Shafieloo (#5)

Gaussian processes have been widely used in cosmology to reconstruct cosmological quantities in a model-independent way. However, the validity of the adopted mean function and hyperparameters, and the dependence of the results on the choice have not been well explored. In this paper, we study the effects of the underlying mean function and the hyperparameter selection on the reconstruction of the distance moduli from type Ia supernovae. We show that the choice of an arbitrary mean function affects the reconstruction: a zero mean function leads to unphysical distance moduli and the best-fit ΛCDM to biased reconstructions. We propose to marginalize over a family of mean functions and over the hyperparameters to effectively remove their impact on the reconstructions. We further explore the validity and consistency of the results considering different kernel functions and show that our method is unbiased.

Tier 1cosmologygaussian-process

Pixel-to-pixel Translation of Solar Extreme-ultraviolet Images for DEMs by Fully Connected Networks

Park et al. (2023)

KASI: Eunsu Park (#1)

In this study, we suggest a pixel-to-pixel image translation method among similar types of filtergrams such as solar extreme-ultraviolet (EUV) images. For this, we consider a deep-learning model based on a fully connected network in which all pixels of solar EUV images are independent of one another. We use six-EUV-channel data from the Atmospheric Imaging Assembly (AIA) on board the Solar Dynamics Observatory (SDO), of which three channels (17.1, 19.3, and 21.1 nm) are used as the input data and the remaining three channels (9.4, 13.1, and 33.5 nm) as the target data. We apply our model to representative solar structures (coronal loops inside of the solar disk and above the limb, coronal bright point, and coronal hole) in SDO/AIA data and then determine differential emission measures (DEMs). Our results from this study are as follows. First, our model generates three EUV channels (9.4, 13.1, and 33.5 nm) with average correlation coefficient values of 0.78, 0.89, and 0.85, respectively. Second, our model generates the solar EUV data with no boundary effects and clearer identification of small structures when compared to a convolutional neural network-based deep-learning model. Third, the estimated DEMs from AI-generated data by our model are consistent with those using only SDO/AIA channel data. Fourth, for a region in the coronal hole, the estimated DEMs from AI-generated data by our model are more consistent with those from the 50 frames stacked SDO/AIA data than those from the single-frame SDO/AIA data.

Tier 3solar/heliophysicsclassical-ML

A multi-band study and exploration of the radio wave-γ-ray connection in 3C 84

Paraschos et al. (2023)

KASI: 김재영 (#3)

Total intensity variability light curves offer a unique insight into the ongoing debate about the launching mechanism of jets. For this work, we utilised the availability of radio and γ-ray light curves over a few decades of the radio source 3C 84 (NGC 1275). We calculated the multi-band time-lags between the flares identified in the light curves via discrete cross-correlation and Gaussian process regression. We find that the jet particle and magnetic field energy densities are in equipartition (kr-=-1.08-±-0.18). The jet apex is located z91.5-GHz-=-22-645-Rs (2---20-×-10-3-pc) upstream of the 3 mm radio core; at that position, the magnetic field amplitude is Bcore91.5-GHz-=-3-10 G. Our results are in good agreement with earlier studies that utilised very-long-baseline interferometry. Furthermore, we investigated the temporal relation between the ejection of radio and γ-ray flares. Our results are in favour of the γ-ray emission being associated with the radio emission. We are able to tentatively connect the ejection of features identified at 43 and 86 GHz to prominent γ-ray flares. Finally, we computed the multiplicity parameter λ and the Michel magnetisation σM, and find that they are consistent with a jet launched by the Blandford & Znajek (1977, MNRAS, 179, 433) mechanism.

Tier 1galaxiesgaussian-process

A sparse regression approach for populating dark matter haloes and subhaloes with galaxies

Icaza-Lizaola et al. (2023)

KASI: M. Icaza-Lizaola (#1)

We use sparse regression methods (SRMs) to build accurate and explainable models that predict the stellar mass of central and satellite galaxies as a function of properties of their host dark matter haloes. SRMs are machine learning algorithms that provide a framework for modelling the governing equations of a system from data. In contrast with other machine learning algorithms, the solutions of SRM methods are simple and depend on a relatively small set of adjustable parameters. We collect data from 35 459 galaxies from the EAGLE simulation using 19 redshift slices between z = 0 and z = 4 to parametrize the mass evolution of the host haloes. Using an appropriate formulation of input parameters, our methodology can model satellite and central haloes using a single predictive model that achieves the same accuracy as when predicted separately. This allows us to remove the somewhat arbitrary distinction between those two galaxy types and model them based only on their halo growth history. Our models can accurately reproduce the total galaxy stellar mass function and the stellar mass-dependent galaxy correlation functions (ξ(r)) of EAGLE. We show that our SRM model predictions of ξ(r) is competitive with those from subhalo abundance matching and might be comparable to results from extremely randomized trees. We suggest SRM as an encouraging approach for populating the haloes of dark matter only simulations with galaxies and for generating mock catalogues that can be used to explore galaxy evolution or analyse forthcoming large-scale structure surveys.

Tier 1galaxiesclassical-ML

Mass Production of 2021 KMTNet Microlensing Planets. III. Analysis of Three Giant Planets

Shin et al. (2023)

KASI: Kyu-Ha Hwang (#4), Chung, Sun-Ju (#8), Jung Youn Kil (#10) +8

We present the analysis of three more planets from the KMTNet 2021 microlensing season. KMT-2021-BLG0119Lb is a ∼6M_Jup planet orbiting an early M dwarf or a K dwarf, KMT-2021-BLG-0192Lb is a ∼2M_Nep planet orbiting an M dwarf, and KMT-2021-BLG-2294Lb is a ∼1.25M_Nep planet orbiting a very-low-mass M dwarf or a brown dwarf. These by-eye planet detections provide an important comparison sample to the sample selected with the AnomalyFinder algorithm, and in particular, KMT-2021-BLG-2294 is a case of a planet detected by eye but not by algorithm. KMT-2021-BLG-2294Lb is part of a population of microlensing planets around very-low-mass host stars that spans the full range of planet masses, in contrast to the planet population at <~ 0.1 au, which shows a strong preference for small planets.

Tier 1exoplanetsclassical-ML

2022 (20 papers)

Deep Learnin-based Fast Spectral Inversion of Hα and Ca II 8542 Line Spectra

Lee et al. (2022)

KASI: Eunsu Park (#3), Hannah Kwak (#5)

A multilayer spectral inversion (MLSI) model has recently been proposed for inferring the physical parameters of plasmas in the solar chromosphere from strong absorption lines taken by the Fast Imaging Solar Spectrograph (FISS). We apply a deep neural network (DNN) technique in order to produce the MLSI outputs with reduced computational costs. We train the model using two absorption lines, Hα and Ca II 8542 A, taken by FISS, and 13 physical parameters obtained from the application of MLSI to 49 raster scans (∼2,000,000 spectra). We use a fully connected network with skip connections and multi-branch architecture to avoid the problem of vanishing gradients and to improve the model’s performance. Our test shows that the DNN successfully reproduces the physical parameters for each line with high accuracy and a computing time of about 0.3-0.4 ms per line, which is about 250 times faster than the direct application of MLSI. We also confirm that the DNN reliably reproduces the temporal variations of the physical parameters generated by the MLSI inversion. By taking advantage of the high performance of the DNN, we plan to provide physical parameter maps for all the FISS observations, in order to understand the chromospheric plasma conditions in various solar features.

Tier 2solar/heliophysicsCNN

Systematic KMTNet Planetary Anomaly Search. VI. Complete Sample of 2018 Sub- prime-field Planets

Jung et al. (2022)

KASI: Jung Youn Kil (#1), Chung, Sun-Ju (#7), Kyu-Ha Hwang (#8) +7

We complete the analysis of all 2018 sub-prime-field microlensing planets identified by the KMTNet AnomalyFinder. Among the 9 previously unpublished events with clear planetary solutions, 6 are clearly planetary (OGLE-2018-BLG-0298, KMT-2018-BLG-0087, KMT-2018-BLG-0247, KMT-2018-BLG-0030, OGLE-2018-BLG-1119, and KMT-2018-BLG-2602), while the remaining 3 are ambiguous in nature. The above ordering of these events is made to facilitate grouping of their Bayesian estimates: the first two are lower-mass gas giants while the last four are Jovian-class planets; the first three most likely lie in the bulge, the last in the disk, and the remaining two are equally likely to be in either population. More specifically, these six planets have host masses ${M}_{\mathrm{host}}=({0.69}_{-0.30}^{+0.34},{0.10}_{-0.05}^{+0.14},{0.29}_{-0.14}^{+0.28},{0.51}_{-0.31}^{+0.43},{0.48}_{-0.28}^{+0.35},{0.66}_{-0.36}^{+0.42}){M}_{\odot }$, planet masses ${M}_{\mathrm{planet}}=({0.14}_{-0.06}^{+0.07},{0.23}_{-0.12}^{+0.32},{2.11}_{-1.04}^{+2.09},{1.45}_{-0.88}^{+1.23},{0.91}_{-0.52}^{+0.66},{1.15}_{-0.63}^{+0.73}){M}_{\mathrm{Jup}}$, and distances ${D}_{L}=({6.54}_{-1.23}^{+0.95},{7.02}_{-1.15}^{+1.03},{6.76}_{-1.24}^{+0.99},{6.48}_{-1.96}^{+1.28},{5.76}_{-2.48}^{+1.43},{4.31}_{-1.84}^{+1.97})\mathrm{kpc}$. In addition, there are 8 previously published sub-prime-field planets that were selected by the AnomalyFinder algorithm. Together with a companion paper on 2018 prime-field planets, this work lays the basis for comprehensive statistical studies. We carry out two such studies, one on caustic topologies and the other on the role of Gaia data. From the first, as expected, half (17/33) of the 2018 planets likely to enter the mass-ratio analysis have non-caustic-crossing anomalies. However, only 1 of the 5 noncaustic anomalies with planet-host mass ratio q < 10-3 was discovered by eye (compared to 7 of the 12 with q > 10-3), showing the importance of the semiautomated AnomalyFinder search. From the second, we find that Gaia has played a major role in the interpretation of 16% of the sample and a supplementary role in 6%.

Tier 2exoplanetsclassical-ML

Joint reconstructions of growth and expansion histories from stage-IV surveys with minimal assumptions: Dark energy beyond Lambda

Calderón et al. (2022)

KASI: Rodrigo Calderon (#1), Arman Shafieloo (#4)

Combining supernovae, baryon acoustic oscillations, and redshift-space distortions data from the next generation of (stage-IV) cosmological surveys, we aim to reconstruct the expansion history up to large redshifts using forward-modeling of f_{DE}(z)~ρ_{DE}(z)/ρ_{DE,0} with Gaussian processes (GP). In order to reconstruct cosmological quantities at high redshifts where few or no data are available, we adopt a new approach to GP which enforces the following minimal assumptions: (a) Our cosmology corresponds to a flat Friedmann-Lemaitre-Robertson-Walker universe; (b) An Einstein de Sitter (EdS) universe is obtained on large redshifts. This allows us to reconstruct the perturbations growth history from the reconstructed background expansion history. Assuming various DE models, we show the ability of our reconstruction method to differentiate them from ΛCDM at <~2σ.

Tier 1cosmologygaussian-process

Improved AI-generated Solar Farside Magnetograms by STEREO and SDO Data Sets and Their Release

Jeong et al. (2022)

KASI: 박은수 (#3), Ji-Hye Baek (#5)

Here we greatly improve artificial intelligence (AI)-generated solar farside magnetograms using data sets from the Solar Terrestrial Relations Observatory (STEREO) and Solar Dynamics Observatory (SDO). We modify our previous deep-learning model and configuration of input data sets to generate more realistic magnetograms than before. First, our model, which is called Pix2PixCC, uses updated objective functions, which include correlation coefficients (CCs) between the real and generated data. Second, we construct input data sets of our model: solar farside STEREO extreme-ultraviolet (EUV) observations together with nearest frontside SDO data pairs of EUV observations and magnetograms. We expect that the frontside data pairs provide historic information on magnetic field polarity distributions. We demonstrate that magnetic field distributions generated by our model are more consistent with the real ones than previously, in consideration of several metrics. The averaged pixel-to-pixel CC for full disk, active regions, and quiet regions between real and AI-generated magnetograms with 8 × 8 binning are 0.88, 0.91, and 0.70, respectively. Total unsigned magnetic flux and net magnetic flux of the AI-generated magnetograms are consistent with those of real ones for the test data sets. It is interesting to note that our farside magnetograms produce polar field strengths and magnetic field polarities consistent with those of nearby frontside magnetograms for solar cycles 24 and 25. Now we can monitor the temporal evolution of active regions using solar farside magnetograms by the model together with the frontside ones. Our AI-generated solar farside magnetograms are now publicly available at the Korean Data Center for SDO (http://sdo.kasi.re.kr).

Tier 3solar/heliophysicsCNN

Generation of Solar Coronal White-light Images from SDO/AIA EUV Images by Deep Learning

Lawrance et al. (2022)

KASI: Eunsu Park (#3)

Low coronal white-light observations are very important to understand low coronal features of the Sun, but they are rarely made. We generate Mauna Loa Solar Observatory (MLSO) K-coronagraph like white-light images from the Solar Dynamics Observatory/Atmospheric Imaging Assembly (SDO/AIA) EUV images using a deep learning model based on conditional generative adversarial networks. In this study, we used pairs of SDO/AIA EUV (171, 193, and 211 A) images and their corresponding MLSO K-coronagraph images between 1.11 and 1.25 solar radii from 2014 to 2019 (January to September) to train the model. For this we made seven (three using single channels and four using multiple channels) deep learning models for image translation. We evaluate the models by comparing the pairs of target white-light images and those of corresponding artificial intelligence (AI)-generated ones in October and November. Our results from the study are summarized as follows. First, the multiple channel AIA 193 and 211 A model is the best among the seven models in view of the correlation coefficient (CC = 0.938). Second, the major low coronal features like helmet streamers, pseudostreamers, and polar coronal holes are well identified in the AI-generated ones by this model. The positions and sizes of the polar coronal holes of the AIgenerated images are very consistent with those of the target ones. Third, from AI-generated images we successfully identified a few interesting solar eruptions such as major coronal mass ejections and jets. We hope that our model provides us with complementary data to study the low coronal features in white light, especially for nonobservable cases (during nighttime, poor atmospheric conditions, and instrumental maintenance).

Tier 3solar/heliophysicsGAN

OGLE-2017-BLG-1038: A Possible Brown-dwarf Binary Revealed by Spitzer Microlensing Parallax

Malpas et al. (2022)

KASI: Andrew P. Gould (#4), Cha, Sang-Mok (#15), Chung, Sun-Ju (#16) +10

We report the analysis of microlensing event OGLE-2017-BLG-1038, observed by the Optical Gravitational Lensing Experiment, Korean Microlensing Telescope Network, and Spitzer telescopes. The event is caused by a giant source star in the Galactic Bulge passing over a large resonant binary-lens caustic. The availability of space-based data allows the full set of physical parameters to be calculated. However, there exists an eightfold degeneracy in the parallax measurement. The four best solutions correspond to very-low-mass binaries near $( M_1 = 170^{+40}_{-50} M_{\rm J}$ {\rm and} $M_2 = 110^{+20}_{-30} M_{\rm J} )$, or well below $( M_1 = 22.5^{+0.7}_{-0.4} M_{\rm J} {\rm and} M_2 = 13.3^{+0.4}_{-0.3} M_{\rm J} )$ the boundary between stars and brown dwarfs. A conventional analysis, with scaled uncertainties for Spitzer data, implies a very-low-mass brown-dwarf binary lens at a distance of 2 kpc. Compensating for systematic Spitzer errors using a Gaussian process model suggests that a higher mass M-dwarf binary at 6 kpc is equally likely. A Bayesian comparison based on a galactic model favors the larger-mass solutions. We demonstrate how this degeneracy can be resolved within the next 10 years through infrared adaptive-optics imaging with a 40 m class telescope.

Tier 1othergaussian-process

A new approach to feature-based asteroid taxonomy in 3D color space I. SDSS photometric system

Roh et al. (2022)

KASI: Dong-Goo Roh (#1), MOON, HONG KYU (#2), Min-Su Shin (#3)

The taxonomic classification of asteroids has been mostly based on spectroscopic observations with wavelengths spanning from the visible (VIS) to the near-infrared (NIR). VIS-NIR spectra of ~2500 asteroids have been obtained since the 1970s; the Sloan Digital Sky Survey (SDSS) Moving Object Catalog 4 (MOC 4) was released with ~4 × 105 measurements of asteroid positions and colors in the early 2000s. A number of works then devised methods to classify these data within the framework of existing taxonomic systems. Some of these works, however, used 2D parameter space (e.g., gri slope vs. z-i color) that displayed a continuous distribution of clouds of data points resulting in boundaries that were artificially defined. We introduce here a more advanced method to classify asteroids based on existing systems. This approach is simply represented by a triplet of SDSS colors. The distributions and memberships of each taxonomic type are determined by machine learning methods in the form of both unsupervised and semi-supervised learning. We apply our scheme to MOC 4 calibrated with VIS-NIR reflectance spectra. We successfully separate seven different taxonomy classifications (C, D, K, L, S, V, and X) with which we have a sufficient number of spectroscopic datasets. We found the overlapping regions of taxonomic types in a 2D plane were separated with relatively clear boundaries in the 3D space newly defined in this work. Our scheme explicitly discriminates between different taxonomic types (e.g., K and X types), which is an improvement over existing systems. This new method for taxonomic classification has a great deal of scalability for asteroid research, such as space weathering in the S-complex, and the origin and evolution of asteroid families. We present the structure of the asteroid belt, and describe the orbital distribution based on our newly assigned taxonomic classifications. It is also possible to extend the methods presented here to other photometric systems, such as the Johnson-Cousins and LSST filter systems.

Tier 1otherclassical-ML

Systematic KMTNet planetary anomaly search V. Complete sample of 2018 prime-field

Gould et al. (2022)

KASI: Kyu-Ha Hwang (#5), Chung, Sun-Ju (#9), Jung Youn Kil (#10) +9

We complete the analysis of all 2018 prime-field microlensing planets identified by the Korea Microlensing Telescope Network (KMTNet) AnomalyFinder. Among the ten previously unpublished events with clear planetary solutions, eight are either unambiguously planetary or are very likely to be planetary in nature: OGLE-2018-BLG-1126, KMT-2018-BLG-2004, OGLE-2018-BLG-1647, OGLE-2018-BLG-1367, OGLE-2018-BLG-1544, OGLE-2018-BLG-0932, OGLE-2018-BLG-1212, and KMT-2018-BLG-2718. Combined with the four previously published new AnomalyFinder events and 12 previously published (or in preparation) planets that were discovered by eye, this makes a total of 24 2018 prime-field planets discovered or recovered by AnomalyFinder. Together with a paper in preparation on 2018 subprime planets, this work lays the basis for the first statistical analysis of the planet mass-ratio function based on planets identified in KMTNet data. By systematically applying the heuristic analysis to each event, we identified the small modification in their formalism that is needed to unify the so-called close-wide and inner-outer degeneracies.

Tier 3exoplanetsclassical-ML

KMT-2021-BLG-0171Lb and KMT-2021-BLG-1689Lb: two microlensing planets in the KMTNet high-cadence fields with followup observations

Yang et al. (2022)

KASI: Kyu-Ha Hwang (#5), Chung, Sun-Ju (#11), Jung Youn Kil (#13) +10

Follow-up observations of high-magnification gravitational microlensing events can fully exploit their intrinsic sensitivity to detect extrasolar planets, especially those with small mass ratios. To make followup observations more uniform and efficient, we develop a system, HighMagFinder, to automatically alert possible ongoing high-magnification events based on the real-time data from the Korea Microlensing Telescope Network (KMTNet). We started a new phase of follow-up observations with the help of HighMagFinder in 2021. Here we report the discovery of two planets in high-magnification microlensing events, KMT2021-BLG-0171 and KMT-2021-BLG-1689, which were identified by the HighMagFinder. We find that both events suffer the ‘central-resonant’ caustic degeneracy. The planet-host mass-ratio is q ∼ 4.7 × 10^-5 or q ∼ 2.2 × 10^-5 for KMT-2021-BLG-0171, and q ∼ 2.5 × 10^-4 or q ∼ 1.8 × 10^-4 for KMT-2021-BLG-1689. Together with two other events, four cases that suffer such degeneracy have been discovered in the 2021 season alone, indicating that the degenerate solutions may have been missed in some previous studies. We also propose a quantitative factor to weight the probability of each solution from the phase space. The resonant interpretations for the two events are disfavoured under this consideration. This factor can be included in future statistical studies to weight degenerate solutions.

Tier 1exoplanetsclassical-ML

Reconstruction of the regional total electron content maps over the Korean Peninsula using deep convolutional generative adversarial network and Poisson blending

Jeong et al. (2022)

KASI: Se-Heon Jeong (#1), Woo Kyoung Lee (#2), Jeong-Heon Kim (#5) +3

This study reconstructs total electron content (TEC) maps in the vicinity of the Korean Peninsula by employing a deep convolutional generative adversarial network and Poisson blending (DCGAN-PB). Our interest is to rebuild small-scale ionosphere structures on the TEC map in a local region where pronounced ionospheric structures, such as the equatorial ionization anomaly, are absent. The reconstructed regional TEC maps have a domain of 120°-135.5°E longitude and 25.5°-41°N latitude with 0.5° resolution. To achieve this, we first train a DCGAN model by using the International Reference Ionosphere (IRI)-based TEC maps from 2002 to 2019 (except for 2010 and 2014) as a training dataset. Next, the trained DCGAN model generates synthetic complete TEC maps from observation-based incomplete TEC maps. Final TEC maps are produced by blending of synthetic TEC maps with observed TEC data by PB. The performance of the DCGANPB model is evaluated by testing the regeneration of the masked TEC observations in 2010 (solar minimum) and 2014 (solar maximum). Our results show that a good correlation between the masked and model-generated TEC values is maintained even with a large percentage (~80%) of masking. The performance of the DCGAN-PB model is not sensitive to local time, solar activity, and magnetic activity. Thus, the DCGAN-PB model can reconstruct fine ionospheric structures in regions where observations are sparse and distinguishing ionospheric structures are absent. This model can contribute to near real-time monitoring of the ionosphere by immediately providing complete TEC maps.

Tier 3solar/heliophysicsCNNGAN

Systematic KMTNet planetary anomaly search. IV. Complete sample of 2019 prime-field

Zang et al. (2022)

KASI: Lee, Chung-Uk (#4), Chung, Sun-Ju (#11), Kyu-Ha Hwang (#12) +10

We report the complete statistical planetary sample from the prime fields (Γ ≥ 2 h-1) of the 2019 Korea Microlensing Telescope Network (KMTNet) microlensing survey. We develop the optimized KMTNet AnomalyFinder algorithm and apply it to the 2019 KMTNet prime fields. We find a total of 13 homogeneously selected planets and report the analysis of three planetary events, KMT-2019-BLG-(1042,1552,2974). The planet-host mass ratios, q, for the three planetary events are 6.34 × 10-4, 4.89 × 10-3, and 6.18 × 10-4, respectively. A Bayesian analysis indicates the three planets are all cold giant planets beyond the snow line of their host stars. The 13 planets are basically uniform in log-q over the range -5.0 < log-q < -1.5. This result suggests that the planets below qbreak = 1.7 × 10-4 proposed by the MOA-II survey may be more common than previously believed. This work is an early component of a large project to determine the KMTNet mass-ratio function, and the whole sample of 2016-2019 KMTNet events should contain about 120 planets.

Tier 3exoplanetsclassical-ML

Machine-guided exploration and calibration of astrophysical simulations

Oh et al. (2022)

KASI: Sungwook E. Hong (#5)

We apply a novel method with machine learning to calibrate sub-grid models within numerical simulation codes to achieve convergence with observations and between different codes. It utilizes active learning and neural density estimators. The hyper parameters of the machine are calibrated with a well-defined projectile motion problem. Then, using a set of 22 cosmological zoom simulations, we tune the parameters of a popular star formation and feedback model within Enzo to match observations. The parameters that are adjusted include the star formation efficiency, coupling of thermal energy from stellar feedback, and volume into which the energy is deposited. This number translates to a factor of more than three improvements over manual calibration. Despite using fewer simulations, we obtain a better agreement to the observed baryon makeup of a Milky Way (MW)-sized halo. Switching to a different strategy, we improve the consistency of the recommended parameters from the machine. Given the success of the calibration, we then apply the technique to reconcile metal transport between grid-based and particle-based simulation codes using an isolated galaxy. It is an improvement over manual exploration while hinting at a less-known relation between the diffusion coefficient and the metal mass in the halo region. The exploration and calibration of the parameters of the sub-grid models with a machine learning approach is concluded to be versatile and directly applicable to different problems.

Tier 3cosmologysimulation-based-inferenceclassical-ML

Millimeter Light Curves of Sagittarius A* Observed during the 2017 Event Horizon Telescope Campaign

Wielgus et al. (2022)

KASI: Do-Young Byun (#59), Taehyun Jung (#121), 김재영 (#127) +3

The Event Horizon Telescope (EHT) observed the compact radio source, Sagittarius A* (Sgr A*), in the Galactic Center on 2017 April 5-11 in the 1.3 mm wavelength band. At the same time, interferometric array data from the Atacama Large Millimeter/submillimeter Array and the Submillimeter Array were collected, providing Sgr A* light curves simultaneous with the EHT observations. These data sets, complementing the EHT very long baseline interferometry, are characterized by a cadence and signal-to-noise ratio previously unattainable for Sgr A* at millimeter wavelengths, and they allow for the investigation of source variability on timescales as short as a minute. While most of the light curves correspond to a low variability state of Sgr A*, the April 11 observations follow an X-ray flare and exhibit strongly enhanced variability. All of the light curves are consistent with a rednoise process, with a power spectral density (PSD) slope measured to be between -2 and -3 on timescales between 1 minute and several hours. Our results indicate a steepening of the PSD slope for timescales shorter than 0.3 hr. The spectral energy distribution is flat at 220 GHz, and there are no time lags between the 213 and 229 GHz frequency bands, suggesting low optical depth for the event horizon scale source. We characterize Sgr A*’s variability, highlighting the different behavior observed just after the X-ray flare, and use Gaussian process modeling to extract a decorrelation timescale and a PSD slope. We also investigate the systematic calibration uncertainties by analyzing data from independent data reduction pipelines.

Tier 1galaxiesgaussian-process

Charting galactic accelerations -- II. How to `learn’ accelerations in the solar neighbourhood

Naik et al. (2022)

KASI: J. An (#2)

Gravitational acceleration fields can be deduced from the collisionless Boltzmann equation, once the distribution function is known. This can be constructed via the method of normalizing flows from data sets of the positions and velocities of stars. Here, we consider application of this technique to the solar neighbourhood. We construct mock data from a linear superposition of multiple ‘quasi-isothermal’ distribution functions, representing stellar populations in the equilibrium Milky Way disc. We show that given a mock data set comprising a million stars within 1 kpc of the Sun, the underlying acceleration field can be measured with excellent, sub-per-cent level accuracy, even in the face of realistic errors and missing line-of-sight velocities. The effects of disequilibrium can lead to bias in the inferred acceleration field. This can be diagnosed by the presence of a phase space spiral, which can be extracted simply and cleanly from the learned distribution function. We carry out a comparison with two other popular methods of finding the local acceleration field (Jeans analysis and 1D distribution function fitting). We show our method most accurately measures accelerations from a given mock data set, particularly in the presence of disequilibria.

Tier 3galaxiesnormalizing-flowclassical-ML

Identifying Lensed Quasars and Measuring Their Time Delays from Unresolved Light Curves

Bag et al. (2022)

KASI: Satadru Bag (#1), Arman Shafieloo (#2)

Identifying multiply imaged quasars is challenging due to their low density in the sky and the limited angular resolution of wide field surveys. We show that multiply imaged quasars can be identified using unresolved light curves, without assuming a light curve template or any prior information. After describing our method, we show using simulations that it can attain high precision and recall when we consider high-quality data with negligible noise well below the variability of the light curves. As the noise level increases to that of the Zwicky Transient Facility (ZTF) telescope, we find that precision can remain close to 100% while recall drops to ~60%. We also consider some examples from the Time Delay Challenge 1 (TDC1) and demonstrate that the time delays can be accurately recovered from the joint light curve data in realistic observational scenarios. We further demonstrate our method by applying it to publicly available COSMOGRAIL data of the observed lensed quasar SDSS J1226-0006. We identify the system as a lensed quasar based on the unresolved light curve and estimate a time delay in good agreement with the one measured by COSMOGRAIL using the individual image light curves. The technique shows great potential to identify lensed quasars in wide field imaging surveys, especially the soon to be commissioned Vera Rubin Observatory.

Tier 1galaxiesclassical-ML

STag: Supernova Tagging and Classification

Davison et al. (2022)

KASI: William Davison (#1), David Parkinson (#2)

Supernovae classes have been defined phenomenologically, based on spectral features and time series data, since the specific details of the physics of the different explosions remain unrevealed. However, the number of these classes is increasing as objects with new features are observed, and the next generation of large surveys will only bring more variety to our attention. We apply the machine learning technique of multi-label classification to the spectra of supernovae. By measuring the probabilities of specific features or "tags" in the supernova spectra, we can compress the information from a specific object down to that suitable for a human or database scan, without the need to directly assign to a reductive "class". We use logistic regression to assign tag probabilities, and then a feed-forward neural network to filter the objects into the standard set of classes, based solely on the tag probabilities. We present STag, a software package that can compute these tag probabilities and make spectral classifications.

Tier 2transients/time-domainclassical-MLCNN

The Galaxy Replacement Technique (GRT): A New Approach to Study Tidal Stripping and Formation of Intracluster Light in a Cosmological Context

Chun et al. (2022)

KASI: Kyungwon Chun (#1), Jihye Shin (#2), Jongwan Ko (#4) +1

We introduce the Galaxy Replacement Technique (GRT) that allows us to model tidal stripping of galaxies with very high mass (mstar = 5.4 × 104 Me h-1) and high spatial resolution (10 pc h-1), in a fully cosmological context, using an efficient and fast technique. The technique works by replacing multiple low-resolution dark-matter (DM) halos in the base cosmological simulation with high-resolution models, including a DM halo and stellar disk. We apply the method to follow the hierarchical buildup of a cluster since redshift ∼8 to now, through the hierarchical accretion of galaxies, individually or in substructures such as galaxy groups. We find we can successfully reproduce the observed total stellar masses of observed clusters since redshift ∼1. The high resolution allows us to accurately resolve the tidal stripping process and well describe the formation of ultralow surface brightness features in the cluster (μV < 32 mag arcsec-2) such as the intracluster light (ICL), shells, and tidal streams. We measure the evolution of the fraction of light in the ICL and brightest cluster galaxy using several different methods. While their broad response to the cluster-mass growth history is similar, the methods show systematic differences, meaning we must be careful when comparing studies that use distinct methods. The GRT represents a powerful new tool for studying tidal effects on galaxies and exploring the formation channels of the ICL in a fully cosmological context and with large samples of simulated groups and clusters.

Tier 3galaxiesother

Systematic KMTNet Planetary Anomaly Search. II. Six New q<2×10^-4 Mass-ratio Planets

Hwang et al. (2022)

KASI: Kyu-Ha Hwang (#1), Chung, Sun-Ju (#9), Jung Youn Kil (#11) +10

We apply the automated AnomalyFinder algorithm of Paper I to 2018-2019 light curves from the ~13 deg^2 covered by the six KMTNet prime fields, with cadences Γ >= 2 hr^-1. We find a total of 11 planets with mass ratios q < 2 × 10^-4, including 6 newly discovered planets, 1 planet that was reported in Paper I, and recovery of 4 previously discovered planets. One of the new planets, OGLE-2018-BLG-0977Lb, is in a planetary caustic event, while the other five (OGLE-2018-BLG-0506Lb, OGLE-2018-BLG-0516Lb, OGLE-2019-BLG-1492Lb, KMT-2019-BLG-0253, and KMT-2019-BLG-0953) are revealed by a “dip” in the light curve as the source crosses the host-planet axis on the opposite side of the planet. These subtle signals were missed in previous by-eye searches. The planet-host separations (scaled to the Einstein radius), s, and planet-host mass ratios, q, are, respectively, (s, q × 10^5) = (0.88, 4.1), (0.96 ± 0.10, 8.3), (0.94 ± 0.07, 13), (0.97 ± 0.07, 18), (0.97 ± 0.04, 4.1), and (0.74, 18), where the “ ± ” indicates a discrete degeneracy. The 11 planets are spread out over the range -5 < log q < -3.7. Together with the two planets previously reported with q ∼ 10^-5 from the 2018-2019 nonprime KMT fields, this result suggests that planets toward the bottom of this mass-ratio range may be more common than previously believed.

Tier 1exoplanetsclassical-ML

Estimation of Photometric Redshifts. II. Identification of Out-of-distribution Data with Neural Networks

Lee & Shin (2022)

KASI: 이준구 (#1), Min-Su Shin (#2)

In this study, we propose a three-stage training approach of neural networks for both photometric redshift estimation of galaxies and detection of out-of-distribution (OOD) objects. Our approach comprises supervised and unsupervised learning, which enables using unlabeled (UL) data for OOD detection in training the networks. Employing the UL data, which is the data set most similar to the real-world data, ensures a reliable usage of the trained model in practice. We quantitatively assess the model performance of photometric redshift estimation and OOD detection using in-distribution (ID) galaxies and labeled OOD (LOOD) samples such as stars and quasars. Our model successfully produces photometric redshifts matched with spectroscopic redshifts for the ID samples and identifies well the LOOD objects with more than 98% accuracy. Although quantitative assessment with the UL samples is impracticable owing to the lack of labels and spectroscopic redshifts, we also find that our model successfully estimates reasonable photometric redshifts for ID-like UL samples and filter OOD-like UL objects.

Tier 2galaxiesclassical-ML

TESS Eclipsing Binary Stars. I. Short-cadence Observations of 4584 Eclipsing Binaries in Sectors 1-26

Prša et al. (2022)

KASI: Lee, Jae Woo (#25)

In this paper we present a catalog of 4584 eclipsing binaries observed during the first two years (26 sectors) of the TESS survey. We discuss selection criteria for eclipsing binary candidates, detection of hitherto unknown eclipsing systems, determination of the ephemerides, the validation and triage process, and the derivation of heuristic estimates for the ephemerides. Instead of keeping to the widely used discrete classes, we propose a binary star morphology classification based on a dimensionality reduction algorithm. Finally, we present statistical properties of the sample, we qualitatively estimate completeness, and we discuss the results. The work presented here is organized and performed within the TESS Eclipsing Binary Working Group, an open group of professional and citizen scientists; we conclude by describing ongoing work and future goals for the group. The catalog is available from http://tessEBs.villanova.edu and from MAST.

Tier 1starsclassical-ML

2021 (22 papers)

Weak-lensing Mass Reconstruction of Galaxy Clusters with a Convolutional Neural Network

Hong et al. (2021)

KASI: Sungwook E. Hong (#1)

We introduce a novel method for reconstructing the projected matter distributions of galaxy clusters with weak-lensing (WL) data based on a convolutional neural network (CNN). Training data sets are generated with ray-tracing through cosmological simulations. We control the noise level of the galaxy shear catalog such that it mimics the typical properties of the existing ground-based WL observations of galaxy clusters. We find that the mass reconstruction by our multilayered CNN with the architecture of alternating convolution and trans-convolution filters significantly outperforms the traditional reconstruction methods. The CNN method provides better pixel-to-pixel correlations with the truth, restores more accurate positions of the mass peaks, and more efficiently suppresses artifacts near the field edges. In addition, the CNN mass reconstruction lifts the mass-sheet degeneracy when applied to our projected cluster mass estimation from sufficiently large fields. This implies that this CNN algorithm can be used to measure the cluster masses in a model-independent way for future wide-field WL surveys.

Tier 2galaxiesCNN

Disruption of Hierarchical Clustering in the Vela OB2 Complex and the Cluster Pair Collinder 135 and UBC 7 with Gaia EDR3: Evidence of Supernova Quenching

Pang et al. (2021)

KASI: Jongsuk Hong (#4)

We identify hierarchical structures in the Vela OB2 complex and the cluster pair Collinder 135 and UBC 7 with Gaia EDR3 using the neural network machine-learning algorithm StarGO. Five second-level substructures are disentangled in Vela OB2, which are referred to as Huluwa 1 (Gamma Velorum), Huluwa 2, Huluwa 3, Huluwa 4, and Huluwa 5. For the first time, Collinder 135 and UBC 7 are simultaneously identified as constituent clusters of the air with minimal manual intervention. We propose an alternative scenario in which Huluwa 1~5 have originated from sequential star formation. The older clusters Huluwa 1~3, with an age of 10-22 Myr, generated stellar feedback to cause turbulence that fostered the formation of the younger-generation Huluwa 4~5 (7~20 Myr). A supernova explosion located inside the Vela IRAS shell quenched star formation in Huluwa 4~5 and rapidly expelled the remaining gas from the clusters. This resulted in global mass stratification across the shell, which is confirmed by the regression discontinuity method. The stellar mass in the lower rim of the shell is 0.32 ± 0.14 Me higher than in the upper rim. Local, cluster-scale mass segregation is observed in the lowest-mass cluster Huluwa 5. Huluwa 1~5 (in Vela OB2) are experiencing significant expansion, while the cluster pair suffers from moderate expansion. The velocity dispersions suggest that all five groups (including Huluwa 1A and Huluwa 1B) in Vela OB2 and the cluster pair are supervirial and are undergoing disruption, and also that Huluwa 1A and Huluwa 1B may be a coeval young cluster pair. N-body simulations predict that Huluwa 1~5 in Vela OB2 and the cluster pair will continue to expand in the future 100 Myr and eventually dissolve

Tier 1starsclassical-ML

Estimation of Photometric Redshifts. I. Machine-learning Inference for Pan-STARRS1 Galaxies Using Neural Networks

Lee & Shin (2021)

KASI: Joongoo Lee (#1), Min-Su Shin (#2)

We present a new machine-learning model for estimating photometric redshifts with improved accuracy for galaxies in Pan-STARRS1 data release 1. Depending on the estimation range of redshifts, this model based on neural networks can handle the difficulty for inferring photometric redshifts. Moreover, to reduce bias induced by the new model's ability to deal with estimation difficulty, it exploits the power of ensemble learning. We extensively examine the mapping between input features and target redshift spaces to which the model is validly applicable to discover the strength and weaknesses of the trained model. Because our trained model is well calibrated, our model produces reliable confidence information about objects with non-catastrophic estimation. While our model is highly accurate for most test examples residing in the input space, where training samples are densely populated, its accuracy quickly diminishes for sparse samples and unobserved objects (i.e., unseen samples) in training. We report that out-of-distribution (OOD) samples for our model contain both physically OOD objects (i.e., stars and quasars) and galaxies with observed properties not represented by training data. The code for our model is available at https://github.com/GooLee0123/MBRNN for other uses of the model and retraining the model with different data.

Tier 2galaxiesCNN

Solar Event Detection Using Deep-Learning-Based Object Detection Methods

Baek et al. (2021)

KASI: Ji-Hye Baek (#1), Sujin Kim (#2), Seonghwan Choi (#3) +2

Research on the detection of the solar events has been conducted over many years. Recently, deep learning and data-driven approaches to solar physics have been applied to solar event recognition. In this study, we present solar event detection using deep-learning-based object detection methods for real-time space weather monitoring. First, we construct a new object detection dataset using imaging data obtained by the Solar Dynamics Observatory with bounding boxes as labels for three representative features: coronal holes, sunspots, and prominences. Second, we train two representative object detection models, the Single Shot MultiBox Detector (SSD), and the Faster Region-based Convolutional Neural Network (R-CNN) with the new dataset. The results show that both models perform similarly well for coronal hole and sunspot detection. For prominence detection, the SSD and Faster R-CNN exhibited relatively low performance. This study demonstrates that deep learning-based object detection can successfully detect multiple types of solar events, and it may be extended to detect other solar events. In addition, we provide the dataset for further achievements of object detection studies in solar physics.

Tier 2solar/heliophysicsCNN

Turbulent Properties in Star-forming Molecular Clouds Down to the Sonic Scale. II. Investigating the Relation between Turbulence and Star-forming Environments in Molecular Clouds

Yun et al. (2021)

KASI: Neal J. Evans, II (#3), Yunhee Choi (#10), Minho Choi (#13) +2

We investigate the effect of star formation on turbulence in the Orion A and Ophiuchus clouds using principal component analysis (PCA). We measure the properties of turbulence by applying PCA on the spectral maps in 13CO, C18O, HCO+ J = 1-0, and CS J = 2-1. First, the scaling relations derived from PCA of the 13CO maps show that the velocity difference (δv) for a given spatial scale (L) is the highest in the integral-shaped filament (ISF) and L1688, where the most active star formation occurs in the two clouds. The δv increases with the number density and total bolometric luminosity of the protostars in the subregions. Second, in the ISF and L1688 regions, the δv of C18O, HCO+, and CS are generally higher than that of 13CO, which implies that the dense gas is more turbulent than the diffuse gas in the star-forming regions; stars form in dense gas, and dynamical activities associated with star formation, such as jets and outflows, can provide energy into the surrounding gas to enhance turbulent motions.

Tier 1ISMclassical-ML

Development of a Deep Learning Model for Inversion of Rotational Coronagraphic Images Into 3D Electron Density

Jang et al. (2021)

KASI: Soojeong Jang (#1), Ryun-Young Kwon (#2), Yeon-Han Kim (#7)

We present, for the first time, a deep learning model that returns the three-dimensional (3D) coronal electron density from coronagraphic images. The intensity of coronagraphic observations arises from the Thomson scattering of photospheric light by the coronal electrons. We use MHD numerical simulations to obtain realistic 3D electron density and construct error-free training sets consisting of input (observation) and target (electron density) images. In the training sets, the input images are directly synthesized from the target 3D electron density by applying the Thomson scattering theory. The input and target images are in the form of latitude-longitude maps given at a radius, often referred to as synoptic maps. Using synoptic maps reduces a tomographic method to an image translation problem. We use pix2pixHD, one of the well-established supervised image translation methods and develop models for six selected heights: 2.0, 2.2, 2.5, 4.0, 6.0, and 12.0 solar radii. All six models have similar performance and the mean absolute percent error of the generated density images is less than 7% with respect to the ground-truth simulated data sets.

Tier 2solar/heliophysicsGAN

Systematic KMTNet Planetary Anomaly Search. I. OGLE-2019-BLG-1053Lb, a Buried Terrestrial Planet

Zang et al. (2021)

KASI: Kyu-Ha Hwang (#2), Sun-Ju Chung (#12), Youn Kil Jung (#14) +10

In order to exhume the buried signatures of “missing planetary caustics” in Korea Microlensing Telescope Network (KMTNet) data, we conducted a systematic anomaly search of the residuals from point-source point-lens fits, based on a modified version of the KMTNet EventFinder algorithm. This search revealed the lowest-mass ratio planetary caustic to date in the microlensing event OGLE-2019-BLG-1053, for which the planetary signal had not been noticed before. The planetary system has a planet-host mass ratio of q = (1.25 ± 0.13) × 10^-5. A Bayesian analysis yielded estimates of the mass of the host star, M_host = 0.61^+0.29_-0.24 M_Sun, the mass of its planet, M_planet = 2.48^+1.19_-0.98 M_Earth, the projected planet-host separation, a = 3.4^+0.5_-0.5 au, and the lens distance, D_L = 6.8^+0.6_-0.9 kpc. The discovery of this very-low-mass-ratio planet illustrates the utility of our method and opens a new window for a large and homogeneous sample to study the microlensing planet-host mass ratio function down to q ∼ 10^-5.

Tier 1exoplanetsclassical-ML

Potential of Regional Ionosphere Prediction Using a Long Short-Term Memory Deep-Learning Algorithm Specialized for Geomagnetic Storm Period

Kim et al. (2021)

KASI: Jeong-Heon Kim (#1), Young-Sil Kwak (#2), Se-Heon Jeong (#5)

In our previous study (Moon et al., 2020), we developed a Long Short Term Memory (LSTM) deep-learning model for geomagnetic quiet days (LSTM-quiet) to perform effective long-term predictions for the regional ionosphere. However, their model could not predict geomagnetic storm days effectively at all. This study developed an LSTM model suitable for geomagnetic storms using the new training data set and re-designing input parameters and hyperparameters. We collected 131 days of geomagnetic storm cases from 1 January 2009 to 31 December 2019, provided by the Japan Meteorological Agency's Kakioka Magnetic Observatory, and obtained the IMF Bz, Dst, Kp, and AE indices related to the geomagnetic storm corresponding to each storm date from the OMNI database. These indices and F2 parameters (foF2 and hmF2) of Jeju ionosonde (33.43˚N, 126.30˚E) were used as input parameters for the LSTM model. To test and verify the predictive performance and the usability of the LSTM model for geomagnetic storms developed in this manner, we created and diagnosed the 0.5, 1, 2, 3, 6, 12, and 24-hour predictive LSTM models. According to the results of this study, the LSTM storm model for 24-hour developed in this study achieved a predictive performance during the three geomagnetic storms about 32% (10%), 34% (17%), and 37% (5%) better in RMSE of foF2 (hmF2) than the LSTM quiet model (Moon et al., 2020), SAMI2, and IRI-2016 models. We propose that the short-term predictions of less than 3 hours are sufficiently competitive than other traditional ionospheric models. Thus, this study suggests that our model can be used for short-term prediction and monitoring of the regional mid-latitude ionosphere.

Tier 2solar/heliophysicsRNN/LSTM

Hubble diagram at higher redshifts: model independent calibration of quasars

Li et al. (2021)

KASI: Ryan E. Keeley (#2), Arman Shafieloo (#3)

In this paper, we present a model-independent approach to calibrate the largest quasar sample. Calibrating quasar samples is essentially constraining the parameters of the linear relation between the log of the ultraviolet (UV) and X-ray luminosities. This calibration allows quasars to be used as standardized candles. There is a strong correlation between the parameters characterizing the quasar luminosity relation and the cosmological distances inferred from using quasars as standardized candles. We break this degeneracy by using Gaussian process regression to model-independently reconstruct the expansion history of the Universe from the latest type Ia supernova observations. Using the calibrated quasar data set, we further reconstruct the expansion history up to redshift of z ∼ 7.5. Finally, we test the consistency between the calibrated quasar sample and the standard Lambda cold dark matter (--CDM) model based on the posterior probability distribution of the GP hyperparameters. Our results show that the quasar sample is in good agreement with the standard --CDM model in the redshift range of the supernova, despite the 2-3σ significant deviations taking place at higher redshifts. Fitting the standard --CDM model to the calibrated quasar sample, we obtain a high value of the matter density parameter -- = 0.382+0.045, which is marginally consistent with the constraints from m -0.042 other cosmological observations.

Tier 1cosmologygaussian-process

Active galactic nuclei catalog from the AKARI NEP-Wide field

Poliszczuk et al. (2021)

KASI: Eunbin Kim (#15)

Context. The north ecliptic pole (NEP) field provides a unique set of panchromatic data that are well suited for active galactic nuclei (AGN) studies. The selection of AGN candidates is often based on mid-infrared (MIR) measurements. Such methods, despite their effectiveness, strongly reduce the breadth of resulting catalogs due to the MIR detection condition. Modern machine learning techniques can solve this problem by finding similar selection criteria using only optical and near-infrared (NIR) data. Aims. The aim of this study is to create a reliable AGN candidates catalog from the NEP field using a combination of optical SUBARU/HSC and NIR AKARI/IRC data and, consequently, to develop an efficient alternative for the MIR-based AKARI/IRC selection technique. Methods. We tested set of supervised machine learning algorithms for the purposes of carrying out an efficient process for AGN selection. The best models were compiled into a majority voting scheme, which used the most popular classification results to produce the final AGN catalog. An additional analysis of the catalog properties was performed as a spectral energy distribution fitting via the CIGALE software. Results. The obtained catalog of 465 AGN candidates (out of 33 119 objects) is characterized by 73% purity and 64% completeness. This new classification demonstrates a suitable consistency with the MIR-based selection. Moreover, 76% of the obtained catalog can be found solely using the new method due to the lack of MIR detection for most of the new AGN candidates. The training data, codes, and final catalog are available via the github repository. The final catalog of AGN candidates is also available via the CDS service. Conclusions. The new selection methods presented in this paper are proven to be a better alternative for the MIR color AGN selection. Machine learning techniques not only show similar effectiveness, but also involve less demanding optical and NIR observations, substantially increasing the extent of available data samples.

Tier 1galaxiesclassical-ML

Charting galactic accelerations: when and how to extract a unique potential from the distribution function

An et al. (2021)

KASI: J. An (#1)

The advent of data sets of stars in the Milky Way with 6D phase-space information makes it possible to construct empirically the distribution function (DF). Here, we show that the accelerations can be uniquely determined from the DF using the collisionless Boltzmann equation, providing the Hessian determinant of the DF with respect to the velocities is non-vanishing. We illustrate this procedure and requirement with some analytic examples. Methods to extract the potential from data sets of discrete positions and velocities of stars are then discussed. Following Green & Ting, we advocate the use of normalizing flows on a sample of observed phase-space positions to obtain a differentiable approximation of the DF. To then derive gravitational accelerations, we outline a semi-analytic method involving direct solutions of the overconstrained linear equations provided by the collisionless Boltzmann equation. Testing our algorithm on mock data sets derived from isotropic and anisotropic Hernquist models, we obtain excellent accuracies even with added noise. Our method represents a new, flexible, and robust means of extracting the underlying gravitational accelerations from snapshots of 6D stellar kinematics of an equilibrium system.

Tier 3galaxiesnormalizing-flowclassical-ML

Identification of Lensed Gravitational Waves with Deep Learning

Kim et al. (2021)

KASI: Kyungmin Kim (#1), Joongoo Lee (#2)

Similar to light, gravitational waves (GWs) can be lensed. Such lensing phenomena can magnify the waves, create multiple images observable as repeated events, and superpose several waveforms together, inducing potentially discernible patterns on the waves. In particular, when the lens is small, -105M⊙ , it can produce lensed images with time delays shorter than the typical gravitational-wave signal length that conspire together to form ``beating patterns''. We present a proof-of-principle study utilizing deep learning for identification of such a lensing signature. We bring the excellence of state-of-the-art deep learning models at recognizing foreground objects from background noises to identifying lensed GWs from noise present spectrograms. We assume the lens mass is around 103M⊙ -- 105M⊙ , which can produce the order of millisecond time delays between two images of lensed GWs. We discuss the feasibility of distinguishing lensed GWs from unlensed ones and estimating physical and lensing parameters. Suggested method may be of interest to the study of more complicated lensing configurations for which we do not have accurate waveform templates.

Tier 2transients/time-domainCNN

Future radio continuum cosmology clustering surveys

Asorey & Parkinson (2021)

KASI: Jacobo Asorey (#1), David Parkinson (#2)

The use of continuum emission radio galaxies as cosmological tracers of the large-scale structure will soon move into a new phase. Upcoming surveys from the Australian Square Kilometre Array Pathfinder (ASKAP), MeerKAT, and the Square Kilometre Array project (SKA) will survey the entire available sky down to an ∼100μJy flux limit, increasing the number of detected extra-galactic radio sources by several orders of magnitude. External data and machine learning algorithms will also enable some low-resolution radial selection (photometric redshift binning) of the sample, increasing the cosmological utility of the sample observed. In this paper, we discuss the flux limit required to detect enough galaxies to decrease the shot-noise term in the error to be 10 per-cent of the total. We show how future surveys of this type will be limited by available technology. The confusion generated by the intrinsic sizes of galaxies may have the consequence that surveys of this type eventually reach a hard flux limit of ∼100-nJy, as is predicted by the current modelling of AGN sizes by simulations such as the Tiered Radio Extragalactic Continuum Simulation (T-RECS). Finally, when considering the multitracer approach, where galaxies are split by type to measure some bias ratio, we find that there are not enough AGN present to achieve a reasonable level of shot noise for this kind of measurement.

Tier 1cosmologyclassical-ML

Deep learning model on gravitational waveforms in merging and ringdown phases of binary black hole coalescences

Lee et al. (2021)

KASI: Joongoo Lee (#1), Kyungmin Kim (#3), Hyung Mok Lee (#7)

The waveform templates of the matched filtering-based gravitational-wave search ought to cover wide range of parameters for the prosperous detection. Numerical relativity (NR) has been widely accepted as the most accurate method for modeling the waveforms. Still, it is well known that NR typically requires a tremendous amount of computational costs. In this paper, we demonstrate a proof-of-concept of a novel deterministic deep learning (DL) architecture that can generate gravitational waveforms from the merger and ringdown phases of the non-spinning binary black hole coalescence. Our model takes O(1) seconds for generating approximately 1500 waveforms with a 99.9% match on average to one of the state-of-the-art waveform approximants, the effective-one-body. We also perform matched filtering with the DL-waveforms and find that the waveforms can recover the event time of the injected gravitational-wave signals.

Tier 3cosmologyCNN

Revealing the Local Cosmic Web from Galaxies by Deep Learning

Hong et al. (2021)

KASI: Sungwook E. Hong (#1), Ho Seong Hwang (#3)

A total of 80% of the matter in the universe is in the form of dark matter that composes the skeleton of the large-scale structure called the cosmic web. As the cosmic web dictates the motion of all matter in galaxies and intergalactic media through gravity, knowing the distribution of dark matter is essential for studying the large-scale structure. However, the cosmic web's detailed structure is unknown because it is dominated by dark matter and warm-hot intergalactic media, both of which are hard to trace. Here we show that we can reconstruct the cosmic web from the galaxy distribution using the convolutional-neural-network-based deep-learning algorithm. We find the mapping between the position and velocity of galaxies and the cosmic web using the results of the state-of-the-art cosmological galaxy simulations of Illustris-TNG. We confirm the mapping by applying it to the EAGLE simulation. Finally, using the local galaxy sample from Cosmicflows-3, we find the dark matter map in the local universe. We anticipate that the local dark matter map will illuminate the studies of the nature of dark matter and the formation and evolution of the Local Group. High-resolution simulations and precise distance measurements to local galaxies will improve the accuracy of the dark matter map.

Tier 2cosmologyCNN

Operational Dst index prediction model based on combination of artificial neural network and empirical model

Park et al. (2021)

KASI: Jaejin Lee (#2), Jong-Kil Lee (#4), Yukinaga Miyashita (#6) +5

In this paper, an operational Dst index prediction model is developed by combining empirical and Artificial Neural Network (ANN) models. ANN algorithms are widely used to predict space weather conditions. While they require a large amount of data for machine learning, large-scale geomagnetic storms have not occurred sufficiently for the last 20 years, Advanced Composition Explorer (ACE) and Deep Space Climate Observatory (DSCOVR) mission operation period. Conversely, the empirical models are based on numerical equations derived from human intuition and are therefore applicable to extrapolate for large storms. In this study, we distinguish between Coronal Mass Ejection (CME) driven and Corotating Interaction Region (CIR) driven storms, estimate the minimum Dst values, and derive an equation for describing the recovery phase. The combined Korea Astronomy and Space Science Institute (KASI) Dst Prediction (KDP) model achieved better performance contrasted to ANN model only. This model could be used practically for space weather operation by extending prediction time to 24 h and updating the model output every hour.

Tier 1solar/heliophysicsclassical-ML

Manually scaling ionograms measured by Icheon and Jeju ionosondes over a 2-year period (2017-2018)

Lee et al. (2021)

KASI: Young-Sil Kwak (#4)

From 1966 until 2009, the ionosonde at the Anyang station has been monitoring the ionosphere over the Korea, Peninsula, and the ionosondes at the Icheon (37.14° N, 127.54° E) and the Jeju (33.43° N, 126.30° E) stations have continued their monitoring missions since 2010 and 2009, respectively. An ionosonde transmits HF radio waves vertically and records the waves reflected from the ionosphere on an ionogram. Ionospheric parameters, such as foF2 (the critical frequency of the F2 layer) and hmF2 (the peak height of the F2 layer), can be extracted from an ionogram using an automatic scaling program. However, an automatic scaling program may result in absurdly inaccurate ionospheric parameters. In this study, we manually scaled total 33,679 ionograms measured at the Icheon and the Jeju ionosondes for 2 years (2017 and 2018). By adopting the manual scaling method from the standard handbook of Wakai et al. (Radio Research Laboratory, Ministry of Posts and Telecommunication, Japan, 1987), we extracted five ionospheric parameters, namely foF2, hmF2, foE (the critical frequency of the E layer), hmE (the peak height of the E layer), and foEs (the critical frequency of the sporadic E layer), and compared them with the automatically scaled values obtained by using ARTIST5002 (the built-in automatic scaling program). We found that the automatic scaling software misinterpreted about 23% and 36% the ionograms taken at the Icheon and the Jeju stations, respectively, thus inaccurately characterizing both the ionospheric E and F layers. We classified those misinterpreted ionograms into eight cases: five cases about the F layer and three cases about the E layer. We discuss the reliability of automatically scaled parameters, emphasizing the need for manual scaling when scientifically investigating the ionosphere with ionosonde data. The manually scaled parameters can be used as “true” values when a deep learning model is trained to interpret measured ionograms.

Tier 1solar/heliophysicsother

Construction of a far-ultraviolet all-sky map from an incomplete survey: application of a deep learning algorithm

Jo et al. (2021)

KASI: Young-Soo Jo (#1), Kwang-Il Seon (#6)

We constructed a far-ultraviolet (FUV) all-sky map based on observations from the Far Ultraviolet Imaging Spectrograph (FIMS) aboard the Korean microsatellite Science and Technology SATellite-1. For the ∼20 percent of the sky not covered by FIMS observations, predictions from a deep artificial neural network were used. Seven data sets were chosen for input parameters, including five all-sky maps of Hα, E(B - V), N(H-I), and two X-ray bands, with Galactic longitudes and latitudes. 70 percent of the pixels of the observed FIMS data set were randomly selected for training as target parameters and the remaining 30 percent were used for validation. A simple four-layer neural network architecture, which consisted of three convolution layers and a dense layer at the end, was adopted, with an individual activation function for each convolution layer; each convolution layer was followed by a dropout layer. The predicted FUV intensities exhibited good agreement with Galaxy Evolution Explorer observations made in a similar FUV wavelength band for high Galactic latitudes. As a sample application of the constructed map, a dust scattering simulation was conducted with model optical parameters and a Galactic dust model for a region that included observed and predicted pixels. Overall, FUV intensities in the observed and predicted regions were reproduced well.

Tier 2ISMCNN

An active galactic nucleus recognition model based on deep neural network

Chen et al. (2021)

KASI: Ho Seong Hwang (#16), Eunbin Kim (#17)

To understand the cosmic accretion history of supermassive black holes, separating the radiation from active galactic nuclei (AGNs) and star-forming galaxies (SFGs) is critical. However, a reliable solution on photometrically recognizing AGNs still remains unsolved. In this work, we present a novel AGN recognition method based on Deep Neural Network (Neural Net; NN). The main goals of this work are (i) to test if the AGN recognition problem in the North Ecliptic Pole Wide (NEPW) field could be solved by NN; (ii) to show that NN exhibits an improvement in the performance compared with the traditional, standard spectral energy distribution (SED) fitting method in our testing samples; and (iii) to publicly release a reliable AGN/SFG catalogue to the astronomical community using the best available NEPW data, and propose a better method that helps future researchers plan an advanced NEPW data base. Finally, according to our experimental result, the NN recognition accuracy is around 80.29 per cent-85.15 per cent, with AGN completeness around 85.42 per cent-88.53 per cent and SFG completeness around 81.17 per cent-85.09 per cent.

Tier 2galaxiesCNN

Reconstructing the Universe: Testing the Mutual Consistency of the Pantheon and SDSS/eBOSS BAO Data Sets with Gaussian Processes

Keeley et al. (2021)

KASI: Ryan E. Keeley (#1), Arman Shafieloo (#2), Hanwool Koo (#5)

We test the mutual consistency between the baryon acoustic oscillation measurements from the eBOSS SDSS final release and the Pantheon supernova compilation in a model-independent fashion using Gaussian process regression. We also test their joint consistency with the ΛCDM model in a model-independent fashion. We also use Gaussian process regression to reconstruct the expansion history that is preferred by these two data sets. While this methodology finds no significant preference for model flexibility beyond ΛCDM, we are able to generate a number of reconstructed expansion histories that fit the data better than the best-fit ΛCDM model. These example expansion histories may point the way toward modifications to ΛCDM. We also constrain the parameters Ωk and H0rd both with ΛCDM and with Gaussian process regression. We find that H0rd = 10,030 ± 130 km s-1 and Ωk = 0.05 ± 0.10 for ΛCDM and that H0rd = 10,040 ± 140 km s-1 and Ωk = 0.02 ± 0.20 for the Gaussian process case.

Tier 1cosmologygaussian-process

The SEDIGISM survey: molecular clouds in the inner Galaxy

Duarte-Cabral et al. (2021)

KASI: M.-Y. Lee (#21)

We use the 13CO (2-1) emission from the SEDIGISM (Structure, Excitation, and Dynamics of the Inner Galactic InterStellar Medium) high-resolution spectral-line survey of the inner Galaxy, to extract the molecular cloud population with a large dynamic range in spatial scales, using the Spectral Clustering for Interstellar Molecular Emission Segmentation (SCIMES) algorithm. This work compiles a cloud catalogue with a total of 10 663 molecular clouds, 10 300 of which we were able to assign distances and compute physical properties. We study some of the global properties of clouds using a science sample, consisting of 6664 well-resolved sources and for which the distance estimates are reliable. In particular, we compare the scaling relations retrieved from SEDIGISM to those of other surveys, and we explore the properties of clouds with and without high-mass star formation. Our results suggest that there is no single global property of a cloud that determines its ability to form massive stars, although we find combined trends of increasing mass, size, surface density, and velocity dispersion for the sub-sample of clouds with ongoing high-mass star formation. We then isolate the most extreme clouds in the SEDIGISM sample (i.e. clouds in the tails of the distributions) to look at their overall Galactic distribution, in search for hints of environmental effects. We find that, for most properties, the Galactic distribution of the most extreme clouds is only marginally different to that of the global cloud population. The Galactic distribution of the largest clouds, the turbulent clouds and the high-mass star-forming clouds are those that deviate most significantly from the global cloud population. We also find that the least dynamically active clouds (with low velocity dispersion or low virial parameter) are situated further afield, mostly in the least populated areas. However, we suspect that part of these trends may be affected by some observational biases (such as completeness and survey limitations), and thus require further follow up work in order to be confirmed.

Tier 1ISMclassical-ML

Cloud structures in M 17 SWex : Possible cloud-cloud collision

Kinoshita et al. (2021)

KASI: Kee-Tae Kim (#10), Hyunwoo KANG (#11)

Using wide-field 13 CO (J = 1-0) data taken with the Nobeyama 45 m telescope, we investigate cloud structures of the infrared dark cloud complex in M 17 with Spectral Clustering for Interstellar Molecular Emission Segmentation. In total, we identified 118 clouds that include 11 large clouds with radii larger than 1 pc. The clouds are mainly distributed in the two representative velocity ranges of 10-20 km s-1 and 30-40 km s-1 . By comparing this with the ATLASGAL catalog, we found that the majority of the 13 CO clouds with 10-20 km s-1 and 30-40 km s-1 are likely located at distances of 2 kpc (Sagit- tarius arm) and 3 kpc (Scutum arm), respectively. Analyzing the spatial configuration of the identified clouds and their velocity structures, we attempt to reveal the origin of the cloud structure in this region. Here we discuss three possibilities: (1) overlapping with different velocities, (2) cloud oscillation, and (3) cloud-cloud collision. In the position velocity diagrams, we found spatially extended faint emission between ∼20 km s-1 and ∼35 km s-1 , which is mainly distributed in the spatially overlapped areas of the clouds. Additionally, the cloud complex system is unlikely to be gravitationally bound. We also found that in some areas where clouds with different velocities overlapped, the mag- netic field orientation changes abruptly. The distribution of the diffuse emission in the position-position-velocity space and the bending magnetic fields appear to favor the cloud-cloud collision scenario compared to other scenarios. In the cloud-cloud collision scenario, we propose that two ∼35 km s-1 foreground clouds are colliding with clouds at ∼20 km s-1 with a relative velocity of 15 km s-1 . These clouds may be substructures of two larger clouds having velocities of ∼35 km s-1 ( 103 M ) and ∼20 km s-1 ( 104 M ), respectively.

Tier 1ISMclassical-ML