KASI's AI Research Capability
A comparative benchmark against Max Planck (MPIA/MPA/MPE), NAOJ, NOIRLab & NRAO — and a recommended adoption path.
KASI AI/ML output 2021–2026 · authoritative 114-paper set (prompt v2, 100% full-text) ·
interactive companion to kasi_ai_benchmark_report.md
1 Executive Summary
KASI produces a respectable volume of ML-assisted astronomy but is structurally behind its peers on every institutional dimension of AI transformation — and the gap is in organization, not talent.
- Volume at the field baseline; depth is not. 114 papers (6.7%) use ML, but only 2.4% use deep learning — below the ~5% astro-ph baseline. No paper reaches frontier level 4–5.
- Real but under-leveraged assets. Five multi-year programs and unique owned facilities (KMTNet, KVN, KASINet, SPHEREx/7DS/K-DRIFT).
- The defining weakness is the facility–algorithm asymmetry. Only 39% of ML papers are KASI-led; on KASI's own KMTNet the algorithms are external.
- Peers have institutionalized what KASI does ad hoc — dedicated units, AI institutes, compute, fellows.
- The opening is national and immediate — Korea's 2026 sovereign-AI buildout, with research institutes eligible.
2 Methodology & Data Provenance
Every KASI SCI paper 2021–2026 (n=1694) was classified by an LLM pipeline (prompt v2) for AI usage, technique, maturity tier (0–4) and whether it develops a method; the 114 genuine-ML papers form a full-text knowledge base (100% coverage) over which 15 flagships were individually re-read and scored 1–5 for frontier-ness. Two systematic errors were corrected before these numbers: 20 of 22 "simulation-based-inference" tags were classical forward-modeling (mock catalogs / MCMC), not neural SBI; and individual name-based mislabels (e.g. GPCAL tagged "gaussian-process"). Peer-institute claims come from a fan-out web-research pipeline with 3-vote adversarial verification. The authoritative figures moved <3% from the original 137-set estimate, so every conclusion is unchanged.
3 KASI's AI Profile
The five (+one nascent) real programs
Sustained, multi-year AI(-adjacent) lines with identifiable owners — as opposed to one-off papers.
Untrained Gaussian-process reconstruction of the expansion/growth history, running from 2021 quasar Hubble diagrams to the DESI DR2 extended dark-energy flagship analysis (2025). KASI's most internationally visible AI-adjacent program, with genuine methodological ownership.
GP · reached DESI DR2KASI-developed cGAN variant (correlation-coefficient loss). Farside magnetograms released as AISFM 3.0, a public data product on KASI's KDC-SDO; plus 3D coronal densities, synthetic EUV, Carrington-event reconstruction. Caveat: the flagship line is KHU-led (Moon group); KASI's equity runs through one person + the hosting infrastructure.
GAN · public data productDCGAN TEC-map completion (2022) → ConvLSTM 24-hr forecasting (2024) → storm-weighted ConvLSTM (2026), on KASI's own KASINet GNSS network. The clearest fully-in-house program end-to-end — but architectures run 5–8 years behind SOTA and one evaluation is circular. Pre-operational.
ConvLSTM · fully in-houseCNN mass mapping beating Kaiser–Squires, validated on real Coma data, with a 2025 LSST-oriented sequel; V-Net density/velocity reconstruction (2023) extended to the Zone of Avoidance (2026); anchored by an earlier cosmic-web CNN (2021). The clearest evidence of self-contained KASI deep-learning capability.
CNN · in-house lineageWorld-leading planet yield, and the one place KASI AI-adjacent software runs in production (HighMagFinder live every 3 hours; AnomalyFinder every season). But the anomaly-search algorithm is Tsinghua/OSU's, LensNet's neural net is MIT's, and AnomalyFinder is a classical grid search — KASI supplies facility, data, labels, operations.
production · algorithms externalKASI-led CNNs on merger waveforms (2021PhRvD.103l3023L, tier 3) and lensed-GW identification (2021ApJ...915..119K, tier 2). Both 2021-vintage and not obviously continued, but they show in-house DL reach beyond the imaging/cosmology programs.
CNN · not yet a coherent lineCross-cutting structural observations
- Facility–algorithm asymmetry. KASI owns world-class facilities and data, provides labels/vetting/operations, while Tsinghua, OSU/MPIA and MIT supply the algorithms and lead the papers. KASI is a data provider to other people's AI.
- Fragmentation. Parallel groups do not share methods — GP cosmology vs CNN cosmology, solar generation vs ionospheric forecasting, space-weather LSTMs vs LensNet. No shared tooling, benchmarks, or ML engineering layer is visible.
- Simulation circularity. Training on and validating against simulations or self-generated data recurs; the sim-to-real gap is addressed seriously in only two papers (domain generalization 2024; GECKO transfer learning 2025).
- Missing modern practice. Uncertainty quantification is nearly absent; interpretability appears in ~2 papers; several flagships lack method baselines entirely.
- Understated operational assets. HighMagFinder runs live on KMTNet; AISFM 3.0 is a public product; the coronal-density GAN targets near-real-time MHD replacement; GECKO's CNN ran in a real O4 campaign. KASI is closer to operational AI than its zero tier-4 count suggests.
- Positive culture signal. A distinctive habit of failure-mode auditing and honest self-critique (OOD overconfidence analysis, storm-forecast failure analysis, AnomalyFinder's documented by-eye miss) — a real asset for trustworthy operational ML.
4 How Frontier Is KASI's Best Work?
5 Peer-Institute Comparison
What each peer actually does
MPIA runs a dedicated, institution-level Data Science department (head Ivelina Momcheva, ~5 staff) embedded in Euclid, Roman, JWST, LSST and Gaia, with UQ and reproducible software as explicit missions. MPA is a methodological leader in field-level and neural simulation-based inference (score-based diffusion posteriors, neural emulators). MPE applies deep learning to its own eROSITA X-ray survey in-house (14/18 authors on a representative paper). Society-wide, MP-AIX pairs every AI PhD with an ML advisor and a domain advisor.
Couples owned compute with owned surveys: the Center for Computational Astrophysics runs ATERUI III (HPE Cray XD2000, 1.99 PF, from Dec 2025), and on top of it sits the Dark Emulator (ML emulator over cosmological parameter space) and in-house CNN classification of ~560,000 Subaru/HSC galaxy images at 97.5% accuracy. NAOJ's AI is anchored to its own instrument and its own compute — the two anchors KASI lacks.
Signature strength is AI in production survey infrastructure: ANTARES, a machine-learning alert broker filtering and classifying Rubin/LSST alerts in real time, plus a 2026 end-to-end Rubin follow-up ecosystem (ANTARES → GOATS → AEON → DRAGONS) built for millions of alerts per night. This is precisely the tier-4 layer at which KASI has zero papers. Also a CosmicAI partner.
Driven by a hard requirement — the ngVLA needs ~50 petaFLOPS (10,000× current ALMA) — NRAO targets it with AI for calibration/imaging/analysis. Founding partner of the $20M NSF–Simons CosmicAI institute (with NOIRLab, UVA, Utah, UCLA), hosting named AI fellows and building a TACC-hosted open AI platform with an LLM-style 'CosmicAI Assistant' and an AI co-pilot for astronomical data.
By subfield
| Subfield | KASI | Peers |
|---|---|---|
| Cosmology / LSS | GP reconstruction (DESI DR2); level-3 DL field reconstruction | MPA: frontier field-level/neural SBI · MPE: eROSITA DL clusters · NAOJ: Dark Emulator |
| Galaxies | Photo-z + OOD (code released); weak-lensing CNN | NAOJ: HSC CNN morphology at 560k scale · MPIA: survey-pipeline ML |
| Solar / space weather | Strongest ownership (Pix2PixCC, TEC forecasting) | Globally crowded: NASA+IBM Surya foundation model, NOAA SWPC, NASA FDL |
| Exoplanets / microlensing | World-leading yield, but algorithms external | MPIA co-owns the algorithm running on KASI data |
| Transients / time-domain | Scattered (GECKO CNN emerging) | NOIRLab: ANTARES broker in production at Rubin scale |
| Radio / VLBI | GPCAL (classical, strong) | NRAO: AI imaging/calibration program for ngVLA scale |
By methodology
| Method | KASI (2021–2026) | Field / peers |
|---|---|---|
| Classical ML | 63 papers — the workhorse | Everywhere, treated as solved plumbing |
| CNN | 31 papers, 2 real lineages | Standard; peers at larger data scale (560k HSC images) |
| Gaussian processes | Genuine depth (DESI DR2) | KASI is competitive here |
| GAN (cGAN era) | 5 papers, real ownership | Field moved to diffusion — KASI has none |
| Neural SBI / field-level | 1 proof-of-concept | MPA / frontier standard for Stage-IV cosmology |
| Transformers | 4 exploratory papers | Foundation-model substrate everywhere else |
| Foundation models | 0 | AION-1: one frozen encoder over 200M+ observations replaces bespoke pipelines |
| LLM agents / RL ops | 0 / 0 | Immature field-wide — an open early-bet lane, taken by CosmicAI |
6 Institution-Wide Efforts & the Korean Opening
The 114-paper corpus shows no evidence of any institution-wide AI initiative at KASI — no dedicated unit, hiring line, training program, compute strategy, or strategy document. The programs above are bottom-up group efforts. (Limitation: internal programs that never surface in publications are invisible to this corpus-based analysis.)
7 Strengths & Weaknesses
Strengths
- Unique data engines — KMTNet, KVN, KASINet, SPHEREx/7DS/K-DRIFT stakes.
- Real solar/space-weather franchise (Pix2PixCC, AISFM 3.0) — though globally crowded.
- Internationally visible GP-cosmology culture (reached DESI DR2).
- Self-contained DL capability exists (photo-z + OOD, weak-lensing CNN lineage).
- Operational temperament + honest failure-mode auditing.
Weaknesses
- No frontier presence — zero papers rated 4–5; no foundation/diffusion/neural-SBI/LLM work.
- Facility–algorithm asymmetry — only 39% of ML papers KASI-led; KMTNet algorithms are external.
- No institutional machinery — no AI unit, compute strategy, fellowship, or strategy doc.
- Stalled trajectory — AI share flat-to-declining (2025 lowest at 5.4%); DL flat ~2.4%.
- Rigor gaps that block deployment — UQ nearly absent, missing baselines, sim-circularity.
8 Recommended Adoption Path
- Stand up a ~5-FTE Astronomical Data Science unit (MPIA model): shared tooling, benchmarks, methods clinics connecting the existing programs.
- Secure National AI Computing Center / MSIT sovereign-GPU allocation before it hardens; volunteer KASI as the astronomy pilot for the 2026 co-scientist framework.
- Fix the measurement baseline — track genuine-ML vs classical, DL share, KASI-led share, tier-4 count annually.
- Quick wins: retrofit UQ + method baselines into TEC/denoiser/ViT lines; benchmark the solar GAN line against NASA/IBM Surya rather than training cGANs from scratch.
- Put KASI algorithms on KASI facilities: an in-house learned KMTNet vetter trained on KASI's own labels.
- Make upcoming surveys AI-native from day one (K-DRIFT, SPHEREx Ices, 7DS); own the broker/pipeline layer. Target: first 2 tier-4 deployments by 2028.
- Build the people pipeline: MP-AIX-style co-advised PhDs/fellowships with KAIST/SNU/UNIST.
- Bridge GP culture to neural SBI, aimed at 7DS/SPHEREx inference problems.
- Join, don't build, the foundation-model layer: contribute KMTNet/SPHEREx/7DS data to AION-class consortia, then fine-tune.
- Bid to anchor the astronomy node of a Korean national AI-for-science institute.
- Adopt LLM/agent tooling internally early — an early-bet lane aligned with the national co-scientist framework.
9 Sources
Peer institutes & frontier (web, 3-vote adversarially verified)
- MPIA Data Science department
- Max Planck Society MP-AIX doctoral program
- MPE eROSITA deep learning (A&A 2024)
- NRAO CosmicAI partnership
- NRAO AI fellows, ngVLA, TACC platform (AAS 247)
- CosmicAI launch, $20M, partners (UT Austin)
- NOIRLab ANTARES + Rubin follow-up ecosystem
- NAOJ HSC CNN morphology (560k, 97.5%)
- NAOJ ATERUI III (CfCA)
- NAOJ Dark Emulator
- MPA field-level SBI (Perimeter 2026)
- AION-1 foundation model
- Field DL baseline ~5% of astro-ph
- Korea national AI GPU initiative (OECD.AI)
- Korea 2026 buildout + co-scientist (Korea Herald)
- Korea–NVIDIA sovereign AI infrastructure
Solar/space-weather competitive landscape (single-pass verified)
KASI knowledge base
- Full text of all 114 AI/ML papers + per-paper/topic wiki (kasi-kb)
- Quantitative profile: analysis/data/enriched.parquet (prompt v2)
- Classification diff v1→v2: analysis/data/classification_v1_v2_diff.md
kasi_ai_benchmark_report.md · profile charts recomputed from KASI's
enriched publication dataset · narrative content curated from deep-research analysis