CLM-020
Claim
无条件合成 SOTA:LDM 在 CelebA-HQ 256² FID 5.11,并在 FFHQ/LSUN 上与最强 GAN/扩散竞争。
| id | CLM-020 |
|---|---|
| type | claim |
| source | SRC-2112.10752 |
| scope | 无条件合成, 256x256; CelebA-HQ (SOTA 论断所在); FFHQ / LSUN-Churches / LSUN-Bedrooms (competitive, 部分低于最强 GAN/ADM); LDM-4 500 步 / LDM-8 200 步等 |
| epistemic | supported |
| lifecycle | RECONCILED |
| created | 2026-09-03 |
| updated | 2026-09-03 |
Proposition
LDM 在无条件图像合成上达到新的 SOTA(CelebA-HQ 256²:500 DDIM 步 FID 5.11,优于此前 likelihood-based 与 GAN),并在 FFHQ / LSUN-Churches / LSUN-Bedrooms 上与最强 GAN 和扩散模型竞争(Bedrooms 上接近 ADM,参数量减半、训练资源少 4 倍);且在 Precision 与 Recall 上系统性优于 GAN——印证其 likelihood(mode-covering)目标相对对抗训练的分布覆盖优势。
Evidence refs
Notes
主表(ms_tables.tex, FID/Precision/Recall):CelebA-HQ — VQGAN+T 10.2、PGGAN 8.0、LSGM 7.22、UDM 7.16;LDM-4 5.11(prec 0.72/recall 0.49)。FFHQ — LDM-4 4.98(prec 0.73/recall 0.50)。LSUN-Churches — StyleGAN2 3.86;LDM-8 4.02(prec 0.64/recall 0.52)。LSUN-Bedrooms — StyleGAN2 2.35、ADM 1.90、ProjectedGAN 1.52;LDM-4 2.95(prec 0.66/recall 0.48)。§4.2 原文:"we report a new state-of-the-art FID of 5.11, outperforming previous likelihood-based models as well as GANs"、"LDMs consistently improve upon GAN-based methods in Precision and Recall, thus confirming the advantages of their mode-covering likelihood-based training objective over adversarial approaches"。注意 SOTA 措辞限于 CelebA-HQ;Bedrooms 的 ADM 1.90 优于 LDM 2.95,"close to ADM despite half its parameters and 4× less resources"。
引用(2)
- EVI-020 无条件合成表:LDM 对 GAN/扩散基线的 FID/Precision/Recall(支撑 CLM-020)。
- SRC-2112.10752 Latent Diffusion Model 原始论文(Rombach et al.);Stable Diffusion 学术源头,DiT 所引『latent diffusion』。
被引用(4)
- EVI-020 无条件合成表:LDM 对 GAN/扩散基线的 FID/Precision/Recall(支撑 CLM-020)。
- EVI-023 §10 Limitations + Fig. firststagecomparison:AE 重建瓶颈与迭代采样局限(支撑 CLM-023)。
- FRM-2112.10752 LDM 五层重建:感知压缩 AE + 潜空间扩散的两段分解,换取训练/采样成本大降;多任务多分辨率(256²/512²→~1024²)。
- VER-020 无条件合成核验:CelebA-HQ 5.11 SOTA、FFHQ 4.98、LSUN-Bedrooms 2.95 vs ADM 1.90 与无条件表一致,SOTA 限 CelebA-HQ、P/R 逐数据集。supported。